Microsoft: West US Azure Outage Was Caused By Network Maintenance Error

A configuration issue during planned maintenance removed network routes and disrupted traffic entering and leaving the region.

Datacenter networking servers

Key Takeaways:

  • A bug in Microsoft’s maintenance process incorrectly affected more network devices than intended.
  • The outage disrupted connectivity to Azure and other Microsoft cloud services in the West US region.
  • Microsoft plans to improve validation and change-management controls to prevent similar incidents.

Microsoft has disclosed that a major outage in its West US Azure region (California) was caused by an error during routine network maintenance. The incident disrupted access to Azure and several other Microsoft cloud services for customers in the affected region.

In its preliminary post-incident review report, Microsoft mentioned that the outage began at 14:44 UTC on July 23rd (07:44 AM Pacific Time), which left customers unable to access critical services for nearly five hours. It occurred when Microsoft engineers started scheduled maintenance on network infrastructure intended to isolate specific devices without affecting service availability.

“In this case, a bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended,” Microsoft explained. “The routes were removed between our datacenter and wide-area network, impacting traffic entering or exiting the region.”

According to Microsoft, applications and services running entirely within the West US region generally continued to operate as expected. However, traffic entering or leaving the region was severely impacted, which caused widespread connectivity problems for customers.

How did Microsoft restore services?

Microsoft detected unusual network behaviour very quickly and directed the service, networking, and incident-response teams to investigate the problem. Microsoft’s engineers analyzed routing anomalies, packet loss, and recent infrastructure changes and identified that the disruption was caused by a recent fiber maintenance activity.

Microsoft then started reversing the faulty configuration changes, and network functionality was largely restored by approximately 18:26 UTC. The company confirmed that all affected Azure services had recovered by 19:41 UTC.

Microsoft outlines steps to prevent similar Azure outages

Microsoft acknowledged that safeguards designed to ensure maintenance would be non-disruptive failed because the request-conversion system incorrectly expanded the list of affected devices. In response, the company plans to enhance its validation and change-management procedures to prevent similar disruptions in the future.

Microsoft also emphasized the importance of deploying mission-critical applications across multiple cloud regions. The company noted that a multi-region strategy can improve resilience and reduce the impact of outages affecting a single region.