Skip to main content
CybersecurityInfrastructure

Microsoft Outage Exposes Flaw in Automated Maintenance Process

Technicians inspect rows of computer servers and networking equipment in a brightly-lit server room.

"We will be preforming a full analysis focusing on safety checks, automated maintenance request change process, and more as we progress through our post mitigation internal retrospective," explained Microsoft.

How the outage unfolded on July 23

The disruption began at 10:44 AM ET on Thursday, July 23, when engineers detected large-scale failures affecting Microsoft 365 customers who accessed services through network infrastructure tied to Microsoft’s West US Azure region. By 11:11 AM ET Downdetector had recorded 2,403 outage reports — far above its normal baseline of 29 — with SharePoint accounting for 78% of complaints, Excel 11%, and the Microsoft 365 Admin Center 6%.

Services confirmed affected under incident ID MO1437424

Microsoft tracked the disruption under incident ID MO1437424 and said multiple Microsoft 365 and Azure services experienced problems. Microsoft 365 impacts included intermittent access to OneDrive; "Something went wrong" errors in SharePoint Online; degraded Microsoft Teams chat functionality (including images not loading); slow or unavailable Microsoft 365 Admin Center pages; Power Automate flows failing to load; intermittent delays or failures in Copilot Chat; and Microsoft Loop pages failing to open or load. Other affected products named by Microsoft included Fabric and Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender, where some customers saw delays receiving responses from Microsoft Defender Experts and failures in workflows triggered through Threat Explorer and Advanced Hunting.

Root cause Microsoft identified: a maintenance-request conversion bug

In a preliminary Post Incident Review for the Azure incident, Microsoft said the outage was triggered during routine device maintenance in its West US Azure region. The company described a maintenance process that converts human-readable requests into system-readable instructions and that ordinarily checks that at least one of two redundant paths remains healthy before proceeding. A bug in that request-conversion system, Microsoft said, mistakenly marked additional network devices as part of the maintenance event.

As a result, IP routes were removed from more devices than intended between a West US datacenter and Microsoft's wide-area network. Microsoft said the removed routes disrupted traffic entering or leaving the West US region, while traffic that remained entirely within the region was not affected. The issue initially manifested to engineers as large-scale route churn in Microsoft's WAN; investigators later traced the removals to the datacenter and linked them to the recent maintenance activity.

Mitigation, rollback, and recovery timeline

Microsoft initially attempted mitigation by rerouting traffic through alternate network paths; the company reported that this helped some customers but that many services remained affected. After identifying the recent networking change as the cause, Microsoft initiated a rollback of the maintenance change at 1:45 PM ET and completed the reversion at 2:26 PM ET. The rollback restored the affected network infrastructure and allowed Microsoft 365 services to recover; Microsoft said some Azure services continued to recover after the fix, and that all affected services had fully recovered by 3:41 PM ET.

Before the cause was known, Microsoft advised customers they might need to review business continuity and disaster recovery plans and take actions appropriate for their environments. The company also said it will publish a final Post Incident Review after completing its investigation, which Microsoft noted is usually published within 14 days.

What this means for technologists, enterprise admins, and end users

  • Technologists and security teams: The incident highlights the role of automated maintenance systems in network availability — Microsoft singled out safety checks and the automated maintenance-request change process for a full analysis, which is likely to inform post-incident controls inside the company.
  • Enterprise administrators and procurement leaders: The outage underlined dependency on regional networking paths; Microsoft’s warning to review business continuity and disaster recovery plans was explicit, and affected admins will need to evaluate routing, failover, and regional access strategies for customers tied to West US infrastructure.
  • End users and operators of Defender services: Some Microsoft Defender customers experienced delayed responses and failures in investigations and remediation actions through specific Defender tools such as Threat Explorer and Advanced Hunting, showing that defensive workflows can be disrupted when platform networking is impaired.

Microsoft’s immediate timeline — detection at 10:44 AM ET, rollback initiated at 1:45 PM ET, rollback completed at 2:26 PM ET, and full recovery by 3:41 PM ET — frames the company’s next steps. The preliminary review pins the outage on an automated conversion bug and removed IP routes; the promised internal retrospective and the forthcoming final Post Incident Review will be the next factual milestones to watch.

Original reporting: https://www.bleepingcomputer.com/news/microsoft/microsoft-blames-massive-microsoft-365-outage-on-maintenance-bug/