A significant Microsoft 365 outage that disrupted services for thousands of users on July 23 was caused by a bug in the company’s automated network maintenance request system, according to Microsoft’s preliminary incident review. The Microsoft 365 outage began at 10:44 AM ET and primarily impacted customers accessing services through network infrastructure connected to the West US Azure region.
Downdetector recorded 2,403 outage reports by 11:11 AM ET, a dramatic increase from the normal baseline of 29 reports. SharePoint accounted for 78% of user complaints, followed by Excel at 11% and the Microsoft 365 Admin Center at 6%. Microsoft tracked the incident under ID MO1437424.
Which Services Were Affected by the Outage?
The disruption impacted numerous Microsoft 365 services simultaneously. OneDrive users experienced intermittent access problems, while SharePoint Online users encountered “Something went wrong” error messages. Microsoft Teams chat functionality was degraded, with images failing to load properly. The Microsoft 365 Admin Center loaded slowly or failed to load entirely for administrators.
Additional affected services included Power Automate flows that wouldn’t load, Copilot Chat experiencing delays and failures during queries, and Microsoft Loop pages that users couldn’t open. Microsoft Defender customers faced delays receiving responses from Microsoft Defender Experts, while investigations and remediation actions through Threat Explorer and Advanced Hunting could fail. Other impacted services included Fabric, Power BI, Power Apps, Copilot Studio, Windows 365, and numerous Azure services.
What Caused the Maintenance System Failure?
Microsoft revealed that the outage originated during routine device maintenance in its West US Azure region, where specific network paths were being isolated. The company’s maintenance process is designed to convert requests into system-readable instructions and verify that at least one of two redundant paths remains healthy before proceeding.
However, a bug in the request conversion system incorrectly marked additional network devices as part of the maintenance event. This error caused IP routes to be removed from more devices than intended between Microsoft’s West US datacenter and its wide-area network. The removed routes disrupted network traffic entering or leaving the West US region, though traffic remaining entirely within the region was unaffected.
How Did Microsoft Restore Services?
Microsoft initially attempted to mitigate the problem by rerouting traffic through alternate network paths, which provided partial relief. Engineers traced the issue to route removals in the West US datacenter linked to the maintenance activity. The company initiated a rollback of the maintenance change at 1:45 PM ET, completing it at 2:26 PM ET. Microsoft confirmed full service recovery by 3:41 PM ET.
The company is now conducting a comprehensive internal review focused on safety checks and automated processes used to execute maintenance requests. A final Post Incident Review will be published within 14 days after completing the investigation.
Source: BleepingComputer