The 10 Biggest Cloud Outages Of 2026 (So Far)
Microsoft, AWS, Google, Verizon and Cloudflare have experienced the biggest cloud outages of 2026 so far.
Cloud outages remain a plague on modern IT environments, disrupting access to the AI tools and cloud services businesses increasingly rely on.
A report published earlier this year by Cisco’s Splunk division puts downtime costs for the Global 2000 up 50 percent over the past two years, with companies losing an average $300 million a year to unplanned outages. Companies on average see 3.4 percent stock price drops after a single incident.
Network and IT environment issues account for 43 percent of downtime events, according to the Splunk report. About 30 percent of events are cybersecurity-driven, and 24 percent come from application or infrastructure failures.
[RELATED: 2026 Year So Far]
Ranking The Biggest Cloud Outages Of 2026 So Far
Every minute of downtime can cost $15,000, or about $900,000 an hour, according to Splunk. Downtimes appear to be getting more expensive over time across most direct cost categories, with lost revenue of $95 million nearly twice as high compared with 2024.
During one of the outages that made this list, Corey Kirkendoll, CEO of Allen, Texas-based Microsoft solution provider 5K Technical Services—a member of CRN’s MSP 500—told CRN in an interview the outage served as a lesson as to why MSPs and customers need to keep a human in the loop as they explore what workloads and workflows they can offload onto cutting-edge AI tools that leverage the cloud.
Incidents like these are why 5K goes through business continuity planning with customers, Kirkendoll said. “Treat AI tools like any other critical SaaS and never let a single vendor become a single point of failure,” he said.
Read on for details on the biggest cloud outages so far this year, ranked from least impactful to most impactful based on a mix of factors including number of users affected, duration of outage and the influence of the companies involved on the channel.
And be sure to check out CRN’s other 2026 Year So Far articles, including The 7 Hottest Data Storage Startups Of 2026 (So Far) and The 10 Coolest AI Observability And Governance Tools Of 2026 (So Far).
U.S. Virgin Islands Power Grid Failure
The U.S. Virgin Islands saw traffic for local provider VI Powernet drop to near zero on March 24 starting around 12:15 local time, 16:15 UTC, and recovery starting at 14:45 local time.
The disruption came from underground cable damage causing a power outage in St. Croix and St. Thomas plus power generation loss at the Richmond Power Plant, according to a Cloudflare blog post.
Traffic from St. Thomas fell by 60 percent and St. Croix by 40 percent, avoiding a complete outage due to the presence of other providers.
The Virgin Islands has faced an increasingly fragile power grid with aging generators, equipment shortages and years of deferred maintenance, The Associated Press said in an article in May. The U.S. federal government has invested more than $100 million into the territory, whose residents already pay about twice the U.S. average for electricity. About 40,000 people live on the main island, St. Thomas.
TalkTalk Broadband Outage
U.K. broadband provider TalkTalk experienced service disruption on March 25.
TalkTalk traffic dropped about 50 percent compared with the prior week starting around 7:00 local time, according to a Cloudflare report on the incident. TalkTalk restored service by about 8:15.
Broadband price comparison website Uswitch reported that more than 5,000 customers saw issues during the event. The website said that 22.4 million people experienced a broadband outage over the prior 12 months, with a total of 238.7 million hours of connectivity lost nationwide.
The provider published to X at 2:43 a.m. Pacific on March 25 acknowledging that customers had issues connecting to their Wi-Fi, calling on those users to refresh their browsers and reboot routers if the issues persisted.
“We’re sorry for any inconvenience this has caused,” the company said.
Iran Internet Shutdowns, AWS Data Center Damage
Cloudflare called Iran’s internet shutdown one of the longest sustained internet disruptions in recent years and described damage to an Amazon Web Services data center during the conflict as an “unusual” disruption.
The vendor called the shutdowns examples of governments weaponizing internet access for political control and the drone strikes “an unprecedented escalation as active military conflict directly and physically damaged major cloud infrastructure, with disastrous consequences for the websites and applications hosted there.”
On Jan. 8, Iran saw IPv6 traffic go to zero, according to Cloudflare, likely due to filtering. That shutdown started at 20:00 local time, 16:30 UTC. Traffic stayed near zero except for brief restorations on Jan. 21 and 25, with a more aggressive recovery starting Jan. 27.
The vendor credited a second internet shutdown Feb. 28 to aggressive filtering and whitelists and white SIM cards restricting access to approved websites as military strikes on Iran escalated. A traffic drop started around 10:30 local time, or 7:00 UTC, with traffic falling to less than 1 percent of previous levels.
AWS data infrastructure in the Middle East has been attacked during the conflict, with reports that the Iranian Islamic Revolutionary Guard Corps attacked the vendor’s Bahrain data center. Back in March, a nearby strike took me-south-1 in Bahrain offline and directly struck me-central-1 in the United Arab Emirates, causing a fire.
The vendor encouraged users at the time with workloads in regions in the Middle East to back up their data or migrate to other regions given the volatility there. The me-south-1 region experienced disruption again on March 23 due to drone activity.
AWS Thermal Event Knocks Out Services
In May, one of Amazon Web Services’ busiest data centers sustained an outage due to a “thermal event,” with key AWS services like Elastic Compute Cloud (EC2) instance and Elastic Block Store (EBS) volumes impacted.
AWS confirmed the outage was due to overheating caused by a cooling failure at its North Virginia data center, US-East-1, which is a heavily used region. Several hours after the thermal event occurred, AWS had restored power to its Availability Zone in its US-EAST-1 data center region.
Key AWS services such as Amazon Redshift, Amazon SageMaker, ElastiCache, AWS IoT Core, Amazon Elastic Kubernetes Service and AWS Nat Gateway were impacted.
Cryptocurrency exchange platform Coinbase was impacted by the outage for roughly seven hours.
Anthropic Claude Outage
On April 15, Claude maker Anthropic experienced elevated error rates across its chatbot, the Claude Code coding assistant and its API.
Anthropic has more than 300,000 business customers, according to the vendor.
Anthropic, which has been making a major push into enterprise technology and the channel this year, reported at 14:53 UTC seeing increased errors in those products. The API fully recovered by 16:01 UTC. The company fully resolved the incident by 17:42 UTC, according to its report.
Downdetector showed about 2,000 users reporting issues with Claude as of 1:12 p.m. Eastern, down from about 6,000 users at 10:42 a.m. Around 500 users were reporting issues by 1:34 p.m., according to CNBC.
Although that incident garnered mainstream news attention, a ThousandEyes report on Anthropic pointed out that the company saw multiple outages during April, including some models failing due to inference errors on April 10 and an authentication failure on April 13 leading to user inability to log into Claude for less than an hour starting at 11:45 a.m. Eastern.
ThousandEyes pointed out that AI services, ever-increasing in enterprise adoption, have more distinct functional layers than cloud, each requiring a different diagnostic lens. In the AI age, users are moving from client/server architecture to dynamic environments where the model is a continuously evolving dependency.
“Reliability for AI services feels different from traditional cloud outages because it is different,” ThousandEyes said in its report. “Traditional web services rely on redundancy and elastic, interchangeable compute. If one region fails, you route to another. AI inference doesn’t work that way. The hardware is specialized, expensive, and capacity constrained. You cannot spin up new inference capacity as quickly as you can a standard web server.”
Google Gemini Outage
On June 10 starting at 3:30 Pacific, Google Gemini services across web, mobile, Google Chrome integrations and other surfaces lost availability and saw elevated error rates for about seven hours.
Google Gemini has about 950 million monthly active users.
Google fully restored service after almost 15 hours, the Mountain View, Calif.-based technology giant said in an incident report. “We sincerely apologize for the disruption this incident caused to your business,” according to the report. “We know how much you rely on Google Cloud, and we regret the impact on your productivity. We are working to address the root cause and prevent this from occurring in the future.”
Google blamed the outage on extreme read contention within the foundational database service that manages tool deployment metadata. An index design issue within the database where a column used for tracking deployment expirations contained a high volume of rows with similar values proved to be the primary cause for the failure. Traffic “hotspotted”—that is, concentrated on a small number of database shards, overwhelming them.
To prevent the issue from happening again, Google restructured the database index, implemented better cache policies to prevent load amplification during back-end stress periods and improved uneven database load distribution monitoring and alerting among other fixes.
Microsoft Copilot Outages Hit Enterprise AI Users In June
In June, Microsoft saw two major outages of its Copilot artificial intelligence tool, which has grown to more than 20 million paid seats.
At the start of the month, Copilot experienced a four-hour, 25-minute outage, with certain users unable to access Copilot desktop or the web application.
The Redmond, Wash.-based tech giant published a note online saying that Microsoft restored Copilot service at 5:35 p.m. UTC June 1. The outage started at 1:10 p.m. UTC, with users experiencing app load and timeout errors when accessing the tool.
Signs of a Copilot outage peaked at 486 reports at 8:46 a.m. Pacific on online outage detection website Downdetector. Reports of a Microsoft 365 outage reached 812 by 6:07 a.m. Pacific June 1 but had fallen to under 30 by 1:22 p.m. Copilot users can buy the AI tool as part of the M365 package of applications and services.
On June 11, Microsoft posted on X at 2:07 p.m. Pacific Thursday about “looking into a potential problem impacting Microsoft 365 Copilot chat.” About 20 minutes later, Microsoft published another post explaining that the vendor had “identified an issue with a recent deployment for Microsoft Copilot, which impacts users’ ability to access Copilot chat and http://portal.office.com.”
Reports of Copilot problems reached 1,878 on Downdetector as of 1:47 p.m. Pacific. That marked an increase over the 257 reports logged by 1:02 p.m.
At 3:09 p.m. Pacific, Microsoft posted on X to say that it “successfully reverted the change and confirmed with previously affected users that the issue is resolved.”
Microsoft also saw issues with Copilot on Jan. 15, acknowledging the problem at 7:42 p.m. Pacific and called it “resolved” at 8:24 p.m. The vendor blamed the issue on a configuration change to the service and reverted the change to resolve impact.
Microsoft 365 Outages Disrupt Outlook, Teams
Microsoft saw back-to-back outages with some of its most popular productivity applications at the start of the year, first with an outage across Teams, Outlook and other Microsoft 365 services at 9:11 a.m. Pacific on Jan. 21.
Microsoft has more than 450 million users of its Microsoft 365 Commercial product.
The vendor called the issues “resolved” at 10:29 a.m. Pacific and blamed a third-party network issue after determining that the Microsoft service environment was healthy.
The next day, Jan. 22,Microsoft acknowledged another outage affecting access to Outlook, Defender, Purview and other services in North America.
At 11:37 a.m. Pacific on Jan. 22, the vendor posted to X that it is “investigating a potential issue impacting multiple Microsoft 365 services.”
At 12:17 p.m., the vendor said it “identified a portion of service infrastructure in North America that is not processing traffic as expected.” At 1:14 p.m., Microsoft said it had “restored the affected infrastructure to a (healthy) state” but “further load balancing is required to mitigate impact.”
“We’re directing traffic to alternate infrastructure to achieve recovery,” the vendor said.
Downdetector logged 12,380 reports of an outage in Microsoft’s Outlook email service by 12:15 p.m. Jan. 22; 15,745 reports of an outage in the Microsoft 365 suite of cloud applications as of 12:17 p.m.; and 2,246 reports for the Microsoft Store as of 12:29 p.m.
Downdetector also showed 598 reports of an outage with Microsoft’s Teams communications application and 395 reports of an outage with the Azure cloud service as of 12:15 p.m. Pacific—although because the website uses user-submitted information, it’s possible the users didn’t know which Microsoft product actually went down.
Microsoft’s status page showed that Outlook users might receive a “451 4.3.2 temporary server issue” error message when attempting to send or receive email. Users couldn’t send and receive email through Exchange Online, including notification emails from Microsoft Viva Engage, according to the vendor.
Users also saw delays and failures of collecting message traces and searching within SharePoint Online and Microsoft OneDrive. Users also might not have had the ability to access Microsoft Purview, Microsoft Defender XDR, the Microsoft 365 administrator center and other service portals.
Verizon Wireless Outage Leaves Millions Without Service
Verizon had a rough start to the year with an issue on Jan. 14 that affected wireless voice and data services for customers.
Verizon promised $20 credits for any customer affected by the outage. Downdetector said total reports of an outage exceeded 2.3 million. At around 10 p.m. EST that day, Verizon said the outage was officially resolved. Verizon said it saw “no indication of a cyberattack” regarding the cause of the outage.
The New York-based telecommunications giant and largest wireless carrier acknowledged the issue in a 10:07 a.m. PT post to X, noting that “our engineers are engaged and are working to identify and solve the issue quickly.”
Multiple Verizon users on X posted to say their phones were stuck in SOS mode. The outage extended from New York and Philadelphia to North Carolina and Texas, according to posts on Reddit.
At 9:57 a.m. Pacific, Washington, D.C.’s emergency alerts account on X posted that the “nationwide Verizon Wireless outage that may be affecting some users to connect with 911.”
Cloudflare BYOIP Failure Among Biggest Cloud Outages Of 2026
Cloudflare experienced a six-hour, seven-minute outage on Feb. 20 starting at 17:48 UTC for users of its Bring Your Own IP service.
A change to how Cloudflare managed IP addresses on-boarded through its BYOIP pipeline caused the outage.
About 1,100 BYOIP prefixes were withdrawn from the Cloudflare network before its engineers could revert the change. Some users resolved their own service with the Cloudflare dashboard. Eventually Cloudflare restored all prefix configurations.
Changes Cloudflare promised to implement after the incident include improving its API schema for better standardization for better testing and validation, a redesign to improve rollbacks and introduce layers between customer configuration and production plus better monitoring to detect when changes happen too fast or too broadly.
“We deeply apologize for this incident today and how it affected the service we provide our customers, and also the Internet at large,” Cloudflare said in the post. “We aim to provide a network that is resilient to change, and we did not deliver on our promise to you. We are actively making these improvements to ensure improved stability moving forward and to prevent this problem from happening again.”
Users at the time might not have had the ability to identify Cloudflare as the problem source due to the BYOIP architecture, instead blaming the actual service they wanted to reach, according to a Thousand Eyes report on the incident. And the owners of those services likely looked inward to start because they couldn’t reach their own IP space due to hidden dependency.
The failures also probably looked like function-level disruptions instead of organizational outages and made the connection to Cloudflare unclear.