Many enterprises have adopted multi-model AI strategies, assuming a second provider will protect them if the first fails. On Sept. 3, that assumption faced its first real test and did not hold up.

That morning, OpenAI’s ChatGPT, Anthropic’s Claude, and xAI’s Grok all experienced significant outages within hours of one another. According to CIO, ChatGPT models were down for roughly two hours, Claude models for four hours, and Grok for nearly three and a half hours. Organizations whose contingency plan amounted to “switch to the other one” found the alternative unavailable as well.

Three Outages, No Single Story

The three vendors did not report the same cause. OpenAI cited a routing error, xAI pointed to a problem at its Memphis compute center, and Anthropic logged elevated errors across several Claude models before deploying a fix. As Memeburn noted, none of the companies has identified a common trigger.

The root cause remains unconfirmed. Axios reported that Microsoft Azure, which provides cloud services for all three, was experiencing issues that morning, and Computing cited reports of ingress failures in Azure’s East US region around the same time. Computing also identified a possible contributing factor, model hopping: when ChatGPT went down, users shifted to Claude and Grok, adding load to services that may already have been under strain.

One further detail stands out. xAI apologized not only to users but also to its affected compute partners, a reminder that the capacity behind these models is shared in ways most customers never see.

A Backup Only Works If It Fails Differently

Most enterprises think about a second AI provider the way they think about a second ISP. The difference is that with an ISP, you can usually see where the lines run. With AI, the weak points that matter—cloud region, edge network, identity provider, shared compute capacity—are largely invisible. Memeburn framed the central question as whether a backup model is truly independent or relies on the same infrastructure as the primary. Most vendors don’t publish enough detail for customers to answer that question with any confidence.

Concentration risk is nothing new to IT, and every major cloud hyperscaler experienced significant outages in 2026. What’s changed is how much work now rests on these models. Technology analyst Carmi Levy told CIO that as AI moves into broader workflows, enterprises could find themselves “uncomfortably exposed” when it fails. With agents now handling tasks once performed by people, an outage can halt work rather than merely slow it.

Where to Start

None of this is an argument for pulling back on AI. It’s an argument for treating AI like any other critical dependency, which means planning for the day it isn’t available. CIOs’ sources recommended that enterprises revisit business continuity plans to account for AI interruptions, including whether alternate or self-hosted models could step in. The following steps offer a practical starting point.

  1. Inventory your AI dependencies. Document every workflow that relies on a single provider, including those buried inside SaaS tools. Coding assistants, CRM copilots, and customer support bots all count, and many of them call the same handful of models.
  2. Ask vendors about failure domains. Find out which cloud regions, CDNs, and identity providers underlie the models you use. A vendor’s inability or unwillingness to answer is itself a meaningful data point.
  3. Put an abstraction layer in front of your models. Routing requests through a gateway lets you swap providers without rewriting applications, though this only pays off if the alternate model runs on genuinely different infrastructure. For the most critical tasks, a self-hosted or open-weight model may be worth the extra overhead.
  4. Write the manual fallback. Decide in advance which processes pause, which revert to manual work, and who makes that call. Teams that have handed entire workflows to agents need this most, since the people who once did that work may no longer remember how to do it.
  5. Test it. Run a tabletop exercise, or pull the plug on a noncritical workflow to see what breaks. A continuity plan that has never been rehearsed offers little real assurance.

Plan Before the Next Outage

The Sept. 3 outage lasted a few hours, making it easy to dismiss as a minor disruption. The next one may run longer and could occur during a quarter close or a product launch. Identifying single points of failure now costs far less than discovering them in the middle of the next incident.

Share Button