TL;DR
On Thursday morning US time, the cloud AI services run by OpenAI, Anthropic, Google and xAI all degraded within the same few hours. The overlap is rare, and it punctures a common assumption: that spreading work across several model providers constitutes a continuity plan.
The timeline
Anthropic moved first, flagging a “partial outage” at 9.23am Eastern, with error rates up across three of its models: Claude Opus 5, Fable 5.1 and Mythos 5.1. It identified the cause about fifteen minutes later, deployed a fix, and closed the incident by 12.16pm. A separate short-lived problem affected Claude Sonnet 5 just after midday.
OpenAI followed at 10.43am, citing “elevated errors across ChatGPT and Codex” that were degrading performance. A mitigation followed roughly half an hour after that, with resolution at 12.55pm. Grok was still showing users an error message as the report was published, with xAI saying it was working to restore service. Reports to DownDetector for Grok climbed from fewer than ten just before 9am to 1,365 by 9.45am, easing to 273 later.
Why this matters more than a normal outage
Multi-model architectures are usually sold on resilience: if one provider fails, route to another. Thursday shows the correlation risk that argument ignores. These providers share cloud regions, network paths and, increasingly, the same accelerator supply. Independent vendors are not the same thing as independent failure modes.
For UK businesses that have moved customer-facing or operational workflows onto these APIs, the practical question is what happened to your service during a three-hour window in which four fallbacks were simultaneously unavailable. Most firms cannot answer it, because the failure mode was never modelled.
Looking forward
This lands the day after Nvidia agreed to buy Hugging Face, consolidating the open-model layer under the dominant compute vendor — a second reduction in genuine independence within the same week. The continuity plan that survives this is not a longer vendor list. It is knowing which processes must keep running without any model at all, and having the non-AI path documented before you need it.