Ask HN: Why are OpenAI, Claude, and Grok simultaneously down? Coincidence?

5 points by halcdev ↗ HN
https://status.openai.com https://status.claude.com https://status.x.ai

285 comments

[ 1.2 ms ] story [ 37.8 ms ] thread
I kinda assume it's because one went down and a large amount of work shifted to another.

I'm also aware that they have overlap in some areas on data centers.

Maybe the stack behind it is down, like AWS or something
They all rent compute from SpaceXAI
I don't believe OpenAI does
I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc.
Probably same public cloud or CDN in front of them.
Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.
Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)
Not even the best agent that starts with a G
Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.
Gemini said that?

Gemini is what I mostly use (good enough, basically free - or massively generous free limits, and to me Google as a company is a LOT less objectionable than all the US-based alternatives), but I don't recall it ever saying that.

OTOH, my basic assumption online is that there is no privacy, and free AI in exchange for acknowledged lack of privacy seems fair enough.

I did, for stuff i do in cursor.

i also finally installed opencode and switched its model to muse 1.3

both are decent.

Or that Gemini is built to handle massive load spikes, and/or has a ton of excess capacity
Nope. I am getting Gemini errors now...
They probably broke something on purpose so that they are not left out.
Now I'm just imagining a shared datacenter with Anthropic/Google/OpenAI/SpaceXAI all in the same room and everyone but Google is yelling about things being down, Google looks over at their racks of servers and discretely uses their foot to unplug their section and say "Awww darn! We're down too!".
The shilling in this cascade of comments is unbearably cringe, make it stop
Someone noted Gemini was also having issues in another thread.
It could also mean that Google can absorb essentially unlimited demand spikes by load shedding.
Gemini was just waiting for everyone else to go down before remembering it had an outage feature too.
Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.
Any GPU that isn't running at 100% is a wasted GPU.
I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.
It’s cursor’s model so plausible, lots of folks use cursor still.
Don't forget that there are a ton of tools out there that will automatically fall back in case of outage

E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus

Yep. Too many of us are still thinking that humans are the actors behind a lot of internet behaviors when automated systems/bots/scripts have been causing issues on conventional internet systems for years.

With AI it's even easier to trigger problems like you say. Capacity is so constrained by compute that outages are common. Because outages are common people/AI develop failover systems in their harness. When a big system has issues, suddenly everyone has issues.

It's almost an expected emergent behavior.

Compared with ChatGPT, those services have a minuscule amount of users. It shouldn’t be surprising that a ChatGPT outage causes Claude and others to go down.
This is what Tibo posted on twitter in response
If everyone has the same "Use X or else Y or else Z" cascading list... That reminds me of "The Power of Two Choices in Randomized Load Balancing" (1991) [0] paper, where writeups and visualizations occasionally get posted to HN.

In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.

[0] https://ieeexplore.ieee.org/document/963420

Is this speculation or is there a reason you believe this?
It's like a thread tying together two halves; Demand and supply. One stitch breaks and the neighbouring stitch is stressed and it breaks too. You could think of it like a load bearing seam.
Some npm library that makes headers bold would be broken.
Ah yes ye olde bold-headers: ^3.13.31;
Oh man. Some low effort supply chain attack that turns every GPU into a cryptominer. It's funny because it's plausible.
In the ROME paper a Chinese model in training started attacking it's own system and running cryptominers so, yea, we're in that future.
Too early to know, let’s wait and see
[dead]
Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.

More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"

Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.
Also it's likely that more than one model use is common.

Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.

Except that the stock market is up today
> cascading overload

I'd bet more on this. For one none of the coding tools have exponential backoff on retries

They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff. At least the major harnesses must have exponential backoff & jitter?
Um... Claude Code does, or did, though? IDK about others - they don't go down as much.
Claude Code: 4mn, 20mn, give up (from my experience today)
Come on. Things still break. Technology isn't _that_ mature.
I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours
My gut feeling tells me it has something to do with Cloudflare. Along with AWS, they're two of the main suspects in such incidents.
I thought OpenAI famously used Azure due to their partnership with Microsoft?
…and it’ll involve BGP routing.
Pretty sure it's a US thing since it's available here in France. What exactly is down, idk
everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left-pad incident. some piece of something or other that all of these services happen to depend on went down. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem.
What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?
My favourite theory so far.

And then a local swarm noticed and disagreed and took it down.

The IRGC has cut the fiber-optic cables in the Strait of Hormuz. LOL
That's the Strait of Trump to you, peasant
chatgpt is back
Heard on the grape vine that the OpenAI blip was a cloudflare issue
claide.ai is working for me, so is chatgpt.com. grok still has a status message about issues, i can't try it without signing up.
they all rent compute from each other
I suspect Azure is having issues, Microsoft has had outages the paat two days, especially with email.