It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.
To give you an idea just how much phone traffic was being routed through that building, they had a major power failure at that site and it took down most of the north eastern US's phone network:
On September 17, 1991, management failure, power equipment failure, and human error combined to disable AT&T's central office switch at 33 Thomas. More than five million calls were blocked, and the Federal Aviation Administration private lines were also interrupted, disrupting air traffic control to 398 airports serving most of the northeastern United States.
Were there significant API outages too? I didn’t notice any on my production workflows, and I’d assume what you’re implying would cover API routes too, otherwise it seems kinda pointless.
What makes you think they'd need to touch every datacenter? All of these endpoints use existing providers with decades-long history at this point, and network monitoring is already a proven 'feature' of the agencies they'd need to co-exist with over their lifetimes.
If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.
I'd say Occam's Razor leans easily to the former as well, given the history of projects that Snowden revealed and were never shut down, plus all of the cooperation with the federal government that's being touted in recent announcements from both companies.
I'm confused by your statement. Are you suggesting that these companies that have invested tens of billions in their networks and datacenters decided to create a shared single point of failure?
Wouldn't be surprised: Snowden's revelations 10 plus years ago already showed how the NSA was injected into the data centers of Social Media, it's only logical that they would now demand to be injected into the biggest, most information providing data stream of the planet of the present: LLM services.
I read Nowhere to Hide recently, really worth it if you can get past Greenwald sticking himself in the middle (start halfway through).
The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now.
Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist under these conditions.
So following this to the obvious conclusion, the NSA is responsible for the closure of the Straight of America, high tarrifs, dropping employment, and high gas prices?
If this hypothesis were true then a spy agency may want to rewrite responses. Every tool in an agent's harness becomes an remote procedure call you can make on that machine. Including a tool to execute a shell command, in many. A harness is completely isometric to a backdoor, it's the same code written with a different intention.
There's a million ways a bad configuration can take down a network. Especially if the part that gets squirrley is a black box.
Even if it isn't precisely this, the fact that no one is saying anything is quite surprising.
Edit: I don't want do contribute to FUD, so want to call out this comment and its replies that identify the shared layer as probably being xAI's infra: https://news.ycombinator.com/item?id=49568622
Historical precedent repeated over and over, and the US ties to DoW work, and the national security implications. It's actually the Occam's Razor explanation if you know the history.
Maybe it was a power hub. And we're not allowed to know where the DC is. If we knew it was a power issue, then with other information (perhaps over time) we could determine the DC location. Or something along these lines.
Very likely yes. I wouldn't be surprised if they were hosted from the same datacenters even. There has been a story every few weeks about how Musk has sublet X.ai capacity for one company or another.
This whole thing makes me thing about a passage in Dune where they mentioned the Spacing Guild transported entire fleets of ships in isolated compartments and leaving said compartments was a capital offense. This way, entire militaries of mortal enemies were shipped to battlefield, with nothing but bulkheads separating each other.
The economics of this kind of warfare make no sense to me. The Harkonens must have been incredibly wealthy by running the spice trade, but even after years of saving and plotting said that the transit fees to move the armies by the Guild were ruinous.
How could anyone wage war like this? The defenders will always outnumber the attackers.
https://acoup.blog/2026/02/24/collections-warfare-in-dune-pa... Lays out an interesting take on this: that armies are quite small, even on developed worlds, because the cost of equipping them is enormous — and the technological advantage is absurd enough that you don’t need a large army. When 300 men can hold a planet, why would you need 1000?
So less guys with knives and more WH40k Space Marines. Impossibly expensive elites who cannot be equaled by a mortal.
Still leaves me questioning how anyone could wage war. If the Harkonens could barely afford it, nobody can.
Also wondering if the Guild charges different rates for goods vs military. What is worth the brutal intergalactic shipping prices that you would not develop local industry to create it? Surely nobody is moving grain or ore, yet the Guild ships are portrayed as comically massive.
But they accepted the transit fees, which means they endorsed the action.
The guild would not allow any action to jeopardize the flow of spice. Thus any attack on Dune is implicitly sanctioned, notwithstanding their spice trade with the Fremen.
Even today (and back when Dune was written), modern military tech is invincible to an enemy below a level of sophistication. The US used to hammer insurgents with Predator drones who had no ability to retaliate.
This playbook has only flipped due to a lot of cheap tech (and the machinery to make them) coming in uncotrollably to these countries has put them on a much more equal playing field.
Since the Guild controls what comes in, they can avoid this scenario from happening. They can control who can have what, and who can make what. Which I think is one of the true, deep explanations of why military tech in Dune is so weird and inefficient - essentially the efficiency and power of a weapon is dictated by how much the Guild charges for shipping them if they allow you to, at all. Which is how you end up with dudes with swords. It gives an idea of the level of technology in the universe that you can equip said dude with a nigh-impenetrable shield that fits in a belt buckle. I guess that's the result of optimizing for Guild transit fees, not real-world power.
Really it's not exactly a novel idea, but as I think more about it, it's uncanny how Dune (at least initially) is basically the fantasy Suez crisis and what followed. It explores the power dynamics very well, how factions can hold humanity hostage without firing a single shot, and how what they don't bother controlling comes to bite them in the ass.
> How could anyone wage war like this? The defenders will always outnumber the attackers.
Which is exactly what ends up happening, but it does take a bunch of extraordinary events.
Why would you default to that explanation? That’s not at all a reasonable default, and I say that as something pretty paranoid with regards to US surveillance
Looking at history, it's much more reasonable to assume there's surveillance, since there are whole branches of the government that exist for "national security", which this easily falls into. See Marissa Mayer explaining that it's not an option to refuse [1]. I assume this is just the same ole' Room 641A [2].
We already know there is surveillance, but why would you default to that explanation for a downtime? As said by others monitoring the traffic isn’t done by routing through the NSA monitoring system, it’s not a bottleneck that would take down all those services. If you look at more details such as the timing it’s even less likely to be the case, but even as a default it’s not a reasonable explanation given the symptoms
It could be as simple as a new model (astra) was released which takes more resources combined with a surge in usage due to novelty took down OpenAI. Meanwhile everyone at big companies have the ability to switch models and moved to Anthropic pushing it too over the edge.
The OpenAI outage lasted only 15 minutes and when it happened everyone started to use the other models which created super heavy load for them. This then cascaded into them all being down.
Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.
Also, OpenAI is saying what caused it:
> "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"
Anthropic stated their issue started earlier:
> "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”
I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.
If it's a "thundering herd" problem where everyone's harness falls back to less popular providers that don't normally see that much demand, I'd say the probability is pretty good. Classic cascading failure.
This would be consistent with providers failing 80 minutes apart instead of simultaneously.
If you ballpark it as a single 3 hour downtime window per week and iid Poisson, then overlapping downtime probability of 2 providers is approximately the expected occurrence rate per 3 hours, 1/56. Not particularly surprising at all.
"All of these" is two. OpenAI had a router issue. Anthropic had a separate issue. Anthropic uses a lot of SpaceX compute, so an Anthropic issue and a SpaceX issue can be one in the same, as was likely the case this time.
And they weren't down at precisely the same time. Anthropic's issue started ~1 hour before OpenAI's.
Watched it happen. Doesn't look like traffic moving off OpenAI took out the others. Looked like the opposite. The Anthropic thread was initially chock full of marketing accounts claiming "Last straw! I finally moved to Codex and I'm so happy! No problems there!" - and then they all got deleted when Codex went down too.
I thought the consensus on here yesterday was that it was likely caused by cascading failures. OpenAI had an issue during their GPT-6 rollout, taking down their service. This caused a lot of OpenAI users to push their requests (or a larger share of their requests) to Claude and/or Grok, which pushed their load high enough to cause outages.
We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.
Yet another blatant proof that those companies are not smarter than anyone else. They might be focusing on intelligence and yet in practice we can all see they are not doing better than most.
Not doing better than most at routine IT/platform stuff, they seem to be doing fantastically at co-opting the US Gov into their "vision" and ramming through DC's getting built.
OpenAI was rushing and making last minute changes for a big product launch. Easy to mess something up in that situation. The outage lasted like 20 minutes.
Anthropic is down a lot regardless, and in this case only specific models were affected.
xAI probably couldn't handle the extra traffic it was getting from the other two.
I work at OpenAI and I was the Incident Commander for yesterday's outage.
We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.
Note: Incident Command is almost certain an allusion (or implementation) of the Incident Command System [1]
It is a common system in all kinds of emergency response scenarios, including local emergency services (fire/police/ambulance) and it scales all the way to massive disasters.
It is especially useful to clarify command structures when multiple response entities need to coordinate. That is true even within organizations like public companies, where the reporting structures may be distinct.
115 comments
[ 0.24 ms ] story [ 111 ms ] threadIt was extremely weird... and if it was a load thing they probably would have explained it by now?
why?
https://theintercept.com/2016/11/16/the-nsas-spy-hub-in-new-...
https://control.fandom.com/wiki/Oldest_House
https://en.wikipedia.org/wiki/Control_(video_game)
On September 17, 1991, management failure, power equipment failure, and human error combined to disable AT&T's central office switch at 33 Thomas. More than five million calls were blocked, and the Federal Aviation Administration private lines were also interrupted, disrupting air traffic control to 398 airports serving most of the northeastern United States.
If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.
It is not often the Executives or Legal even know, but sometimes they did. AT&T bent over backwards to help.
This is standard behavior by the CIA and NSA, and has been for a long time.
https://www.propublica.org/article/nsa-documents-suggest-clo...
https://www.nytimes.com/2015/08/16/us/politics/att-helped-ns...
https://www.theguardian.com/world/2014/mar/19/us-tech-giants...
https://theintercept.com/2018/06/25/att-internet-nsa-spy-hub...
The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now.
Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist under these conditions.
Even if it isn't precisely this, the fact that no one is saying anything is quite surprising.
Edit: I don't want do contribute to FUD, so want to call out this comment and its replies that identify the shared layer as probably being xAI's infra: https://news.ycombinator.com/item?id=49568622
Gemini, too.
https://arstechnica.com/ai/2026/09/four-major-ai-models-suff...
They still do, but they used to, too.
This whole thing makes me thing about a passage in Dune where they mentioned the Spacing Guild transported entire fleets of ships in isolated compartments and leaving said compartments was a capital offense. This way, entire militaries of mortal enemies were shipped to battlefield, with nothing but bulkheads separating each other.
How could anyone wage war like this? The defenders will always outnumber the attackers.
Still leaves me questioning how anyone could wage war. If the Harkonens could barely afford it, nobody can.
Also wondering if the Guild charges different rates for goods vs military. What is worth the brutal intergalactic shipping prices that you would not develop local industry to create it? Surely nobody is moving grain or ore, yet the Guild ships are portrayed as comically massive.
The guild would not allow any action to jeopardize the flow of spice. Thus any attack on Dune is implicitly sanctioned, notwithstanding their spice trade with the Fremen.
This playbook has only flipped due to a lot of cheap tech (and the machinery to make them) coming in uncotrollably to these countries has put them on a much more equal playing field.
Since the Guild controls what comes in, they can avoid this scenario from happening. They can control who can have what, and who can make what. Which I think is one of the true, deep explanations of why military tech in Dune is so weird and inefficient - essentially the efficiency and power of a weapon is dictated by how much the Guild charges for shipping them if they allow you to, at all. Which is how you end up with dudes with swords. It gives an idea of the level of technology in the universe that you can equip said dude with a nigh-impenetrable shield that fits in a belt buckle. I guess that's the result of optimizing for Guild transit fees, not real-world power.
Really it's not exactly a novel idea, but as I think more about it, it's uncanny how Dune (at least initially) is basically the fantasy Suez crisis and what followed. It explores the power dynamics very well, how factions can hold humanity hostage without firing a single shot, and how what they don't bother controlling comes to bite them in the ass.
> How could anyone wage war like this? The defenders will always outnumber the attackers.
Which is exactly what ends up happening, but it does take a bunch of extraordinary events.
No need. They simply cut off the offending faction from all space travel.
And is anyone keeping a table of correlations between outages? Sounds like valuable data.
[1] https://www.cnet.com/tech/services-and-software/yahoo-report...
[2] https://en.wikipedia.org/wiki/Room_641A
Does this really need an explanation?
Also, OpenAI is saying what caused it:
> "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"
Anthropic stated their issue started earlier:
> "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”
I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.
And not only that, when one goes down a bunch of API traffic switches over to the other, spiking demand and knocking it down.
This would be consistent with providers failing 80 minutes apart instead of simultaneously.
And they weren't down at precisely the same time. Anthropic's issue started ~1 hour before OpenAI's.
Biggest competitor goes down and all of a sudden you have a lot more traffic...
In fact, it feels pretty ordinary.
Anthropic: 6:23 am PT.
OpenAI: 7:43 am PT.
Watched it happen. Doesn't look like traffic moving off OpenAI took out the others. Looked like the opposite. The Anthropic thread was initially chock full of marketing accounts claiming "Last straw! I finally moved to Codex and I'm so happy! No problems there!" - and then they all got deleted when Codex went down too.
"Incident with Grok 4.6 Copilot AI Model Provider"
at 7:20 am PT
We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.
Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
https://news.ycombinator.com/item?id=49551096
Anthropic is down a lot regardless, and in this case only specific models were affected.
xAI probably couldn't handle the extra traffic it was getting from the other two.
Not everything is a conspiracy.
We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.
It is a common system in all kinds of emergency response scenarios, including local emergency services (fire/police/ambulance) and it scales all the way to massive disasters.
It is especially useful to clarify command structures when multiple response entities need to coordinate. That is true even within organizations like public companies, where the reporting structures may be distinct.
1. https://en.wikipedia.org/wiki/Incident_Command_System
This country is conducting a test of the Emergency SAIfguard System.
THIS ONLY A TEST
In the event of a real emergency you would have been given instructions on how to grab your ankles and kiss your butt goodbye.
...
This concludes our test of the Emergency SAIfguard System."