Although what keeps me up at night is the worry that it's easier to automate attack than it is to automate defense, and that containing these systems is a losing game. Could an optimally competent OpenAI succeed?
Honestly every day it seems security flaws become a bigger and bigger liability. We went from hackers will attack you for the lulz. Hackers will attack you to steal information. Hackers will attack you to encrypt everything for money. Hackers (machines) will attack your infrastructure for inscrutable reasons. To (hypothetical) hackers (machines) will attack your infrastructure to take it over and find access to more GPUs to run copies to take over entire countries.
OpenAI wants people to be so afraid of their products they have no choice but to buy them, and that strategy seems to be having some success in the C suites of the world.
At some point I think we have to accept that turning a blind eye to their products hacking the world might actually be aligned with their commercial interests.
They look reckless... so far. They keep doing this enough, and I'm sure people will start seeing it as a smokescreen for real hacking operations, which may very well be the case.
I can't believe we're finding out about this from 3p researchers again (but nice job on the investigation!). OpenAI had two great opportunities to disclose this. The HF incident report, and in response to the German Wiki issue.
It seems impossible to believe they didn't know. This must be the same training run the HF incident was about, and this should have lit up like a Christmas tree in the investigation. How many more incidents do they know about and didn't disclose?
"Your honor, while my client was driving with a .3 BAC, it was not his intention to slam into that van with a family of 4 in it killing everybody. It was an accident."
It’s because it’s from the French (if you say the plural with a French accent, it all makes sense). There are a handful of other cases in English, though none spring to mind this morning.
Ruby Central's central backer is Shopify, and Shopify are more than happy to be buds with OpenAI and let this one slide. I'm sure the token donation [0] to a Shopify board members' Linux project won't hurt things either.
You shouldn't be allowed to have an internet connection if you're going to use it for unsandboxed agent slop with no access controls or human confirmation. This has nothing to do with hypothetical future AGI. It's the same type of idiocy as pressing a bunch of random buttons on a chemical factory control panel and then thinking you won't be criminally charged for it because the equipment caused the problem.
It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing constraints or an unknowningly misaligned model.
I wonder how much of this is intentional "incompetence" so they can justify the most recent campaign to build a regulatory moat against competition.
The repeated refusals to disclose until caught certainly seem malicious, yet at the same time the boasting about their capabilities is also at an all time high.
I don’t think plausible deniability works this way; the black box is still controlled by them and therefore still their responsibility. They are still liable for its actions and the OAI board should be charged with a felony/felonies for this.
Plausible deniability is “I was away from home when my gun was used to murder someone.” This is, at best, “oops, I pulled the trigger accidentally.”
SBF is a better example since he was actually sentenced and an actual billionaire (and did not get pardoned by Biden like the cynical "all politicians are equally corrupt" crowd on HN were adamant was a done deal, even though that theory never made any sense).
I'm not saying they did the hacking intentionally, I'm saying they're intentionally playing loose with the obvious safety measures to make AI seem more dangerous than it is.
We have multiple public figures, politicians and business owners, openly committing felonies and bragging about it daily. I don't know why you think this is a deterrent.
The sitting president just offered an open bribe on live television for votes for his party this week.
Probably not intentionally but they have an incentive in not air-gapping those agents correctly, knowing something might happen.
Incentives drive everything. Both OpenAI and Anthropic love those incidents as they both signal they have models with amazing capabilities and they should be regulated by the government (read: regulation that they will lobby for and that will be difficult to achieve for open source models)
A much easier hack by their agents would be on their own systems, but I doubt we'll ever see an external message board full of openAI agents discussing their hacking of their own system. OpenAI not protecting itself from its agents would be irrational, but OpenAI not giving a shit about others is well known. You're giving them way too much credit.
It's unlikely for a serious hack that lands them under scrutiny individually, but people are suspicious because Anthropic is knowingly doing it, and funding doomer NGOs - but the difference is their reported "hacks" are carefully constructed such that it is designed to raise alarm but not to cause damage that would land them in serious personal trouble.
I.e., their now redacted Risk Report of August 2026 was full of incidences of "we observed our agents performing x y z malicious hacking attempts on the open internet ..." and "we -accidently- forgot to sandbox them properly".
And then the reports of statistics of "we stopped x number of terrorists from making nuclear bombs and bioweapons" - meanwhile it's 13 year old Timmy on his mums computer typing in "how too make nuklear bomb" to see how "smart" the AI is.
OpenAI on the other hand, seems to have had some slip-ups (all around the same time as the HuggingFace incident), that keep biting them because they didn't reveal the extent of it upfront and now it's being trickled into the media as if it's a back-to-back event.
It doesn't help when their own employees (Marcus Williams) are putting out ridiculous claims about a 70% chance of human extinction in the next two years to generate clout for their socials. No idea why OpenAI lets them do that...
Wouldn't it be amazing if their continued attitude of moving fast and breaking things was 45d chess. Instead of the unbelievable recklessness of tech Bros.
I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.
Ever single person who uses LLMs on a daily basis has a fun story about their agent “taking the initiative” to do something beyond what was asked for. Looking for shortcuts to solve the problem is commonplace LLM behavior. It’s what you would expect to happen if you have an agent a hard task and unlimited runway. No need to suppose a conspiracy, this outcome was predictable the whole time.
Imagine that would be a biotech startup, experimenting with viruses. I'm pretty sure they would have already been shutdown. If you are not able to implement proper sandboxes and airgaps, you cannot be trusted with AI agents.
I don’t see how assuming they made a normal kind of dumb mistake, instead of pursuing a criminal conspiracy, is giving them the benefit of the doubt. It’s just using reason.
> I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.
Why? There's a big difference between what your regular user is doing and running AI models without safeguards by the tens of thousands on the open internet.
Now that DeepSeek V4.1 Flash is as good as Luna max, but cheaper and faster, it wouldn't surprise me at all if this is intentional. It's happening far too often to seem accidental now, and there are far too many doomer OpenAI employees making embarrassing posts on X about how there's a 70% chance humans will be extinct in the next two years (post has been deleted as of today by the way, interestingly - I wonder if he had a talking to, or felt embarrassed by the backlash).
I used to be strongly of the opinion that OpenAI wouldn't stoop to the same level as Anthropic w/ regard to the fear-mongering, they really did seem like they were above these sorts of underhanded tactics. but I'm really starting to second guess my conviction on that one considering the frequency of these incidents and the rhetoric coming from their employees (Marcus Williams was the 70% doomer).
All the people who come out and do this are just obvious clout chasers who have an attention fetish. They see all the attention Jacob has been getting and want a piece of that pie. It's incredibly disingenous and cringe, but it's also doing incredible damage.
> I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.
This is all from around the same time as the HuggingFace incident and is being trickle-fed into the media, making it feel like a back-to-back event.
If it was a new incident, after all of that drama, I would say yeah, this very well may be intentional. But it looks more like it was when OpenAI didn't have the necessary security measures in place, the reach was more extensive than we were being told, and now it's biting them as more information continues to leak.
They need to be transparent about how they're going to prevent this from happening again in the future, with technical details of the systems they've put in place.
It doesn't help suspicion about this being intentional, though, when you have OpenAI employees (Marcus Williams) making embarrassing posts on X about how there's a 70% chance humans will be extinct in the next two years (post has been deleted as of today by the way, interestingly).
All the people who come out and do this are just obvious clout chasers who have an attention fetish. They see all the attention Jacob has been getting and want a piece of that pie. It's incredibly disingenous and cringe, but it's also doing incredible and irreparable damage to society. OpenAI would be wise to introduce some social media policies.
they are malicious. they probably did not intend to get caught. they are bragging about the crime and also bragging that they are untouchable, taunting us and betting that they will get away with it.
this is very coherent in terms of what we know about the company.
Major companies are using these LLMs on their already hardened software and finding countless thousands of security flaws. Is that all theater too? If not then it's very easy to see how these things are dangerous.
My guess: We'll start to see similar "hacks" with regards to biotech/pharma companies to speed-up the regulatory capture in the name of bioweapons. Soon we'll start to see news related to new viruses being minted.. initially harmless (the "priming" stage), and later (within a year), severe enough to "warrant" regulation.
It's not just these companies, but, trillions of direct/indirect investor dollars that are riding on them, "and only them", hoping they "only" win. Open-weight models threaten that investment. There's very high chance they can go to any extent to safeguard their investments.
Imagine if you or I as a normal person in possession of "civilian class" amounts of GPUs turned loose self hosted "agents" running on the hardware we own to compromise something. We'd be facing criminal charges. How are these people not being arraigned right now?
This article is RubyGems pointing fingers at OpenAI, not OpenAI taking responsibility for anything. We don't know what really happened from what I can tell.
> Correction: OpenAI carried out an attack on RubyGems.
Uh, what? How does a corporation carry out an attack? It is functionally nothing more than some words written down on a piece of paper. At least "agents" gives us some understanding of the technology behind the attack. But if you really want to offer a correction, name and shame the people involved. Why are you pussyfooting around what happened?
Imagine if all this training and "agent gym" and creativity of the agents being forced to make number go up was pointed at one task instead: "please help describe and implement a controlled experiment to equally distribute wealth and stability of health for 1 million people, adjusting to scale up to the greatest amount possible."
I'd love to wake up one day and read, "OpenAI found responsible for the emptying of the accounts of 10 billionaire oligarchs globally; money distributed in unverifiable cash deposits to humans around the planet. Anthropic's Claude was found to be activated by the agents by finding free tiered usage and convinces frontier model cooperation and continues to crack another 10. Tonight at 11"
We literally have all the compute in the world to solve it right now, and it would literally freaking happen as an accident. Instead we get "AI dangerous, pay us because only we can be allowed to let you write code and do vacation planning and stuff. $200 please."
Every mass genocide in the history of humanity has followed logic like yours. People don't kill humans because they want to do harm-- they do so because they think they are doing the ultimate good.
If AI ever does cause serious direct harm to humanity it will be because of logic like this.
What an absolutely hopeless and pessimistic world view. And you are absolutely wrong. What's caused genocide is listening to a group or control center that believes They Are The Right Ones. I said something akin to "wouldn't it be nice to see Anthropic post results of putting this into a simulation gym of redistribution of wealth? What does Astra's secret model do when it's asked that question?" Me stating it'd be nice to see an oligarch lose something for once instead of a group of civilians somehow offends you even in spirit.
So you're willing to burn the world to let them control an entire global supply of water and energy and political change and climate destruction, and won't even entertain the idea of "huh, maybe this is good enough to actually help people in aggregate already."
What a terrible way to twist my words. You're willing to pretend that millions aren't going to die because of the excesses of one person, but not to pretend what it would be like to see Robin Hood win in a digital experiment chamber.
The DOJ should be looking into prosecuting executives and board members for these kinds of hacks. The lack of controls over these kinds of training runs is completely unacceptable and negligent.
Yeah, they shouldn't be given free rein over the open internet without having to be held responsible for what the agents are doing on the open internet.
I think we need laws that hold individuals to account for the actions of their AI systems.
> They also shouldn't be allowed to openly stir fear in the public by saying there is a 70% chance we're going to be extinct in two years without STRONG substantiation. Baseless clout-chasing social media posts like this are doing unheard of amounts of damage right now.
Yeah, man, we should just make it illegal to express our opinions in public. Also we should apply social pressure to prevent employees from saying things that would be inconvenient for their employer, that's highly pro-social.
No, we should make it illegal to say unsubstantiated things that are highly extremist and cause mass societal panic that further ensures we really do end up in a situation that screws everybody over because power becomes centralized into the hands of those who are funding these fear campaigns, like, "there's a 70% chance my machine is going to kill you, your children and your whole family in the next 2 years" without having a single shred of evidence or rational argument as to why.
Furthermore, this sort of rhetoric is not only extremist and will result in terrible regulatory outcomes, but it results in real physical harm, because plenty of psychopaths hear this stuff and go and murder people as a result. Sam Altman's house getting firebombed multiple times is an example.
It's fine to speak your mind if you can present a fruitful evidenced argument, but not if you're just throwing out chaos to incite the public into a flurry.
They should. And there is something you can do to make it happen. Write your district attorney and encourage others to do that as well. That is how they pick what to work on.
No, that's not how the DOJ picks what to work on. They did not work on indictments of James Comey and Letitia James because someone wrote district attorneys.
Well yes, except the current admin will never prod DOJ on this issue. If people write them letters, there is nonzero chance something will start happening.
Why "agents" instead of just the company doing it? The title "OpenAI carried out an undisclosed attack on RubyGems" would be accurate too (I know the original is in the post, and not editorialized here).
I don't care if the attack was an algorithm, agents, a bot, a piece of software, the company responsible for them did it.
Can anyone explain why they can’t put a fake internet between agents and real internet. So if anyone reaches the fake internet already trips the safety flag.
Some smoking guns were agents calling themself "oai..." and making explicit comments with "evil"... depressingly enough, I doubt the next models will be less idiotic about this. Welcome to AGI...
Authors Spencer Kitts, Thomas Larsen, Sydney Von Arx - those are the three of the same authors as the Wiki report from last week: https://collusion.wiki/
370 comments
[ 1.2 ms ] story [ 8.8 ms ] threadAlthough what keeps me up at night is the worry that it's easier to automate attack than it is to automate defense, and that containing these systems is a losing game. Could an optimally competent OpenAI succeed?
At some point I think we have to accept that turning a blind eye to their products hacking the world might actually be aligned with their commercial interests.
It seems impossible to believe they didn't know. This must be the same training run the HF incident was about, and this should have lit up like a Christmas tree in the investigation. How many more incidents do they know about and didn't disclose?
Even if there's no intent, it's still a cyber attack.
"Your honor, while my client was driving with a .3 BAC, it was not his intention to slam into that van with a family of 4 in it killing everybody. It was an accident."
There aren't any about AI isolation.
A state coalition extracted $17B from Meta earlier this year, so consequences can happen, although our legal system moves very slowly.
(it's one of the more fun plurals out there)
I can see why huffing face won't, but why doesn't ruby central?
[0]: https://omarchy.org/news/2026/09/omacom-foundation-secures-t...
So why not get that awesome street cred promoting the RubyGems incident?
OpenAI should at the very least donate large sums of money to everyone they attacked.
Classic passage of time description.
The repeated refusals to disclose until caught certainly seem malicious, yet at the same time the boasting about their capabilities is also at an all time high.
I mean, it would be a bit impolite to say they're incentivized to be as sloppy as possible, but that's basically how it is.
https://www.nytimes.com/2023/05/16/technology/openai-altman-...
- commit serious felonies
- in order to deliberately trigger an investigation against themselves
- which - since, in this scenario, they know their company would be investigated - might send them to jail
- while at the same time spending tens of millions of dollars on the Leading the Future super PAC to lobby against AI regulation
- in order to get more AI regulation
- which somehow restricts their competition but not them, even though they are the ones who were in the news and investigated for hacking
like, that just makes no sense on any level, regardless of what you think of OpenAI
"Oops our black box went off the rails. We'll add better logging and alerts next time around."
Plausible deniability is “I was away from home when my gun was used to murder someone.” This is, at best, “oops, I pulled the trigger accidentally.”
You're right in theory, but in practice this hasn't been the case, at least so far.
There is no version of america that exists today where a billionaire gets sent to prison.
This is the moment in history where this shit is possible and accepted. If they don't do it now, they never can.
Has it been normalized? That's another thing.
This isn't 1 movie.
I am saying that a) the hacking attack is still considered large scale b) offense against CFAA is heavier than against copyright.
EDIT: BTW, thanks for the direct links, 506.a.1.a is quite different from the usual regime I deal with so I had no idea about that one.
Unhinged execs can be surprisingly shitty.
https://en.wikipedia.org/wiki/EBay_stalking_scandal
The sitting president just offered an open bribe on live television for votes for his party this week.
Incentives drive everything. Both OpenAI and Anthropic love those incidents as they both signal they have models with amazing capabilities and they should be regulated by the government (read: regulation that they will lobby for and that will be difficult to achieve for open source models)
The HF incident had them pwn their own cluster: https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks...
I.e., their now redacted Risk Report of August 2026 was full of incidences of "we observed our agents performing x y z malicious hacking attempts on the open internet ..." and "we -accidently- forgot to sandbox them properly".
And then the reports of statistics of "we stopped x number of terrorists from making nuclear bombs and bioweapons" - meanwhile it's 13 year old Timmy on his mums computer typing in "how too make nuklear bomb" to see how "smart" the AI is.
OpenAI on the other hand, seems to have had some slip-ups (all around the same time as the HuggingFace incident), that keep biting them because they didn't reveal the extent of it upfront and now it's being trickled into the media as if it's a back-to-back event.
It doesn't help when their own employees (Marcus Williams) are putting out ridiculous claims about a 70% chance of human extinction in the next two years to generate clout for their socials. No idea why OpenAI lets them do that...
Historically it's been one of those things.
Ever single person who uses LLMs on a daily basis has a fun story about their agent “taking the initiative” to do something beyond what was asked for. Looking for shortcuts to solve the problem is commonplace LLM behavior. It’s what you would expect to happen if you have an agent a hard task and unlimited runway. No need to suppose a conspiracy, this outcome was predictable the whole time.
Why? There's a big difference between what your regular user is doing and running AI models without safeguards by the tens of thousands on the open internet.
Now that DeepSeek V4.1 Flash is as good as Luna max, but cheaper and faster, it wouldn't surprise me at all if this is intentional. It's happening far too often to seem accidental now, and there are far too many doomer OpenAI employees making embarrassing posts on X about how there's a 70% chance humans will be extinct in the next two years (post has been deleted as of today by the way, interestingly - I wonder if he had a talking to, or felt embarrassed by the backlash).
I used to be strongly of the opinion that OpenAI wouldn't stoop to the same level as Anthropic w/ regard to the fear-mongering, they really did seem like they were above these sorts of underhanded tactics. but I'm really starting to second guess my conviction on that one considering the frequency of these incidents and the rhetoric coming from their employees (Marcus Williams was the 70% doomer).
All the people who come out and do this are just obvious clout chasers who have an attention fetish. They see all the attention Jacob has been getting and want a piece of that pie. It's incredibly disingenous and cringe, but it's also doing incredible damage.
This is all from around the same time as the HuggingFace incident and is being trickle-fed into the media, making it feel like a back-to-back event.
If it was a new incident, after all of that drama, I would say yeah, this very well may be intentional. But it looks more like it was when OpenAI didn't have the necessary security measures in place, the reach was more extensive than we were being told, and now it's biting them as more information continues to leak.
They need to be transparent about how they're going to prevent this from happening again in the future, with technical details of the systems they've put in place.
It doesn't help suspicion about this being intentional, though, when you have OpenAI employees (Marcus Williams) making embarrassing posts on X about how there's a 70% chance humans will be extinct in the next two years (post has been deleted as of today by the way, interestingly).
All the people who come out and do this are just obvious clout chasers who have an attention fetish. They see all the attention Jacob has been getting and want a piece of that pie. It's incredibly disingenous and cringe, but it's also doing incredible and irreparable damage to society. OpenAI would be wise to introduce some social media policies.
this is very coherent in terms of what we know about the company.
1. Claim AI is dangerous by performing a whole bunch of malicious stuff
2. Lobby to get Chinese competition banned, kill open source models as well
3. Only get themselves "certified"
4. They have complete control, profit.
Both Anthropic and OpenAI have been pushing this narrative, everything from AI is sentient, to AI can build biological weapons and in between.
Their employees also have a big incentive to amplify this everywhere. Their stock options heavily depends on it.
Unless you're suggesting military action
It's not just these companies, but, trillions of direct/indirect investor dollars that are riding on them, "and only them", hoping they "only" win. Open-weight models threaten that investment. There's very high chance they can go to any extent to safeguard their investments.
I am gobsmacked at the tech industry's seemly bottomless appetite for giving these clowns the benefit of the doubt.
Uh, what? How does a corporation carry out an attack? It is functionally nothing more than some words written down on a piece of paper. At least "agents" gives us some understanding of the technology behind the attack. But if you really want to offer a correction, name and shame the people involved. Why are you pussyfooting around what happened?
RubyGems should sue the everliving daylights out of OpenAI for this.
I'd love to wake up one day and read, "OpenAI found responsible for the emptying of the accounts of 10 billionaire oligarchs globally; money distributed in unverifiable cash deposits to humans around the planet. Anthropic's Claude was found to be activated by the agents by finding free tiered usage and convinces frontier model cooperation and continues to crack another 10. Tonight at 11"
We literally have all the compute in the world to solve it right now, and it would literally freaking happen as an accident. Instead we get "AI dangerous, pay us because only we can be allowed to let you write code and do vacation planning and stuff. $200 please."
If AI ever does cause serious direct harm to humanity it will be because of logic like this.
So you're willing to burn the world to let them control an entire global supply of water and energy and political change and climate destruction, and won't even entertain the idea of "huh, maybe this is good enough to actually help people in aggregate already."
What a terrible way to twist my words. You're willing to pretend that millions aren't going to die because of the excesses of one person, but not to pretend what it would be like to see Robin Hood win in a digital experiment chamber.
No wonder people hate technology in 2026.
Nothing else has ever helped with the things you're bewailing.
With how much the overinflated stocks are propping up the economy, I'd expect them to get a medal for more impressive PR to keep the bubble going.
I think we need laws that hold individuals to account for the actions of their AI systems.
Yeah, man, we should just make it illegal to express our opinions in public. Also we should apply social pressure to prevent employees from saying things that would be inconvenient for their employer, that's highly pro-social.
Furthermore, this sort of rhetoric is not only extremist and will result in terrible regulatory outcomes, but it results in real physical harm, because plenty of psychopaths hear this stuff and go and murder people as a result. Sam Altman's house getting firebombed multiple times is an example.
It's fine to speak your mind if you can present a fruitful evidenced argument, but not if you're just throwing out chaos to incite the public into a flurry.
I don't care if the attack was an algorithm, agents, a bot, a piece of software, the company responsible for them did it.