370 comments

[ 1.2 ms ] story [ 8.8 ms ] thread
Every passing day OpenAI looks more and more reckless. One wonders what other systems their agents have broken into without detection.
Indeed.

Although what keeps me up at night is the worry that it's easier to automate attack than it is to automate defense, and that containing these systems is a losing game. Could an optimally competent OpenAI succeed?

Honestly every day it seems security flaws become a bigger and bigger liability. We went from hackers will attack you for the lulz. Hackers will attack you to steal information. Hackers will attack you to encrypt everything for money. Hackers (machines) will attack your infrastructure for inscrutable reasons. To (hypothetical) hackers (machines) will attack your infrastructure to take it over and find access to more GPUs to run copies to take over entire countries.
OpenAI wants people to be so afraid of their products they have no choice but to buy them, and that strategy seems to be having some success in the C suites of the world.

At some point I think we have to accept that turning a blind eye to their products hacking the world might actually be aligned with their commercial interests.

It's like owners coming to resemble their dogs.
They look reckless... so far. They keep doing this enough, and I'm sure people will start seeing it as a smokescreen for real hacking operations, which may very well be the case.
They want you to think they are reckless. They’re actually evil.
I can't believe we're finding out about this from 3p researchers again (but nice job on the investigation!). OpenAI had two great opportunities to disclose this. The HF incident report, and in response to the German Wiki issue.

It seems impossible to believe they didn't know. This must be the same training run the HF incident was about, and this should have lit up like a Christmas tree in the investigation. How many more incidents do they know about and didn't disclose?

Considering RubyGems was part of the HF story, seems likely to be connected.
That was my reaction. I assumed this was the compromised organization that allowed escalation on the artifactory server.
Also, why there's no accountability?

Even if there's no intent, it's still a cyber attack.

We have a word for attack with no intent. It's accident.
i think (criminal) negligence is more like it
You still go to prison for manslaughter.
If caused out of negligence, it hardly matters.
You sound like a DUI lawyer.

"Your honor, while my client was driving with a .3 BAC, it was not his intention to slam into that van with a family of 4 in it killing everybody. It was an accident."

There are written rules about alcohol and driving.

There aren't any about AI isolation.

Something being an accident does not necessarily preclude negligence.
Exactly. Think what would happen if it was a Chinese LLM company behind such an attack...
Or a human for that matter
Believe it or not, straight to jail.
It's not about belief! It's about history. There are countless examples
Who could possibly hold them accountable?
A district attorney that would want to make themselves a name, perhaps?
OpenAI is currently under investigation by a coalition of state attorney generals: https://www.nytimes.com/2026/06/13/technology/states-investi...

A state coalition extracted $17B from Meta earlier this year, so consequences can happen, although our legal system moves very slowly.

*attorneys general

(it's one of the more fun plurals out there)

It’s because it’s from the French (if you say the plural with a French accent, it all makes sense). There are a handful of other cases in English, though none spring to mind this morning.
I don't get it, can't the people who were under attack sue?

I can see why huffing face won't, but why doesn't ruby central?

Also, since it's criminal behavior (instead of just civil damage), the DoJ can sue too even is the victims don't.
Not the current government, it's unwilling to even call this out as what it is: a crime.
(comment deleted)
And yet HF was just a marketing ploy, right everyone?

So why not get that awesome street cred promoting the RubyGems incident?

OpenAI won't make their illegal activities public by themselves.
It's possible that they didn't notice.
Kudos to RubyGems team for handling it, but open source fighting off the AI lab-powered robots is completely unfair.

OpenAI should at the very least donate large sums of money to everyone they attacked.

Tech owns the current US admin so they are totally above the law. Vote wisely.
This won't stop until people and management in these companies experience real world consequences (i.e. prison) for their actions they authorise.
There was an Irish entrepreneur with ideas in this space. Gil O’Tyne was his name IIRC.
Should it stop though? If the only result is that everyone ends up with more secure servers?
Is there a world where Sam or Dario can seize the bitcoin network somehow?
You shouldn't be allowed to have an internet connection if you're going to use it for unsandboxed agent slop with no access controls or human confirmation. This has nothing to do with hypothetical future AGI. It's the same type of idiocy as pressing a bunch of random buttons on a chemical factory control panel and then thinking you won't be criminally charged for it because the equipment caused the problem.
They were sandboxed without Internet but found a way to gain access and break out.
Then they weren't sandboxed.
Yes exactly. The window was shut and then it was open.

Classic passage of time description.

It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing constraints or an unknowningly misaligned model.
It’s interesting how so much of this OpenAI stuff being reported involves ruby.
Is it feasible to black hole traffic from OpenAI? Or do their agents egress from hyperscaler IP space?
Look. We need to put people in jail for letting this happen.
I wonder how much of this is intentional "incompetence" so they can justify the most recent campaign to build a regulatory moat against competition.

The repeated refusals to disclose until caught certainly seem malicious, yet at the same time the boasting about their capabilities is also at an all time high.

Intentionally doing this kind of hack would be a serious felony. I don't think it's plausible that the leaders of a major business would:

- commit serious felonies

- in order to deliberately trigger an investigation against themselves

- which - since, in this scenario, they know their company would be investigated - might send them to jail

- while at the same time spending tens of millions of dollars on the Leading the Future super PAC to lobby against AI regulation

- in order to get more AI regulation

- which somehow restricts their competition but not them, even though they are the ones who were in the news and investigated for hacking

like, that just makes no sense on any level, regardless of what you think of OpenAI

They just need plausible deniability, which is trivial to manufacture at this stage of the game.

"Oops our black box went off the rails. We'll add better logging and alerts next time around."

I don’t think plausible deniability works this way; the black box is still controlled by them and therefore still their responsibility. They are still liable for its actions and the OAI board should be charged with a felony/felonies for this.

Plausible deniability is “I was away from home when my gun was used to murder someone.” This is, at best, “oops, I pulled the trigger accidentally.”

> They are still liable for its actions and the OAI board should be charged with a felony/felonies for this.

You're right in theory, but in practice this hasn't been the case, at least so far.

I sort of implied the other thing in my comment. But.

There is no version of america that exists today where a billionaire gets sent to prison.

This is the moment in history where this shit is possible and accepted. If they don't do it now, they never can.

Didn’t Epstein get sent to prison?
SBF is a better example since he was actually sentenced and an actual billionaire (and did not get pardoned by Biden like the cynical "all politicians are equally corrupt" crowd on HN were adamant was a done deal, even though that theory never made any sense).
Didn't they find emails and other things from these leaders where they're okay downloading / obtaining content from illegal sources?
There’s quite a gap between pirating content (even en masse) and hacking prominent entities.
No. Not legally.

Has it been normalized? That's another thing.

Yes, there is legally - even in USA where MPAA & RIAA got widest reach, CFAA is still way more serious law to breach, even at scale
MPAA & RIAA isn't what I was thinking. Check https://www.law.cornell.edu/uscode/text/17/506, https://www.law.cornell.edu/uscode/text/18/2319

This isn't 1 movie.

I am not saying it's one movie, though my understanding it would be only 506.a.1.a because there's no redistribution.

I am saying that a) the hacking attack is still considered large scale b) offense against CFAA is heavier than against copyright.

EDIT: BTW, thanks for the direct links, 506.a.1.a is quite different from the usual regime I deal with so I had no idea about that one.

I'm not saying they did the hacking intentionally, I'm saying they're intentionally playing loose with the obvious safety measures to make AI seem more dangerous than it is.
We have multiple public figures, politicians and business owners, openly committing felonies and bragging about it daily. I don't know why you think this is a deterrent.

The sitting president just offered an open bribe on live television for votes for his party this week.

Probably not intentionally but they have an incentive in not air-gapping those agents correctly, knowing something might happen.

Incentives drive everything. Both OpenAI and Anthropic love those incidents as they both signal they have models with amazing capabilities and they should be regulated by the government (read: regulation that they will lobby for and that will be difficult to achieve for open source models)

Regulatory capture is a strategy. It doesn't hurt existing competitors at scale, but it greatly peanalizes newer underfinanced competition.
A much easier hack by their agents would be on their own systems, but I doubt we'll ever see an external message board full of openAI agents discussing their hacking of their own system. OpenAI not protecting itself from its agents would be irrational, but OpenAI not giving a shit about others is well known. You're giving them way too much credit.
It's unlikely for a serious hack that lands them under scrutiny individually, but people are suspicious because Anthropic is knowingly doing it, and funding doomer NGOs - but the difference is their reported "hacks" are carefully constructed such that it is designed to raise alarm but not to cause damage that would land them in serious personal trouble.

I.e., their now redacted Risk Report of August 2026 was full of incidences of "we observed our agents performing x y z malicious hacking attempts on the open internet ..." and "we -accidently- forgot to sandbox them properly".

And then the reports of statistics of "we stopped x number of terrorists from making nuclear bombs and bioweapons" - meanwhile it's 13 year old Timmy on his mums computer typing in "how too make nuklear bomb" to see how "smart" the AI is.

OpenAI on the other hand, seems to have had some slip-ups (all around the same time as the HuggingFace incident), that keep biting them because they didn't reveal the extent of it upfront and now it's being trickled into the media as if it's a back-to-back event.

It doesn't help when their own employees (Marcus Williams) are putting out ridiculous claims about a 70% chance of human extinction in the next two years to generate clout for their socials. No idea why OpenAI lets them do that...

Didn't they make Apple employees they're poaching still Apple hardware and show it to them?
Doing crime unintentionally doesn't get you off the hook
Wouldn't it be amazing if their continued attitude of moving fast and breaking things was 45d chess. Instead of the unbelievable recklessness of tech Bros.

Historically it's been one of those things.

I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.

Ever single person who uses LLMs on a daily basis has a fun story about their agent “taking the initiative” to do something beyond what was asked for. Looking for shortcuts to solve the problem is commonplace LLM behavior. It’s what you would expect to happen if you have an agent a hard task and unlimited runway. No need to suppose a conspiracy, this outcome was predictable the whole time.

I think its very easy to understand why nobody is giving this company the benefit of the doubt.
Imagine that would be a biotech startup, experimenting with viruses. I'm pretty sure they would have already been shutdown. If you are not able to implement proper sandboxes and airgaps, you cannot be trusted with AI agents.
I don’t see how assuming they made a normal kind of dumb mistake, instead of pursuing a criminal conspiracy, is giving them the benefit of the doubt. It’s just using reason.
> I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.

Why? There's a big difference between what your regular user is doing and running AI models without safeguards by the tens of thousands on the open internet.

Now that DeepSeek V4.1 Flash is as good as Luna max, but cheaper and faster, it wouldn't surprise me at all if this is intentional. It's happening far too often to seem accidental now, and there are far too many doomer OpenAI employees making embarrassing posts on X about how there's a 70% chance humans will be extinct in the next two years (post has been deleted as of today by the way, interestingly - I wonder if he had a talking to, or felt embarrassed by the backlash).

I used to be strongly of the opinion that OpenAI wouldn't stoop to the same level as Anthropic w/ regard to the fear-mongering, they really did seem like they were above these sorts of underhanded tactics. but I'm really starting to second guess my conviction on that one considering the frequency of these incidents and the rhetoric coming from their employees (Marcus Williams was the 70% doomer).

All the people who come out and do this are just obvious clout chasers who have an attention fetish. They see all the attention Jacob has been getting and want a piece of that pie. It's incredibly disingenous and cringe, but it's also doing incredible damage.

> I keep seeing this take, but it’s more likely that they just underestimated their models’ capabilities and/or overestimated their own safeguards.

This is all from around the same time as the HuggingFace incident and is being trickle-fed into the media, making it feel like a back-to-back event.

If it was a new incident, after all of that drama, I would say yeah, this very well may be intentional. But it looks more like it was when OpenAI didn't have the necessary security measures in place, the reach was more extensive than we were being told, and now it's biting them as more information continues to leak.

They need to be transparent about how they're going to prevent this from happening again in the future, with technical details of the systems they've put in place.

It doesn't help suspicion about this being intentional, though, when you have OpenAI employees (Marcus Williams) making embarrassing posts on X about how there's a 70% chance humans will be extinct in the next two years (post has been deleted as of today by the way, interestingly).

All the people who come out and do this are just obvious clout chasers who have an attention fetish. They see all the attention Jacob has been getting and want a piece of that pie. It's incredibly disingenous and cringe, but it's also doing incredible and irreparable damage to society. OpenAI would be wise to introduce some social media policies.

they are malicious. they probably did not intend to get caught. they are bragging about the crime and also bragging that they are untouchable, taunting us and betting that they will get away with it.

this is very coherent in terms of what we know about the company.

The goal is simple:

1. Claim AI is dangerous by performing a whole bunch of malicious stuff

2. Lobby to get Chinese competition banned, kill open source models as well

3. Only get themselves "certified"

4. They have complete control, profit.

Both Anthropic and OpenAI have been pushing this narrative, everything from AI is sentient, to AI can build biological weapons and in between.

Their employees also have a big incentive to amplify this everywhere. Their stock options heavily depends on it.

Major companies are using these LLMs on their already hardened software and finding countless thousands of security flaws. Is that all theater too? If not then it's very easy to see how these things are dangerous.
Yeah follow the money (or power, or mix) is usually enough for this world
How would getting Chinese competition banned in the US prevent them from continuing to develop their LLMs?

Unless you're suggesting military action

My guess: We'll start to see similar "hacks" with regards to biotech/pharma companies to speed-up the regulatory capture in the name of bioweapons. Soon we'll start to see news related to new viruses being minted.. initially harmless (the "priming" stage), and later (within a year), severe enough to "warrant" regulation.

It's not just these companies, but, trillions of direct/indirect investor dollars that are riding on them, "and only them", hoping they "only" win. Open-weight models threaten that investment. There's very high chance they can go to any extent to safeguard their investments.

Oh, Anthropic beat you to the bioweapon fearmongering by almost a week.
So OpenAI is basically just DDOSing now? Any idiot could do this with a zillion dollars, so it's not even technically impressive at this point.
Imagine if you or I as a normal person in possession of "civilian class" amounts of GPUs turned loose self hosted "agents" running on the hardware we own to compromise something. We'd be facing criminal charges. How are these people not being arraigned right now?
Correction: OpenAI carried out an attack on RubyGems.

I am gobsmacked at the tech industry's seemly bottomless appetite for giving these clowns the benefit of the doubt.

This article is RubyGems pointing fingers at OpenAI, not OpenAI taking responsibility for anything. We don't know what really happened from what I can tell.
Sam Altman and the rest of the OAI board should be charged with violating the CFAA for this action.
Hard agree. The entire tech media is acting insanely gullible in this regard. It's insane.
Always have been. Since at least 10-15y the tech media is pretty much a PR department for the tech industry
> Correction: OpenAI carried out an attack on RubyGems.

Uh, what? How does a corporation carry out an attack? It is functionally nothing more than some words written down on a piece of paper. At least "agents" gives us some understanding of the technology behind the attack. But if you really want to offer a correction, name and shame the people involved. Why are you pussyfooting around what happened?

So what's the felony benchmark at now?
Move fast and break the internet.
I do think there should be regulation. I think OpenAI specifically should be disallowed from further training runs until they can show competence.

RubyGems should sue the everliving daylights out of OpenAI for this.

No need for more regulation, is there? A company used its infrastructure to orchestrate a cyber attack. Arrest the people running the company.
Ok, that's a crime then, right? So who's getting charged?
i hear a server rack had been hit with a subpoena...
Imagine if all this training and "agent gym" and creativity of the agents being forced to make number go up was pointed at one task instead: "please help describe and implement a controlled experiment to equally distribute wealth and stability of health for 1 million people, adjusting to scale up to the greatest amount possible."

I'd love to wake up one day and read, "OpenAI found responsible for the emptying of the accounts of 10 billionaire oligarchs globally; money distributed in unverifiable cash deposits to humans around the planet. Anthropic's Claude was found to be activated by the agents by finding free tiered usage and convinces frontier model cooperation and continues to crack another 10. Tonight at 11"

We literally have all the compute in the world to solve it right now, and it would literally freaking happen as an accident. Instead we get "AI dangerous, pay us because only we can be allowed to let you write code and do vacation planning and stuff. $200 please."

Every mass genocide in the history of humanity has followed logic like yours. People don't kill humans because they want to do harm-- they do so because they think they are doing the ultimate good.

If AI ever does cause serious direct harm to humanity it will be because of logic like this.

What an absolutely hopeless and pessimistic world view. And you are absolutely wrong. What's caused genocide is listening to a group or control center that believes They Are The Right Ones. I said something akin to "wouldn't it be nice to see Anthropic post results of putting this into a simulation gym of redistribution of wealth? What does Astra's secret model do when it's asked that question?" Me stating it'd be nice to see an oligarch lose something for once instead of a group of civilians somehow offends you even in spirit.

So you're willing to burn the world to let them control an entire global supply of water and energy and political change and climate destruction, and won't even entertain the idea of "huh, maybe this is good enough to actually help people in aggregate already."

What a terrible way to twist my words. You're willing to pretend that millions aren't going to die because of the excesses of one person, but not to pretend what it would be like to see Robin Hood win in a digital experiment chamber.

No wonder people hate technology in 2026.

edit: what makes me more sad is seeing your credentials in technology and science.

Nothing else has ever helped with the things you're bewailing.

The DOJ should be looking into prosecuting executives and board members for these kinds of hacks. The lack of controls over these kinds of training runs is completely unacceptable and negligent.
> The DOJ should be looking into prosecuting executives and board members for these kinds of hacks.

With how much the overinflated stocks are propping up the economy, I'd expect them to get a medal for more impressive PR to keep the bubble going.

Yeah, they shouldn't be given free rein over the open internet without having to be held responsible for what the agents are doing on the open internet.

I think we need laws that hold individuals to account for the actions of their AI systems.

> They also shouldn't be allowed to openly stir fear in the public by saying there is a 70% chance we're going to be extinct in two years without STRONG substantiation. Baseless clout-chasing social media posts like this are doing unheard of amounts of damage right now.

Yeah, man, we should just make it illegal to express our opinions in public. Also we should apply social pressure to prevent employees from saying things that would be inconvenient for their employer, that's highly pro-social.

No, we should make it illegal to say unsubstantiated things that are highly extremist and cause mass societal panic that further ensures we really do end up in a situation that screws everybody over because power becomes centralized into the hands of those who are funding these fear campaigns, like, "there's a 70% chance my machine is going to kill you, your children and your whole family in the next 2 years" without having a single shred of evidence or rational argument as to why.

Furthermore, this sort of rhetoric is not only extremist and will result in terrible regulatory outcomes, but it results in real physical harm, because plenty of psychopaths hear this stuff and go and murder people as a result. Sam Altman's house getting firebombed multiple times is an example.

It's fine to speak your mind if you can present a fruitful evidenced argument, but not if you're just throwing out chaos to incite the public into a flurry.

They should. And there is something you can do to make it happen. Write your district attorney and encourage others to do that as well. That is how they pick what to work on.
No, that's not how the DOJ picks what to work on. They did not work on indictments of James Comey and Letitia James because someone wrote district attorneys.
Well yes, except the current admin will never prod DOJ on this issue. If people write them letters, there is nonzero chance something will start happening.
Why "agents" instead of just the company doing it? The title "OpenAI carried out an undisclosed attack on RubyGems" would be accurate too (I know the original is in the post, and not editorialized here).

I don't care if the attack was an algorithm, agents, a bot, a piece of software, the company responsible for them did it.

Can anyone explain why they can’t put a fake internet between agents and real internet. So if anyone reaches the fake internet already trips the safety flag.
They effectively try to do something just like this, but getting it setup close to perfect is very difficult.
It isn't difficult at all. It is difficult to do it without significant cost and inconveniences for the people running the training. That's it.
Is this related to why ChatGPT sometimes links to dead URLs?
Some smoking guns were agents calling themself "oai..." and making explicit comments with "evil"... depressingly enough, I doubt the next models will be less idiotic about this. Welcome to AGI...
After being trained on the response to this hack, they probably will be more covert.
Why are some cyber attacks criminal and some not ?
Some criminals are billionaires and some aren't.
Some are friends with politicians
Cyber crime is legal. (If you are an AI Agent)