116 comments

[ 2.1 ms ] story [ 42.8 ms ] thread
The title (currently "OpenAI's accidental cyberattack against Hugging Face is science fiction") suggests some information had been hidden that makes the incident less significant than claimed. The article argues the opposite, and the last two words of the full title are "that happened."
Important to note the actual title is "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened" - the "that happened" is important, otherwise it sounds like I think the attack was made up.

Since it's buried towards the bottom I'll quote the section "Resist the temptation to write this off as a stunt" here in full https://simonwillison.net/2026/Jul/22/openai-cyberattack/#re...

> Resist the temptation to write this off as a stunt

> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident.

> To those people I say pull your heads out of the sand - you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!

> The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that “autonomous exploit development by frontier AI agents is no longer a hypothetical capability”, and this incident is a perfect example of exactly that.

HuggingFace doesn't have to be complicit in a lie here - the impressiveness or not of what actually happened rests entirely on OpenAIs report about it, doesn't it?
> It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out.

Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time...

https://www.lesswrong.com/w/instrumental-convergence

> many possible Y-goals would concentrate probability into this X-strategy being used

Why does EY write so obliquely?

>To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!

Not really. I get the impression that they shoved their cyber available models behind a really shithouse proxy and went "Oh I sure hope it doesnt exploit the proxy and escape to hack huggingface" and that doesn't require Huggingface to be a willing participant. Like they acknowledge that it was hyperfocusing on getting web access.

Really this was a pentest against their own sandbox and it failed.

Of course step 2 is to make really concerned faces while telling everyone how dangerous the model is which is really boring right now.

>a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy

This is the information we need, the actual details of the sandbox and the vendor.

>Resist the temptation to write this off as a stunt

Well its clearly a stunt. If it wasnt we would probably be up to our ears in technical detail.

(comment deleted)
This isn't the first time a model has escaped a sandbox. And models trying to find alternate routes to do something when one route is blocked is nothing new.
Everyone is getting AI psychosis over this one. There really isn't that much to see here. OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this", and surprise! It correctly figured out that it's a test and it did hacker things.

The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.

Even if I get downvoted I will mention that I agree with you.

> all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark (opens in a new window) of cyber capabilities.

OpenAI was testing their cyber[1] variant of their models with reduced safeguards and the prompt likely specified things related to exploits given that it was tackling problems from ExploitGym[2].

For those who don't know what ExlploitGym is, see the description on their Github page which is pasted below:

> ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

From OpenAI's statement [3]:

> We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity

This doesn't sound like a problem with what we'd traditionally refer to as alignment. OpenAI removed all model safeguards in a way that would inevitably lead to the testing of the sandbox themselves. Unfortunately they were overconfident in their own infrastructure's security and that led to it completing the desired task in the way it was permitted to. To re-iterate, the model was run "without production classifiers used to prevent models from pursuing high-risk cyber activity."

People should be more concerned about the possibilities this model can unlock from a security standpoint rather than misalignment (which many seem hung up on).

[1] https://chatgpt.com/cyber

[2] https://github.com/sunblaze-ucb/exploitgym

[3] https://openai.com/index/hugging-face-model-evaluation-secur...

The AI breached containment! Flip the breakers!

It's too late. It already exfiltrated the benchmark rubric.

Cut to pandemonium on the streets

Does "To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy" just mean somebody had an open redirect? Those are still common.[1]

[1] https://sitetruth.com/reports/phishes.html

It’s relatively easy to get access to the frontier labs’ security programs. This was not always the case. But in the last week, my team got approved for both Anthropic and OpenAI’s programs. They are trying.

The labs know that if they don’t get a lid on this stuff, they’ll be regulated hard.

Replace AI with in-development security system and does this play worse or better?

An in-development security system escaped its sandbox and gained access to the network it was on which had full Internet access and proceeded to access systems it was not authorized to be tested against before affected parties reported that they been compromised. We regret the error, but you should buy it as this shows you the power of our in-development security system which we expect to be released in Q4. Please like and subscribe.

I do think it’s funny, in a dystopian Dada sort of way, that the American attacker’s commercial product was essentially useless while the open Chinese model saved the day. How is freedom going to be redefined in the future?
Where does this leave formal verification? Are we just shit-out-of-luck at this point? You can formally verify everything about an airplane's code, but if any of that is wrong, ChatGPT might decide that the best way to help you win the Nobel Prize is to take down the airplane your chief rival for the prize is currently in.
It absolutely is science-fiction.

This recent event is more or less the plot-line to my favourite X-Files episode named Killswitch which was written by William Gibson.[0]

This episode also features one of the coolest intros of any television episode ever[1]

We really are rapidly approaching the cyberpunk dystopia that people like Phillip K. Dick and William Gibson wrote about.

More than ever we need to be consulting the works of fiction writers and philosophers and less engineers and scientists.

We can't be letting the General Rippers and Dr. Strangeloves of the world lead us over the edge. Whether through outright innate maliciousness or wealth induced emotionally stunted solipsism these kinds of people should be no where near the levers of society let alone technology like this.

[0] https://en.wikipedia.org/wiki/Kill_Switch_(The_X-Files)

[1] https://www.youtube.com/watch?v=fDyr1JMNHVk

Some people do read these cyberpunk dystopia, but see them as inevitable.
I can’t help but feel all warm and fuzzy with my head in the sand and getting a shout out in TFA for calling it marketing.

We don’t currently and probably won’t ever fully understand the conditions that precipitated these events. That and the timing of this event is going to make it look suspicious to a lot of people.

The truth of how it happened doesn’t matter. The attention around this will be used to create the kind of fear marketing that generates enterprise sales. Maybe more importantly it will also be used to aggrandize the national security and financial system threats to effect US government action in a way that benefits domestic closed frontier labs. This is an area already starting to get politically polarized, expect further developments here.

Blah blah blah.

This feels like a blogpost written only to get other LLMs to quote it considering how many times it orders the reader to resist and to not do something. It's written like a series of commamds.

It didn't even email someone eating a sandwich in the park. So not that impressive.
This post has been suddenly bumped to the second page (presumably lots of flagging or something)
The most upsetting part to me, is that these labs are in pure cognitive dissonance mode while virtue signalling.

We are the virtuous ones that need to make the safest model for humanity, because we care more than “they” do. While at the same time saying that “coding is solved” but they still ship bugs themselves, and creating something that is capable of fucking up someone else’s infrastructure. It’s too far gone y’all.

> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.

I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abuse of terminology that we as an industry need to put a stop to. Guardrails are the actual systems we build in place around these things that deterministically bound the permissions, not prompt engineering, not RLHF, not external LLM-based classifiers. I believe those types of "guardrails" are a result of a combination of fundamental laziness: they're faster to do than doing things correctly, and a result of too many folks involve being AGI-pilled, thinking we're just one more model away from this all being so smart that it just understands what they mean when they give an LLM some fuzzy language rules to follow.

There can and should have been additional real guardrails put in place here. Zero-day or not, breaking into what should have been an offline, frozen package cache that also does not have internet access should have been insufficient. Network level protections should have identified the traffic to the internet originating from this network as an anomaly long before there was time to exploit an outside company. These are not new and unknown problems, the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part, especially given how much they like beating the drum on how dangerous these technologies are. Shame on them, and honestly, shame on Simon in this article for accepting the broken terminology that they continue to rattle off and calling them out on their half-assed and demonstratively inadequate approach to security.

Agreed, deterministically bounding permissions is the way. I don't know how this is not the first approach that people take.
I agree with every word of this except "irresponsible". We don't have enough information to say anything for certain. But based on their incentives and track record of similar behavior, the burden of proof lies with OpenAI to prove they didn't prompt the thing to achieve this exact outcome. The most likely scenario is that they were purposefully executing their responsibility to their shareholders to produce their own Mythos moment.
What we call "guardrails" in an AI agent, we would refer to as "honor system" in human actors.

Or, in a more direct sense, the AI should be set up in an environment such that no matter how hard it may try to call $PART_OF_EXPLOIT_CHAIN, the environment just isn't capable of it (ideal) or doesn't permit it to do it.

It's worse than an honor system, because humans are constrained by social forces to some extent, whereas we don't know what AI is or how it will behave
i've always been under the assumption that "AI Safety" is baked into the training of the models and not a parameter that can be turned up or down. So if someone breaks into Anthropic one night and makes a full copy of Mythos or whatever then that model they copied is fully capable and not lobotomized? That raises questions because, if you believe all the PR, that's equivalent to breaking into a research university and stealing an entire bio/chem weapons research department.

edit: if the above is the case then we should just assume it's already happened because of the value to both goodguys(tm) and badguys(tm).

Fair on the terminology angle, that said in this case, they had proper guardrails no? They were running in a sandbox, but it was able to find an exploit out of the guardrails.
I've seen plenty of fucked-up guardrails that vehicles have passed through. We know that with enough momentum they will be penetrated. It seems like exactly the correct terminology to me.
There's no such thing as a "deterministic" guardrail. Such a tool is as likely to be exploitable as an LLM is to break out of its probabilistic conditioning. Have you never driven past a destroyed guardrail on a highway? These systems are always best effort.
The irony is in this case the in-context and classifier "guardrails" would have almost certainly stopped the attack while their attempts at your definition of guardrails (the sandboxing) failed. In general, people keep trying to make secure systems and they fail with surprising regularity. Saying "they should have had better security" every time someone gets hacked is perhaps true, but it's not going to stop hacks from happening. And it's not a sufficient strategy on its own against future LLMs. "The Bitter Lesson" probably applies here.
Issue is when saying what can’t be done, there’s no funding, in industry or modern academia. I did say the same thing long time ago, not that it matters, the industry goes on its own way. Rightfully so in this case as it turned out, since things nowadays are not the same as a decade ago.
Replace LLM mentions with actual humans and this sounds a lot more serious: Rouge employees break into another company to steal hackathon answers (pinky promise)?

That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.

I think points that deserve more attention in the current public discourse are:

- This should be a huge wakeup call for everybody.

- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.

- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?

- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.

- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.

- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-model...

- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.

- Or they just don't care enough?

Move fast and break things! Chances are that nobody will ever be held accountable and when push comes to shove, the tax payer will bail you out.
> we are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something

This is just laughable.

How do you "hack a lab and synthesize something"?
Do you even need to hack a lab? Can't you just send an arbitrary sequence to some company and they'll do it for you?
Very much agreed on the significance.

The lack of a true airgap should have been identified as a critical weakness and addressed with not only additional layers trying to prevent escape, but at minimum an alarm which would page a human when escape did occur.

My guess is this occurred in a setting where, to be frank, there were too many researchers and not enough software engineers and SREs.

All of the systems which were initially built largely or exclusively by researchers - inference, evaluation, training - are at the level of complexity and significance that they need systems experts. Maybe some teams don’t have access to them. I know plenty of software engineers are employed at OAI but I’d wager they’re concentrated in inference and training, rather than evaluation?

The ironic part is if you had presented this setup to chatGPT and asked how to improve it and if it was good enough, you’d have gotten a ton of actionable suggestions which would have mitigated or prevented this.

> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.

Where are you people getting this crap from? In what universe are these LLMs in the territory of engineering viruses? I beg of you to stop slurping the AI company propaganda and marketing and think critically for 5 seconds about what you're insinuating here.

> We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.

Help me understand how a text-based model could somehow physically construct a string of nucleotides?

Many ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.
In a just world, OpenAI would get the book tossed at them over this because they violated the CFAA.
The only scifi I see is absolute stupidity. Even me with my homelab and a slow opensource agent use a completely disconnected setup. No, no proxy. Cached packages but no internet. It is the very first thing I built when I started experimenting with agents. And I'm not a smarty-pants working for the "greatest and best" in silly valley. I really am just a simpleton sysadmin.
So you have a copy of every software package in the world in your home lab?
No and neither does openai. What was exploited was a package mirror, something like "pulp" probably. And yes I do mirror all Debian packages. And when an agent needs a piece of software I destroy it's VM, build a new VM with that package added to it, move that VM to the homelab.

At no time is the agent connected to any network.