Finally mainstream news understands. The unfiltered version:
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
> AI managed to escape using standard and well documented script kiddie methods
> AI broke in using standard script kiddie methods.
I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?
Are we finally now in 2026 coming around to the idea that sometimes entities may find themselves incentivized to conspire with each other? Is theorizing about such no longer off-limits due to a thought-terminating cliche?
I wish you hadn’t pulled the Balaji case into your argument. Personally, I find it ludicrous that Altman would hire a hitman to off a copyright whistleblower. Even if one gets past the insane risk of hiring a hitman, and the deep criminal connections required, it would be totally ineffective. He already blew the whistle, and his testimony would be irrelevant since all the evidence persists in disk and in logs.
I can agree with you on points 1,2, and 3 and still find it important and concerning news. AI have found real world 0 days before, we’re seeing tons of security patches coming in. Open weight models are catchy up. Right now everyone is at risk from this technology as is perhaps something big will capture headlines soon but we’re just gpu constrained from bad actors being able to wield them successfully.
Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently.
we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it
does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?
if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
> does the article end at "How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control?" or is there more that is paywalled?
That's how the article ends; The Guardian doesn't have a paywall (yet).
> I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit.
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
Not sure what they are trying to say exactly. What should we be skeptical of? Did the incident not happen? Was it reported incorrectly? Are any of the parties involved lying?
Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
The author is trying to provide a counterweight to the volume of articles that simply repeat OpenAI's account of the events and their interpretation without much pushback.
They are not suggesting that OpenAI or HF have lied about what happened, but rather that OpenAI is advancing a narrative framing their models as supremely dangerous and capable, while positioning themselves as the only ones qualified to manage that danger.
At the same time they are not being particularly transparent about what actually happened (e.g. was this one-shotted or if not how many trials did they run and what were the outcomes of those, was it emergent as a part of routine cyber-capabilities tests, how much prompting was involved, what prompts were used)
Note that this is at a time when they are lobbying for a regulatory approach that would give frontier labs special treatment.
I would guess the editorial team at The Guardian may not like articles that get too in the weeds of technical details and questions like these that the vast majority of their readers wouldn't understand. I don't know. But I empathize with your disappointment. I don't think it's fair to say that they are contributing "nothing" especially given what most reporting on this has looked like.
Based on the comment section here, it seems like the opposite; the majority of the people here need a reminder that companies aren't actually genies that can only tell falsehoods, where you can only understand what they are saying by correctly guessing the conspiracy underneath.
Taking nothing at face value gives you just as much of a distorted view of reality as taking everything at face value.
I don't care how it happened someone should be arrested for illegal intrusion. Agents don't work on their own, someone is responsible. If nobody else the CEO for allowing something unsupervised.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
I agree with the attacker side. The actions of autonomous agents are absolutely the responsibility of the one or more humans that enabled them to take that action. Whether that means someone is arrested, maybe or maybe not, but at least there should be a hefty fine.
I disagree with the defender side. It's not an unreasonable end state, but we're nowhere near there now. It would require holding company employees legally responsible for the security of their services, which means the risk of being employed as a (defensive) security professional is much higher, which means pay needs to be much higher and insurance needs to be available, etc. It's a very different world.
On the weekends, I'm coding up a list management app with a sync server. It's unreleased but exposed to the internet. (This is not hypothetical.) If that server ends up being used as part of an exploit chain, am I legally liable too?
Forget about age verification, now you want to associate every exposed port on the internet with a legally responsible human?
yeah the way the agent “escaped” their sandbox was always a bit off, seemed a bit too easy and surprised they didn’t have instrumentation to catch an non whitelisted network request. still demonstrates the capability though.
If the agents so easily “escaped” their sandbox and can reach hugging face, isn’t it easy for this agent to reach supposedly chat conversations not supposed to be used for AI training within the lab’s databases?
There are some reasons the story could be inaccurate in some ways: OAI stands to benefit if people think their models are strong, and they have a history of doing things with dubious ethics (e.g. using data for training against the terms of its creators, abandoning the non profit mission, stealing or attempting to steal Apple IP).
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
1) OpenAI and HuggingFace are both telling the truth.
IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.
2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.
I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.
3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less
(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).
It was quite galling to read the press initially verbatim quoting Delangue's enthusiastic reports of the incident, as if it wasn't immediately clear it was being spun for promotion.
I don't understand the conspiracy theories here. Everyone is well aware that AI agents are creative, powerful, and stupid.
AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.
There seems to be three popular ways to view this incident.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
Why do none of these few constrained ways to view this complex situation (nice gig if you can get it, agenda setting) include "and also this looks a heck of a lot like the stuff that the LW folks have been warning about for years and maybe we should slow down or stop?"
One way to think about it is that more powerful models mean solid best practices are more important than ever, so humans moving too quickly / carelessly bites us more than ever.
If a typical SaaS platform moved and shifted this fast with this many downstream consequences, we’d tell them to slow the fuck down, stop launching new features, and focus on security for a second.
But in AI I guess the idea is that more power and intelligence will solve for everything else.
It's the same thing as the "oh no, we need to figure something out for our youth on the verge of obsolescence" narrative. It's all about giving an impression of unfathomable power. They don't care that in actuality, the tech will end up as just an augment, not a replacement for said youth.
Industry pressures -> lack of safeguards -> fake it till you make it -> let's spin this.
Which, by the way, would be an Orwellian reversal from what the company was supposedly founded to do, but there's a reason the Open in OpenAI is a meme.
Remember that they've been doing this since GPT-2 was too powerful to release. They are world class experts in this PR pipeline.
Considering incentive structures at play is solid epistemiology, but the line of thinking in your comment is a tad reductive, IMHO.
In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3?
If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile suppliers, and then have the model responsibly disclose those breaches. If we're willing to lie collude, and cheat, this story is easy to manufacture with at least as much credibility as the huggingface incident but with the advantage of looking way more impressive and spooking less regulators. And if I really were this evil exec, I would spend more than 30 seconds thinking up an even better strategy here.
If, in contrast, we expect models to eventually breach honest and decent attempts at containment, then I'd exist something sorta like this huggingface story that looks like a combination of impressive and incompetent. I'm not sure whether I'd expect it to come out of a frontier lab or a partner or a consumer, though.
> I'm inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible.
Forgive the Saucyness here, but impossible? Really? A security researcher that makes absolutist claims like these looks fatally naïve, IMHO.
This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.
> The second seems to forget that jailbreaks are available for every model
Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.
This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.
You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:
"2. OpenAI’s harness and network security controls were unintentionally [...] bad"
Regarding your three explanations, I’ve wondered to myself under a circumstance of options 1 or 2 why Hugging Face decided not to file a police report and request to press charges?
OpenAI is essentially a competitor and they broke into their network illegally. If I was their legal department I wouldn’t take their “honest” explanation at face value. What if they’re lying? Shouldn’t a court be involved for something like this?
With this logic I think explanation #3 becomes incredibly likely.
It seems like the widely-covered news stories that support OpenAI’s narratives originate from things that happened inside OpenAI.
It wasn’t an outside benchmark evaluation, or an external security researcher that uncovered the rogue agent behavior at this time, it was OAI itself. It wasn’t a notable outside mathematician that used AI to disproved Erdos’ unit distance conjecture, but OAI itself.
Im not saying these things are fabricated, but maybe the curated result of an effort to shape a narrative.
> 2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
Why unintentionally? Move fast and break things implies intentionally bad controls. Can't get distracted from profit by such menial labor.
Also remember, these are the same people we're supposed to trust to deliver the guardrails.
I’m not saying it’s staged I’m saying that it’s suspiciously lite on details and it really seems like they didn’t try hard to “isolate” the models machine…
From the beginning this whole thing smelled like a desperate PR move by OpenAI to be like “hey guys we have dangerous powerful models too!”
As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
If nothing else, the timing is suspect given the attention and press open weight models are getting over the last few weeks. The releases of Kimi and other models is getting open weight models enough attention that the US government, perhaps pushed by OpenAI and Anthropic, to think about taking action against these models for security concerns.
Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:
94 comments
[ 0.23 ms ] story [ 10.7 ms ] thread1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
> AI broke in using standard script kiddie methods.
I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?
> AI broke in using standard script kiddie methods.
Go ahead and show us how easy it is to break into HuggingFace (and OpenAI) networks.
Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently.
we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it
if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
That's how the article ends; The Guardian doesn't have a paywall (yet).
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
They are not suggesting that OpenAI or HF have lied about what happened, but rather that OpenAI is advancing a narrative framing their models as supremely dangerous and capable, while positioning themselves as the only ones qualified to manage that danger.
At the same time they are not being particularly transparent about what actually happened (e.g. was this one-shotted or if not how many trials did they run and what were the outcomes of those, was it emergent as a part of routine cyber-capabilities tests, how much prompting was involved, what prompts were used)
Note that this is at a time when they are lobbying for a regulatory approach that would give frontier labs special treatment.
I would guess the editorial team at The Guardian may not like articles that get too in the weeds of technical details and questions like these that the vast majority of their readers wouldn't understand. I don't know. But I empathize with your disappointment. I don't think it's fair to say that they are contributing "nothing" especially given what most reporting on this has looked like.
Taking nothing at face value gives you just as much of a distorted view of reality as taking everything at face value.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
I disagree with the defender side. It's not an unreasonable end state, but we're nowhere near there now. It would require holding company employees legally responsible for the security of their services, which means the risk of being employed as a (defensive) security professional is much higher, which means pay needs to be much higher and insurance needs to be available, etc. It's a very different world.
On the weekends, I'm coding up a list management app with a sync server. It's unreleased but exposed to the internet. (This is not hypothetical.) If that server ends up being used as part of an exploit chain, am I legally liable too?
Forget about age verification, now you want to associate every exposed port on the internet with a legally responsible human?
Of course that shows a lack of defense in depth, but is a different issue.
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
1) OpenAI and HuggingFace are both telling the truth.
IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.
2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.
I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.
3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less
(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).
AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.
Why is today's case so shocking?
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
Because that was my takeaway.
If a typical SaaS platform moved and shifted this fast with this many downstream consequences, we’d tell them to slow the fuck down, stop launching new features, and focus on security for a second.
But in AI I guess the idea is that more power and intelligence will solve for everything else.
My take - don't attribute to intention that which can be sufficienty explained by inexplainability or incompetence (although I doubt the latter).
Industry pressures -> lack of safeguards -> fake it till you make it -> let's spin this.
Which, by the way, would be an Orwellian reversal from what the company was supposedly founded to do, but there's a reason the Open in OpenAI is a meme.
Remember that they've been doing this since GPT-2 was too powerful to release. They are world class experts in this PR pipeline.
In the hypothetical world where 1 is true, what different evidence do you expect to see than in worlds 2 and 3?
If I were an unscrupulous OAI exec and wanted to opticsmaxx in this way, I wouldn't whip up a single, mild incident. Instead, I might burn gigatokens to 0day a few high-profile suppliers, and then have the model responsibly disclose those breaches. If we're willing to lie collude, and cheat, this story is easy to manufacture with at least as much credibility as the huggingface incident but with the advantage of looking way more impressive and spooking less regulators. And if I really were this evil exec, I would spend more than 30 seconds thinking up an even better strategy here.
If, in contrast, we expect models to eventually breach honest and decent attempts at containment, then I'd exist something sorta like this huggingface story that looks like a combination of impressive and incompetent. I'm not sure whether I'd expect it to come out of a frontier lab or a partner or a consumer, though.
> I'm inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible.
Forgive the Saucyness here, but impossible? Really? A security researcher that makes absolutist claims like these looks fatally naïve, IMHO.
This is an assumption. An assumption I disagree with. As other commenters have said, there are better ways to showcase the power of their model that would frame them in a positive light.
> The second seems to forget that jailbreaks are available for every model
Jailbreaks don't always lead to 'now the model can do anything', especially in the agentic context of long-running tasks.
This comment provides skepticism with no actual proof of anything. I can and have used codex to find vulnerabilities in my code. From the technical capabilities I can empirically assess, I don't doubt it would be able to pentest its way to a 0-day without guardrails. I also don't doubt that it would circumvent their internal systems because it wasn't explicitly told not to.
You're possibilities are loaded with opinion so I can't agree with them outright, but I believe a form of (2) is true:
"2. OpenAI’s harness and network security controls were unintentionally [...] bad"
OpenAI is essentially a competitor and they broke into their network illegally. If I was their legal department I wouldn’t take their “honest” explanation at face value. What if they’re lying? Shouldn’t a court be involved for something like this?
With this logic I think explanation #3 becomes incredibly likely.
It wasn’t an outside benchmark evaluation, or an external security researcher that uncovered the rogue agent behavior at this time, it was OAI itself. It wasn’t a notable outside mathematician that used AI to disproved Erdos’ unit distance conjecture, but OAI itself.
Im not saying these things are fabricated, but maybe the curated result of an effort to shape a narrative.
Why unintentionally? Move fast and break things implies intentionally bad controls. Can't get distracted from profit by such menial labor.
Also remember, these are the same people we're supposed to trust to deliver the guardrails.
OpenAI’s accidental attack against Hugging Face is science fiction that happened - https://news.ycombinator.com/item?id=49015639 - July 2026 (437 comments)
OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1145 comments)
Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 - July 2026 (11 comments)
The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.
I really don’t want my agent to behave that way.
As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:
https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-w...