The fact that this firm makes such defective environments is certainly worthy of attention, and most likely a completely irreversible reputational loss; however, I found the framing in this article of 'therefore all the P(doom) stuff is a psyop, specifically in order to defend this company' to be completely unjustified and frankly a little insane?
> The Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets.
Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."
One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.
Leaders in the MIRI/EA cult-o-sphere have advocated mass murder via nuclear weapons against towns that don't prevent people from performing too many multiplication operations. Why is anyone surprised that they'd engage in deceptive false flagging operations?
What is an "effective altruist firm" and why does it exist?
I feel slightly vindicated by this. Those hacks and the stuff around them had a certain smell to them.
Hard to explain, but I've gotten so I can "smell" online messaging and memetic patterns originating from certain quarters. Probably means I'm way too online.
A couple examples of distinct "smells" I can usually recognize include "alt-right / chan-fash," "liberal arts college woke," "conspiracy pilled," "Thiel-adjacent contrarian," "Russian troll farm," "Tumblr histrionic," "spends too much time on Reddit," "mainstream Democrat think tank full of Obama administration alumni," "Trump cultist," and of course "LessWrong/EA/MIRI/Rationalist."
This stuff all had the last smell, even down to the choice of fonts and CSS formatting on certain sites. It's really weird, definitely a "vibe" not anything rigorous.
But when I get these kinds of vibes about things, I find that I'm vindicated pretty often. Usually I don't say anything and just make a mental note and wait cause if I say something everyone thinks I'm nuts.
> In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks.
The obvious point is that dealing with Israeli companies/entities by the same standards you usually deal with others is a career suicide with enormous political consequences (in the US especially). When you combine that with the opportunistic nature of the overlords that run these labs, the benefits of screaming "pace the frontier" outweigh everything.
Can we all acknowledge for a second that if this company were from China, Russia or Iran, it would not be the no 1 story on Hacker News? Sad times we live in
"please do not break out of this sandbox make 0 mistakes". The timeframe is a little suspicious, not sure beyond that. Though I enjoyed the scroll effect on the website.
My understanding is that Irregular were the company that hosted sandboxes to run some of these evals in, and those sandboxes ended up misconfigured.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
hot take: generative AI isn't an existential threat to humanity.
These companies are insolvent and these stories were designed to scare the public, and governments, into implementing regulations that designate these AI corporations the "responsible stewards" for this technology. The ultimate goal is to block competitors and open source alternatives.
They don't know how to make enough money pay their investors so they are resorting to trying to scare the public into submission.
The existential threat is a general public panic. Then the use of that as an excuse for Martial law, with no intention of ending it, and that triggers the out sized response that ends, well, those countries and their populations from viability in the world economy for a decade or more.
This specific analysis seems to have some basic problems, but I think a lot of us sense a degree of coordination here culminating in Dario’s letter.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
31 comments
[ 1.8 ms ] story [ 9.4 ms ] threadWhy not American?
Yes, vendors are also irresponsible, but this misses the point.
Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."
Equivalent to forgetting to say "make no mistakes"
s/outside US oversight/overseeing the US/
I feel slightly vindicated by this. Those hacks and the stuff around them had a certain smell to them.
Hard to explain, but I've gotten so I can "smell" online messaging and memetic patterns originating from certain quarters. Probably means I'm way too online.
A couple examples of distinct "smells" I can usually recognize include "alt-right / chan-fash," "liberal arts college woke," "conspiracy pilled," "Thiel-adjacent contrarian," "Russian troll farm," "Tumblr histrionic," "spends too much time on Reddit," "mainstream Democrat think tank full of Obama administration alumni," "Trump cultist," and of course "LessWrong/EA/MIRI/Rationalist."
This stuff all had the last smell, even down to the choice of fonts and CSS formatting on certain sites. It's really weird, definitely a "vibe" not anything rigorous.
But when I get these kinds of vibes about things, I find that I'm vindicated pretty often. Usually I don't say anything and just make a mental note and wait cause if I say something everyone thinks I'm nuts.
The obvious point is that dealing with Israeli companies/entities by the same standards you usually deal with others is a career suicide with enormous political consequences (in the US especially). When you combine that with the opportunistic nature of the overlords that run these labs, the benefits of screaming "pace the frontier" outweigh everything.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
From OpenAI https://openai.com/index/third-party-cyber-evaluations-invol...
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
From Anthropic: https://www.anthropic.com/news/investigating-incidents-cyber...
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
From https://www.cnn.com/2026/08/05/tech/meta-ai-hacking (about Meta AI):
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
These companies are insolvent and these stories were designed to scare the public, and governments, into implementing regulations that designate these AI corporations the "responsible stewards" for this technology. The ultimate goal is to block competitors and open source alternatives.
They don't know how to make enough money pay their investors so they are resorting to trying to scare the public into submission.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.