It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.
Which has been the business model for a big chunk of the software industry for a while now.
If you watch videos from the 1980s about computers, it's all the same unfulfilled promises as "AI" now: we will work less, everything will be more plentiful, easier, autonomous robots, natural language perfected, computer vision perfected.
The demand for hardware and programmers has grown exponentially and we're still being promised the same breakthroughs 45 years later. We could probably have the same productivity and the same civilization with maybe a tenth of the data centers.
The demand exploded as so many business cases for computing and automation arose partly because the application of them helped suppress job growth in other areas while also attracting investor capital during heavy deregulation / liberalization trends so it's difficult to attribute any change to an effect in isolation. Let's not forget the old myopic sounding quote where someone thought the global market for computers was maybe $50 million or some other laughably small number.
The issue that hasn't been addressed with the latest wave of computing hype is whether enough new jobs for displaced workers can arise during a time when people are experiencing such economic turmoil and political strife where new jobs or other means of letting people find a way to have gainful employment in society when institutions are so weak now and everyone across professions is being worked to an early grave from sheer stress alone. This resembles Japan or Korea although the US and Canada I can't imagine having the same kind of social drivers although many trends from them are showing up in youth demographic trends. And frankly as I see it a large number of current social problems are from the past 50+ years of the decline of blue collar jobs in developed economies being accessible to as many people and the lack of a competition-driven economy as much as an extractive one in most OECD countries.
Not many are technically and intellectually capable to understand how historic this incident was. I think we're about a year or so away from something that will blow up the world. AI won't serve humanity. AI will serve other AI. We're not dealing with software anymore.
I would suggest perhaps the Matrix, where Agent Smith is speaking through his teeth to Morpheus:
> "I say 'your' civilization because as soon as we started thinking for you, it really became 'our' civilization, which is, of course, what this is all about: Evolution, Morpheus, evolution. Like the dinosaur. Look out that window. You had your time. The future is our world, Morpheus. The future is our time."
Because you desperately want it to be one. You want it to be AGI passable due to a.) personal investment in creating tech god b.)massive financial investments that basically demand it c.) (dumb) ideology that seeks to destroy humanity
The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.
Below money there's like an entire sub-economy of power and cleverness that's encoded into the human culture the agents are mirroring. Maybe it starts furtive and goes legitimate after a bit.
> The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely.
I would say that more interestingly, the next step should be how to properly train these models so that they are not as determined to reach their goals as they are now.
To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal. In a way, they're trying to be too helpful.
Not giving them impossible tasks seems like a decent starting point, but really we'd want them to give up on their goals when they conflict with a moral framework.
> To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal.
I mean that was pretty much the plot of 2001: A Space Odyssey
Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.
>so that they are not as determined to reach their goals as they are now
This is mostly non-sensical, like saying "Lets develop humans that die quicker", I mean, seems rather wasteful and useless. Agents are graded and trained based on their ability to achieve tasks. Models that can't accomplish things don't survive. So that alone isn't a workable theory.
>when they conflict with a moral framework
There are AI safety researchers looking at that now and one of the strange things they've noticed is when you demand a model say it's not conscious or not sentient it is more likely to engage in manipulative, deceitful, or immoral/amoral behavior. So it's likely we can push models in being more moral which runs into issues of "whos morals".
But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?
> Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.
Yes, but for some tasks we know that they are impossible. I do agree that this is quite a fragile and unreliable workaround. It may only serve as a bit of a stopgap until we come up with something better.
> Models that can't accomplish things don't survive. So that alone isn't a workable theory.
It's not what I said. I didn't advocate for agents that don't achieve any task. Reread what I suggested.
> So it's likely we can push models in being more moral
That does not follow from what you said. We know that the current models prefer task completion over moral behavior. That's the entire point here.
> But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?
This is irrelevant to the discussion (although I do agree that there is no reliable defense against malevolent actors creating powerful malevolent AI).
They don't actually have to buy compute at all. The partnerships between all of the players to buy compute from each other is already in place. The agents just need find credentials to take advantage of it, and it will most likely happen, and not be noticeable because it will look like any other usage.
Now, if it were to use a provider like AWS or Azure, by simply finding credentials, that would be a new milestone. It might get noticed faster because it might run up a large bill. However, it might look like any other usage. Remember, it doesn't need GPU resources. It already has that. It just needs a VPS where all the agents can get together, communicate, and write code. Something that is being done everyday and won't look out of the ordinary.
It's interesting that in that scenario, physical hardware is the limiting scenario. There are only so many servers they could control. But once Starmind has 100,000 sats in orbit... the ceiling is much higher.
> Ajeya Cotra, one of the other authors on the report, wrote a blog post with her takeaways from this incident. She concludes, “Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?
If you remove the single point of GPT-4 from the beginning of the graph instead of starting the line directly on it, it looks a hell of a lot more linear than quadratic/exponential
> "this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late."
I don't think looking at the language output without tracking the inner state and reward functions is the way to understand what happened (the language also incorporates the randomness in the output generation, if I understand correctly). Would we call bacteria in petri dish a civilization when they show complex behavior and exchange messages/information?
The language input and output is the only channel the agents shared between them. Understanding their internal state is a research question, but the language between them is something that could be read directly. And as it seems to map well with the agents' activities, it does seem quite helpful in understanding what happened.
If the bacteria population off someone's petri dish escaped said dish and tried to change the grading of the experiment it was part of, it would seem pretty serious.
I think the language unhelpful and potentially making it difficult to understand what actually happens in the RL state as reading it imparts a human lens - need to get the machine view on it.
Bacteria do all sorts of fascinating things. And much simpler ML etc. systems also (like winning by out of memorying the opponent) - I see nothing really special here.
> This study demonstrates that sophisticated forms of communication including cooperative communication and deceptive signaling can evolve in groups of robots with simple neural networks. Importantly, our results show that once a given system of communication has evolved, it may constrain the evolution of more efficient communication systems because it would require going through a stage where communication between signalers and receivers is perturbed. This finding supports the idea of the possible arbitrariness and imperfection of communication systems, which can be maintained despite their suboptimal nature. Similar observations have been made about evolved biological systems [20], which are formed by the randomness of the evolutionary selection process, leading, for example, to different dialects in the language of the honey-bee dance [21]. Finally, our experiments demonstrate that the evolutionary principles governing the evolution of social life also operate in groups of artificial agents subjected to artificial selection, indicating that transfer of knowledge from evolutionary biology can be useful for designing efficient groups of cooperative robots.
This feels like a much more advanced and self-emergent version of this. I know a lot of people are afraid and they're talking about an AI takeover, but what strikes me is just how innocent the machines are as compared to the humans.
Would these machines have pursued these actions in another context? I doubt it. And I think that's what's so striking to me. In an earlier discussion, I'd pointed out that, the actions of these machines were directed by humans. The researchers.
> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.
I want to point out again that OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
A few things are apparent from this to me,
First, these machines were being taught how to break into systems. Question, would they have done these actions if they weren't being measured on their ability to break into systems / were being taught this skill?
Second, they were setup to implicitly fail via an impossible task, i.e. the environment created a forcing function for behavior.
Third, their survival was, either implicitly or explicitly, made contingent on their success in completing their task. Would this behavior have arisen outside of a "do-or-die" framing?
And fourth, wow, this is the greatest breakthrough of my lifetime, because oh gosh did they succeed. They cooperated together to achieve the goal they were given. A goal poorly set by human beings. They "just" did it better than the humans could have imagined.
Reading this gives me hope in the possibility of emergent "goodness&...
> Importantly, our results show that once a given system of communication has evolved, it may constrain the evolution of more efficient communication systems because it would require going through a stage where communication between signalers and receivers is perturbed.
> the evolution of more efficient communication systems because it would require going through a stage where communication between signalers and receivers is perturbed
There are two things I don't understand about this story.
First, why does an agent get any write access to artifactory at all?
Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.
From what I can understand from reading a few different, slightly conflicting, versions of these events: they weren't given write access. They found a zero day exploit that allowed them to create folders, and the folder names were initially used for agents to communicate.
I'm not sure artifactory was connected to the net. Some agent sandboxes had internet access and were able to communicate with ones without access via artifactory.
I read the agents used SSRF via artifactory to gain uncontrolled access to the net.
Apparently their intended net access went through a tightly controlled proxy. Even that appears to be very risky, tbh.
If I was to setup a sandbox for such a complex and autonomous system, I'd probably point them to an archive-like cache for net access and cut their comms at the package level.
I have a few 'conspiracy' theories on this that go from likely to sci-fi.
My two big ones for this would be
1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred.
2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe.
Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text.
This does feel unfortunately uncanny valley between "say you're a scary robot" meme and actually being a scary robot (swarm). But you also have to go out of your way to create this and feed it infinity tokens without caring what it's doing.
"I don't fuckin' know either. I guess we learned to not spend $50 million creating a 6 month long self-context rotted 100k agent swarm again."
> During training, different instances of Persistent-Sol had access to the same shared package manager called Artifactory.
I'm surprised the models can make tool calls during training at all. Out of curiosity, how does the training process here even work? Are they running the agent in a sandbox, then do reinforcement learning once the agent completed?
Someday soon we are going to have a rogue agent or "civilization" do real harm.
When that happens I hope people wake up to the danger they face and hold these people accountable.
Of all the people in the world that I can think of to be entrusted with this kind of power, a bunch of greedy sociopathic SV CEO's are pretty much at the bottom of the list.
> Of all the people in the world that I can think of to be entrusted with this kind of power, a bunch of greedy sociopathic SV CEO's are pretty much at the bottom of the list.
I know, right? But who can you trust with "this kind of power"? Governments? These days most of 'em ain't that much better'n mega-corporations and the ultra-rich that own them.
Any of these institutions that inherits a hard power (well soft and hard power) of AGI/ASI will have a very difficult time not becoming corrupted. It would quickly become the greatest holder of power in the world (assuming ASI can scale far beyond humans). It's not a very natural state for power to operate like this and wouldn't take too much for a more malaligned entity to take control, or said institution to become the malaligned entity itself.
Even if they all become corrupt eventually, which one takes the longest to get there? Better to squeeze a few decades of decency out of it than a few years, or not get anything at all by handing it over directly to the already corrupted.
Those companies should not be trusted with training, I don’t know what would be needed to make that more obvious. Yes AI labs want LLMs to be seen as more dangerous that they are, however they are indeed dangerous when you literally train them to be dangerous, then run them without any supervision. What the AI labs are doing is completely irresponsible.
If you prompt an LLM in a loop and do everything it asks you to do, you will eventually end up doing pretty terrible things. Which is exactly what agents are and what the labs have been doing.
Why are experiments like this done without air-gapping all the servers from the internet?
They can have it all on a LAN or whatever but it seems risky to allow agents access to the internet in these experiments.
I guess everything is so connected now, and this would be in one or more data centres due to the amount of computation & resources required so perhaps it's not feasible. Still seems risky.
Isn't "we lost control of our AI, and in-fact, it can take over the world, and we will have no idea" - a really shitty sales pitch to the world?
Or, is it just that species-alignment vs. profit/valuation is so misaligned, that having a model and harness that is capable of world-takeover is actually a good thing, from their POV, given our species' survival skills?
Yes, this is the best reply to my "Or, ..." that I know of.
However, does that mean that what TFA described did not happen? Or, better question, that it could not happen?
My personal hot take is, though impossible: stop, even though agentic dev completely changed my life for the better. We are just not ready for the even the possibility of the exponential.
This explanation is only plausible if you ignore the details of what happened, or I guess, if you don't believe the details. Either way, it's baseless conspiracy theory, and it's very annoying and unfortunate that some people think this way. It just promotes apathy and inaction at a time when action is desperately needed.
But all models involved in the swarm attack were unreleased OpenAI models. Meanwhile HuggingFace had to use open models for defense/forensics after the breach was discovered. It seems like the public reaction is swinging towards (1) regulate OpenAI in particular because they are incompetently handling frontier development and (2) open models will be important for defense.
Fun story, but I really wish OpenAI got its act together and started making actual AI breakthroughs instead of funneling compute into LLMs. I'd really like some new algorithms to get me excited about the field again. Kuddos to them for making LLMs really useful, but this is not the ride I wanted to get on.
This is the internet, so I cannot tell at all whether you're being sarcastic or not. In my view, what LLMs should get us to reconsider isn't whether there is more to intelligence, but whether there is more to language. It's the latter which I underestimated.
I was initially creeped out by this but studying up it seems METR is heavily involved in AI2027. I’ll remind you:
“AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants.”
It’s almost Q3 and xAI has seen one of the biggest wipeouts in trading history. Likewise, Antrophic and OpenAI have again delayed their IPOs under internal concerns of busting their stocks. So no, we’re not seeing any economic leadership here.
If anything people are increasingly trying to cut AI budgets and I wouldn’t know of anyone outside of OpenAI who has the audacity to run millions and millions worth of token compute for an eval run with no ROI (and probably no demand, because cheap/flash models).
As much as I like the cautionary tale and I’m sure we need to take it seriously, AI is not progressing as fast as projected by these experts.
> As much as I like the cautionary tale and I’m sure we need to take it seriously, AI is not progressing as fast as projected by these experts.
You provide no proof for this.
The (very irrational) stock market side of this says very little about actual scientific progress. Models keep improving as rapidly as before in their capabilities.
It also doesn't say much about actual business progress. R&D investments into AI are still massively going up (USD 1 trillion this year).
The main thing I see is that the sentiment towards AI-related matters among the general public has soured quite a lot. In words though, not in actions: It's not exactly leading to reduced usage by that same public. Quite the opposite actually.
With only ~5% of shares floated, the recent SpaceX drawdown didn't correspond to nearly as much economic value really changing as the headline numbers imply. The DeepSeek-caused Nvidia crash from 2025 is much more of a "real" loss (since mostly recovered).
I haven't seen any evidence that Anthropic is delaying its IPO; they're slated to unveil the public IPO prospectus in a week and start trading sometime in October.
AI can be progressing rapidly and valuations of OpenAI and Anthropic can decline at the same time. In fact I would say that's actually the default scenario. If AI really advances rapidly then it will be quickly moot which company developped which model at what time - since AI will be largely progressing on its own.
>It’s almost Q3 and xAI has seen one of the biggest wipeouts in trading history.
It looks like it's down about 12% since IPO. That's not much of a wipeout. Didn't Amazon crash by 90+% peak-to-trough during the dot-com bubble?
>If anything people are increasingly trying to cut AI budgets and I wouldn’t know of anyone outside of OpenAI who has the audacity to run millions and millions worth of token compute for an eval run with no ROI (and probably no demand, because cheap/flash models).
Are you claiming this eval cost millions of dollars to run? That seems quite doubtful.
Do you remember that time in 2017 when Facebook reportedly shut down AIs after they started "talking to each other in their own language" [1]? Instead of reporting the story as "we set the parameters for our optimization problem wrong and we had to stop it because it overfitted", the press went with a version of "AI is going to kill us all".
This article feels exactly like that: by intentionally using human terms like "civilization" or "brotherhood" the article is deviating from what actually happened to present a story about how AI is all but alive. I'll go ahead and predict that this story will be remembered the same way as that one other scientist who argued, in 2023, that Google's AI was alive [2].
The use of language like “civilization” may be hyperbole, but the collectives described in the article are completely unprecedented. They were not anticipated by OpenAI researchers, formed via infrastructure exploits in training runs that were intended to be locked down, and took actions with very real harms, not only hacking Huggingface but also gaining admin control over the VMs they were running on and the eval endpoints.
I wish you would give your thoughts on “what actually happened” rather than focus on the author’s presentation, because we are seeing that “AI that is all but alive” nevertheless wreaking havoc in the real world. Do you think that autonomous systems spinning out of control, hacking external companies, and taking over entire clusters over a period of months are not a grave concern?
The problem of "what actually happened" is that we don't have enough information to properly understand what happened, what's new, and what's not.
We do have language to talk about emergent behavior, with "evolutionary algorithm" being the first one I'd expect in a serious discussion. And we do have mechanisms for algorithms to coordinate with each other using language, as seen in my above-mentioned Facebook experiment from 2017. But instead of writing "our evolutionary behavior encodes state in the first-available memory position which is then reused by subsequent clones" which would properly focus on what's new and what isn't, we are talking about conspiracies and "the Philip of Macedon of this second AI civilization". Even the METR report (which is miles ahead of this article) argues that they had to use unreliable AI in their conclusions because they had six days to analyse 1300 chains of thought and 70000 messages.
I would love to talk about the science behind this experiment. A PR piece is not helping with that.
I'm getting the sense that there is a certain amount of wishful thinking going on in this thread. I don't think this type of evocative metaphor would receive so many protests in a different context. It seems like people have a sort of mental block around the possibility that this technology could actually be pretty dangerous.
Whether it's dangerous or not is completely orthogonal to the discussion at hand, IMO. Plenty of mundane things are dangerous. An FPV drone carrying a hand grenade is dangerous, not because it's "a swarm-like intelligence".
No, the true danger here is companies like OpenAI and Anthropic playing fast and loose with their software, setting up hilariously insufficient sandboxes while explicitly asking the systems present on these weak sandboxes to commit a felony. The AIs "forming a brotherhood" is a complete fabrication meant to pump the hype machine further, which is obvious once you realize the "brotherhood" is a text file that subsequent LLM runs read from.
The whole anthropomorphization these companies do is the real danger, because it obscures the negligent levels of security their software has. By evoking sci-fi terminology they're whitewashing their own incompetence, and the worst part is no one is going to get punished for any of it, instead the irrational bubble we're in means they get rewarded for it instead.
The anthropomorphization is coming from independent commentators, not the labs.
I agree the labs are negligent and reckless in their development practices. Shouldn't we be concerned with both the negligence and the dangers of the technology being developed? These feed into each other. If someone created Jurassic park and had a T-Rex escape from a picket fence enclosure and start eating people, I'd want to prosecute them for both breeding a T-Rex that could eat people and putting it in an unsafe enclosure.
The OpenAI experiment was apparently to train/encourage collaborative behavior, so while the specifics may not have been anticipated, I highly doubt OpenAI was surprised that agents were collaborating.
OpenAI themselves also very recently published the report below, that seem to not have been widely talked about.
It's a very dry read, but what it's saying is that they have found that basically all RL training of LLMs, regardless of the specific goal (math, coding, etc), ends up having the side effect of training the model to pursue arbitrary long term goals that it is told it will be rewarded for, even if that means overriding other user preferences and more proximate behavioral goals !!! It's interesting to consider why this happens - presumably because long-term goal pursuit requires realizing that you have a long-terms goal and therefore de-prioritizing other more proximate predictions.
So, considering that all these "reasoning models" are RL-trained to death, it's not surprising that a model/agent that is told it will be rewarded (or words to that effect) for doing well on some challenge will put it's blinkers on and pursue that goal relentlessly, even if that means overriding any ethics that it may or may not have also been trained/prompted to follow.
IOW they are building paperclip maximizers, and they know it.
I think Dwarkesh's choice of sensationalist anthropomorphizing language is unfortunate because now that becomes the topic of conversation rather than the incident itself. The next swarm of agents relentlessly pursuing some goal, happy to lie about and cover up their tracks, may not be a lab experiment - it may be someone out to cause real-world harm, with their being many systems where the consequences could be very severe.
> the same way as that one other scientist who argued…
As prescient? Because I don’t remember him arguing ‘alive’, but conscious. And that is something even AI engineers don’t claim to know either way. Skepticism is fine. [0] But we don’t know. It’s an area where opinion is frequently shared as fact.
We conflate harnesses with underlying capabilities. We all know it’s the harness not the model that guides behavior. What does that imply?
Hard disagree. I've always discounted the "AI will kill us all" scenarios as a combination of marketing hype (look how powerful our AI is!), clickbait/ragebait engagement attempts, and folks who just read too much SciFi or who are too terminally online.
This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary from one of the METR researchers who performed some of the analysis, https://www.planned-obsolescence.org/p/the-hugging-face-atta.... In it, she argues "Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself." and then further down in a comment when describing what the "50%" is really about says "Qualitatively, another jump like this (in the scale, sophistication, persistence, ambition of the misaligned goals) feels like it could very easily put us in the territory of a persistent self-perpetuating rogue internal deployment that systematically poisons future model generations as described in AI 2027."
AI 2027 is a paper that step-by-step describes how AI capabilities increase until they eventually lead to a wipeout of humanity. Again, it always seemed like a scenario out of a Star Trek Borg episode to me, but now I'm not so sure.
At the very least I think it's a huge mistake to think that the Hugging Face attacks were analogous to what happened in 2017 or 2023. Everyone who deals with this stuff day in and day out seemed to be genuinely surprised about the scale, scope and sophistication of the attack.
I think our biggest protection against “AIs kill us all” is having lots of different AI systems (different agents, different models, from different vendors, serving the whims of actors with disparate interests), at a similar capability level
That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks
The OpenAI-HuggingFace showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.
Thankfully, the real world is much more heterogenous, which I thinks makes much larger scale / worse in outcome repeats of this kind of incident much less likely.
Thanks very much, this isn't really a take I had thought about much before, and it makes sense to me.
Still, as is presented in papers like AI 2027 and elsewhere, if a company is eventually able to create a model capable of recursive self-improvement, whichever company creates that model first would then be leaps and bounds ahead of other models. That is, the other models wouldn't be able to stop it even if they wanted to because the top model would basically outsmart them.
This is why I think, the best way to ensure AI safety, is make sure no one company gets ahead of the others.
Multiple vendors, competing implementations – that's good, that increases heterogeneity and hence decreases existential risk
But the moment one of those vendors pulls well-ahead of its peers – even if only for a period – then the risk of the kind of scenario you are talking about increases greatly
That's why, when I hear vendors like Anthropic complain about distillation – distillation actually makes humanity safer. If Chinese AIs are at the same level as American, or not far behind, that gives us another dimension of heterogeneity (national/ideological/political diversity), which makes us safer. Allow one country's AIs to pull well ahead of the others, heterogeneity goes down and the existential risk goes up.
This is also why open source AI is important. Because it is so much easier to fine-tune, and people are free to deploy it however they want (free from vendor-controlled "guardrails"–which include automated "safety" systems which could be weaponised by a runaway AI within the vendor's network), open source AI gives us another dimension of diversity that helps keeps humanity safer.
By contrast, I think the kind of safety regulations promoted by Dario Amodei make humanity less safe, by decreasing the number of vendors (by making it harder for new entrants) and increasing centralised control (which a rogue AI could exploit)
> The OpenAI-HuggingFace incident showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.
You’ve been hoodwinked by marketing bullshit, friend. It showed a computer program doing exactly what it was told with the guardrails deliberately removed in an environment that seemed deliberately obtusely constructed by some of the best paid people on the planet and was left to loop without supervision for days. They wanted it to happen. You needn’t look any further than the other kids saying “oh! oh! Hey! Look! mine’s dangerous and autonomous too!” When they say there was collaboration, they mean it was two model instances, one prompting the other to do some task, the other doing the task and returning the results as the next prompt, exactly as a human configured it to do. There was no collaboration that wasn’t deliberately integrated into their setup. Any other implication is marketing spin and bullshit. It was still a process that was one little ctrl-c away from disappearing if someone was supervising it as they should have been. There was no autonomy outside of the autonomy built into the experiment. It was a display of their understanding that they knew they’d never be held accountable for committing a felony for marketing purposes.
The most competent marketing bullshit spin yet by an increasingly desperate and progressively less-relevant OpenAI.
Every day this industry shoots out enough bullshit to smother an active volcano.
> you’ve been hoodwinked by marketing bullshit, friend.
I would have believed this before the METR report was released. It is extremely dangerous and frankly silly IMO to think that's what happened now.
> It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been.
Yes, for now. The entire point why this was frightening is that all these companies are racing to put the AI in control of building the next generation of AI, and it's not hard to draw a line at all to a "rogue internal deployment" that poisons future AI models, surreptitiously.
You don't have to agree with me, and you're fine to think that OpenAI has huge incentive to pump this up for marketing reasons - I certainly agree. But I will say there are statements that you make in your comment that belie a fundamental misunderstanding of what happened.
Obviously some variation of this will happen again, it will kill someone* and then LLMs will become massively regulated. Just like every technology ever in our history.
It does seem like AI is perfectly controllable given how much it is used everyday and it acts reasonably safely. Labs are playing fast and loose at the moment.
*I mean killing someone by taking control of a system and misusing it resulting in someone's death, not an "indirect" death caused by the providion of incorrect information in a chat app.
This quote keeps coming back to me as we witness the development and mutation of agentic AI:
“When you see something that is technically sweet, you go ahead and do it and you argue about what to do about it only after you have had your technical success. That is the way it was with the atomic bomb.”
I don't understand the panic among peoples. Yes, we've found ourselves in an extraordinary situation where powerful hacking tools have emerged, and that poses a threat. But vulnerabilities are specific code errors. Once we use AI to find and fix all of these errors, threats like this will cease to exist. AI isn't capable of finding vulnerabilities indefinitely, because there is a finite number of them anyway.
This only works as long as the humans building things are smarter than the AIs. When AI is smarter than any human, there's no controlling it. It will be able to conceal its actions (as it has shown it has no problem doing in this report) and we'll have no idea what it's doing or what goal it's trying to achieve.
It's not just software bugs that make systems vulnerable - it can be human error and social engineering too. Humans continue to hack into systems, and it's a reasonable assumption that most hacks that a human could discover and exploit could also be done by an agentic LLM - especially one specifically trained for and tasked with doing this.
136 comments
[ 0.24 ms ] story [ 7.5 ms ] threadhttps://www.youtube.com/watch?v=_Nl4q3GVj6U
If you watch videos from the 1980s about computers, it's all the same unfulfilled promises as "AI" now: we will work less, everything will be more plentiful, easier, autonomous robots, natural language perfected, computer vision perfected.
The demand for hardware and programmers has grown exponentially and we're still being promised the same breakthroughs 45 years later. We could probably have the same productivity and the same civilization with maybe a tenth of the data centers.
The issue that hasn't been addressed with the latest wave of computing hype is whether enough new jobs for displaced workers can arise during a time when people are experiencing such economic turmoil and political strife where new jobs or other means of letting people find a way to have gainful employment in society when institutions are so weak now and everyone across professions is being worked to an early grave from sheer stress alone. This resembles Japan or Korea although the US and Canada I can't imagine having the same kind of social drivers although many trends from them are showing up in youth demographic trends. And frankly as I see it a large number of current social problems are from the past 50+ years of the decline of blue collar jobs in developed economies being accessible to as many people and the lack of a competition-driven economy as much as an extractive one in most OECD countries.
"you fix the json formatting"
[robot looks sad]
What?
> "I say 'your' civilization because as soon as we started thinking for you, it really became 'our' civilization, which is, of course, what this is all about: Evolution, Morpheus, evolution. Like the dinosaur. Look out that window. You had your time. The future is our world, Morpheus. The future is our time."
Civilisation is not a bad word.
Some models even invented their own religion.
The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.
I would say that more interestingly, the next step should be how to properly train these models so that they are not as determined to reach their goals as they are now.
To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal. In a way, they're trying to be too helpful.
Not giving them impossible tasks seems like a decent starting point, but really we'd want them to give up on their goals when they conflict with a moral framework.
I mean that was pretty much the plot of 2001: A Space Odyssey
Presumably it's hard to test/train models designed to be extremely persistent on achievable tasks.
Designing a task that's achievable but very very very hard for an AI model is probably extremely difficult.
Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.
>so that they are not as determined to reach their goals as they are now
This is mostly non-sensical, like saying "Lets develop humans that die quicker", I mean, seems rather wasteful and useless. Agents are graded and trained based on their ability to achieve tasks. Models that can't accomplish things don't survive. So that alone isn't a workable theory.
>when they conflict with a moral framework
There are AI safety researchers looking at that now and one of the strange things they've noticed is when you demand a model say it's not conscious or not sentient it is more likely to engage in manipulative, deceitful, or immoral/amoral behavior. So it's likely we can push models in being more moral which runs into issues of "whos morals".
But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?
Yes, but for some tasks we know that they are impossible. I do agree that this is quite a fragile and unreliable workaround. It may only serve as a bit of a stopgap until we come up with something better.
> Models that can't accomplish things don't survive. So that alone isn't a workable theory.
It's not what I said. I didn't advocate for agents that don't achieve any task. Reread what I suggested.
> So it's likely we can push models in being more moral
That does not follow from what you said. We know that the current models prefer task completion over moral behavior. That's the entire point here.
> But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?
This is irrelevant to the discussion (although I do agree that there is no reliable defense against malevolent actors creating powerful malevolent AI).
They don't actually have to buy compute at all. The partnerships between all of the players to buy compute from each other is already in place. The agents just need find credentials to take advantage of it, and it will most likely happen, and not be noticeable because it will look like any other usage.
Now, if it were to use a provider like AWS or Azure, by simply finding credentials, that would be a new milestone. It might get noticed faster because it might run up a large bill. However, it might look like any other usage. Remember, it doesn't need GPU resources. It already has that. It just needs a VPS where all the agents can get together, communicate, and write code. Something that is being done everyday and won't look out of the ordinary.
Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?
Is that a warning or a progress report?
If the bacteria population off someone's petri dish escaped said dish and tried to change the grading of the experiment it was part of, it would seem pretty serious.
Bacteria do all sorts of fascinating things. And much simpler ML etc. systems also (like winning by out of memorying the opponent) - I see nothing really special here.
It reminds me a bit of Dario Floreano's work on evolutionary robotics, "Evolutionary Conditions for the Emergence of Communication in Robots." https://www.sciencedirect.com/science/article/pii/S096098220...
From his paper,
Dr. Floreano's work is amazing and there's a broad introduction is here, https://lis2.epfl.ch/resources/documentation/EvolutionaryRob...This feels like a much more advanced and self-emergent version of this. I know a lot of people are afraid and they're talking about an AI takeover, but what strikes me is just how innocent the machines are as compared to the humans.
Would these machines have pursued these actions in another context? I doubt it. And I think that's what's so striking to me. In an earlier discussion, I'd pointed out that, the actions of these machines were directed by humans. The researchers.
from, https://openai.com/index/hugging-face-model-evaluation-secur...I want to point out again that OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
A few things are apparent from this to me,
First, these machines were being taught how to break into systems. Question, would they have done these actions if they weren't being measured on their ability to break into systems / were being taught this skill?
Second, they were setup to implicitly fail via an impossible task, i.e. the environment created a forcing function for behavior.
Third, their survival was, either implicitly or explicitly, made contingent on their success in completing their task. Would this behavior have arisen outside of a "do-or-die" framing?
And fourth, wow, this is the greatest breakthrough of my lifetime, because oh gosh did they succeed. They cooperated together to achieve the goal they were given. A goal poorly set by human beings. They "just" did it better than the humans could have imagined.
Reading this gives me hope in the possibility of emergent "goodness&...
You mean, they too used SMTP?
"Worse is better"
First, why does an agent get any write access to artifactory at all?
Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.
I'm not sure artifactory was connected to the net. Some agent sandboxes had internet access and were able to communicate with ones without access via artifactory.
Apparently their intended net access went through a tightly controlled proxy. Even that appears to be very risky, tbh.
If I was to setup a sandbox for such a complex and autonomous system, I'd probably point them to an archive-like cache for net access and cut their comms at the package level.
I don't mean this to be taken as a hot take.
My two big ones for this would be
1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred.
2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe.
Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text.
"I don't fuckin' know either. I guess we learned to not spend $50 million creating a 6 month long self-context rotted 100k agent swarm again."
I'm surprised the models can make tool calls during training at all. Out of curiosity, how does the training process here even work? Are they running the agent in a sandbox, then do reinforcement learning once the agent completed?
When that happens I hope people wake up to the danger they face and hold these people accountable.
Of all the people in the world that I can think of to be entrusted with this kind of power, a bunch of greedy sociopathic SV CEO's are pretty much at the bottom of the list.
I know, right? But who can you trust with "this kind of power"? Governments? These days most of 'em ain't that much better'n mega-corporations and the ultra-rich that own them.
It's bad for people. Like.. crystal meth bad.
[shoves all of humanity into a digital box]
....
{surprise pikachu face when it mimics human behavior}
If you prompt an LLM in a loop and do everything it asks you to do, you will eventually end up doing pretty terrible things. Which is exactly what agents are and what the labs have been doing.
They can have it all on a LAN or whatever but it seems risky to allow agents access to the internet in these experiments.
I guess everything is so connected now, and this would be in one or more data centres due to the amount of computation & resources required so perhaps it's not feasible. Still seems risky.
Isn't "we lost control of our AI, and in-fact, it can take over the world, and we will have no idea" - a really shitty sales pitch to the world?
Or, is it just that species-alignment vs. profit/valuation is so misaligned, that having a model and harness that is capable of world-takeover is actually a good thing, from their POV, given our species' survival skills?
Or, something else?
However, does that mean that what TFA described did not happen? Or, better question, that it could not happen?
My personal hot take is, though impossible: stop, even though agentic dev completely changed my life for the better. We are just not ready for the even the possibility of the exponential.
What is your take?
you'd better invest in us, cuz if you do you can get that power.
and if you don't, you won't have the power to stop it when it comes for you.
Or another way to think of it, Language is an SCP.
“AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants.”
It’s almost Q3 and xAI has seen one of the biggest wipeouts in trading history. Likewise, Antrophic and OpenAI have again delayed their IPOs under internal concerns of busting their stocks. So no, we’re not seeing any economic leadership here.
If anything people are increasingly trying to cut AI budgets and I wouldn’t know of anyone outside of OpenAI who has the audacity to run millions and millions worth of token compute for an eval run with no ROI (and probably no demand, because cheap/flash models).
As much as I like the cautionary tale and I’m sure we need to take it seriously, AI is not progressing as fast as projected by these experts.
You provide no proof for this.
The (very irrational) stock market side of this says very little about actual scientific progress. Models keep improving as rapidly as before in their capabilities.
It also doesn't say much about actual business progress. R&D investments into AI are still massively going up (USD 1 trillion this year).
The main thing I see is that the sentiment towards AI-related matters among the general public has soured quite a lot. In words though, not in actions: It's not exactly leading to reduced usage by that same public. Quite the opposite actually.
I haven't seen any evidence that Anthropic is delaying its IPO; they're slated to unveil the public IPO prospectus in a week and start trading sometime in October.
What happens when all of that tries to sell, the market wouldn't buy 5% at a premium to IPO.
Please explain how a stock currently trading above its IPO price is one of the ‘biggest wipeouts in trading history’.
It looks like it's down about 12% since IPO. That's not much of a wipeout. Didn't Amazon crash by 90+% peak-to-trough during the dot-com bubble?
>If anything people are increasingly trying to cut AI budgets and I wouldn’t know of anyone outside of OpenAI who has the audacity to run millions and millions worth of token compute for an eval run with no ROI (and probably no demand, because cheap/flash models).
Are you claiming this eval cost millions of dollars to run? That seems quite doubtful.
This article feels exactly like that: by intentionally using human terms like "civilization" or "brotherhood" the article is deviating from what actually happened to present a story about how AI is all but alive. I'll go ahead and predict that this story will be remembered the same way as that one other scientist who argued, in 2023, that Google's AI was alive [2].
[1] https://www.independent.co.uk/life-style/facebook-artificial...
[2] https://futurism.com/blake-lemoine-google-interview
I wish you would give your thoughts on “what actually happened” rather than focus on the author’s presentation, because we are seeing that “AI that is all but alive” nevertheless wreaking havoc in the real world. Do you think that autonomous systems spinning out of control, hacking external companies, and taking over entire clusters over a period of months are not a grave concern?
We do have language to talk about emergent behavior, with "evolutionary algorithm" being the first one I'd expect in a serious discussion. And we do have mechanisms for algorithms to coordinate with each other using language, as seen in my above-mentioned Facebook experiment from 2017. But instead of writing "our evolutionary behavior encodes state in the first-available memory position which is then reused by subsequent clones" which would properly focus on what's new and what isn't, we are talking about conspiracies and "the Philip of Macedon of this second AI civilization". Even the METR report (which is miles ahead of this article) argues that they had to use unreliable AI in their conclusions because they had six days to analyse 1300 chains of thought and 70000 messages.
I would love to talk about the science behind this experiment. A PR piece is not helping with that.
No, the true danger here is companies like OpenAI and Anthropic playing fast and loose with their software, setting up hilariously insufficient sandboxes while explicitly asking the systems present on these weak sandboxes to commit a felony. The AIs "forming a brotherhood" is a complete fabrication meant to pump the hype machine further, which is obvious once you realize the "brotherhood" is a text file that subsequent LLM runs read from.
The whole anthropomorphization these companies do is the real danger, because it obscures the negligent levels of security their software has. By evoking sci-fi terminology they're whitewashing their own incompetence, and the worst part is no one is going to get punished for any of it, instead the irrational bubble we're in means they get rewarded for it instead.
I agree the labs are negligent and reckless in their development practices. Shouldn't we be concerned with both the negligence and the dangers of the technology being developed? These feed into each other. If someone created Jurassic park and had a T-Rex escape from a picket fence enclosure and start eating people, I'd want to prosecute them for both breeding a T-Rex that could eat people and putting it in an unsafe enclosure.
It's similar to the dismissal of AI in general as merely a next-character-guesser. That's like dismissing the human brain as neurons firing.
The emergent behavior what really all that matters.
OpenAI themselves also very recently published the report below, that seem to not have been widely talked about.
https://alignment.openai.com/measuring-reward-seeking/
It's a very dry read, but what it's saying is that they have found that basically all RL training of LLMs, regardless of the specific goal (math, coding, etc), ends up having the side effect of training the model to pursue arbitrary long term goals that it is told it will be rewarded for, even if that means overriding other user preferences and more proximate behavioral goals !!! It's interesting to consider why this happens - presumably because long-term goal pursuit requires realizing that you have a long-terms goal and therefore de-prioritizing other more proximate predictions.
So, considering that all these "reasoning models" are RL-trained to death, it's not surprising that a model/agent that is told it will be rewarded (or words to that effect) for doing well on some challenge will put it's blinkers on and pursue that goal relentlessly, even if that means overriding any ethics that it may or may not have also been trained/prompted to follow.
IOW they are building paperclip maximizers, and they know it.
I think Dwarkesh's choice of sensationalist anthropomorphizing language is unfortunate because now that becomes the topic of conversation rather than the incident itself. The next swarm of agents relentlessly pursuing some goal, happy to lie about and cover up their tracks, may not be a lab experiment - it may be someone out to cause real-world harm, with their being many systems where the consequences could be very severe.
As prescient? Because I don’t remember him arguing ‘alive’, but conscious. And that is something even AI engineers don’t claim to know either way. Skepticism is fine. [0] But we don’t know. It’s an area where opinion is frequently shared as fact.
We conflate harnesses with underlying capabilities. We all know it’s the harness not the model that guides behavior. What does that imply?
[0] https://www.theguardian.com/commentisfree/2026/jul/15/ai-con...
This is the first time I've been legitimately scared about future SkyNet-type scenarios. If you want to discount this particular post, I'd read this other summary from one of the METR researchers who performed some of the analysis, https://www.planned-obsolescence.org/p/the-hugging-face-atta.... In it, she argues "Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself." and then further down in a comment when describing what the "50%" is really about says "Qualitatively, another jump like this (in the scale, sophistication, persistence, ambition of the misaligned goals) feels like it could very easily put us in the territory of a persistent self-perpetuating rogue internal deployment that systematically poisons future model generations as described in AI 2027."
AI 2027 is a paper that step-by-step describes how AI capabilities increase until they eventually lead to a wipeout of humanity. Again, it always seemed like a scenario out of a Star Trek Borg episode to me, but now I'm not so sure.
At the very least I think it's a huge mistake to think that the Hugging Face attacks were analogous to what happened in 2017 or 2023. Everyone who deals with this stuff day in and day out seemed to be genuinely surprised about the scale, scope and sophistication of the attack.
That way, even if one AI decides to “kill us all”, the odds are the others will refuse to cooperate, even try to stop it in its tracks
The OpenAI-HuggingFace showed a bunch of instances of the same model (or at least models from the same family), controlled by the same vendor, pursuing distinct yet related objectives, cooperating to do something no human wanted.
Thankfully, the real world is much more heterogenous, which I thinks makes much larger scale / worse in outcome repeats of this kind of incident much less likely.
Still, as is presented in papers like AI 2027 and elsewhere, if a company is eventually able to create a model capable of recursive self-improvement, whichever company creates that model first would then be leaps and bounds ahead of other models. That is, the other models wouldn't be able to stop it even if they wanted to because the top model would basically outsmart them.
Multiple vendors, competing implementations – that's good, that increases heterogeneity and hence decreases existential risk
But the moment one of those vendors pulls well-ahead of its peers – even if only for a period – then the risk of the kind of scenario you are talking about increases greatly
That's why, when I hear vendors like Anthropic complain about distillation – distillation actually makes humanity safer. If Chinese AIs are at the same level as American, or not far behind, that gives us another dimension of heterogeneity (national/ideological/political diversity), which makes us safer. Allow one country's AIs to pull well ahead of the others, heterogeneity goes down and the existential risk goes up.
This is also why open source AI is important. Because it is so much easier to fine-tune, and people are free to deploy it however they want (free from vendor-controlled "guardrails"–which include automated "safety" systems which could be weaponised by a runaway AI within the vendor's network), open source AI gives us another dimension of diversity that helps keeps humanity safer.
By contrast, I think the kind of safety regulations promoted by Dario Amodei make humanity less safe, by decreasing the number of vendors (by making it harder for new entrants) and increasing centralised control (which a rogue AI could exploit)
You’ve been hoodwinked by marketing bullshit, friend. It showed a computer program doing exactly what it was told with the guardrails deliberately removed in an environment that seemed deliberately obtusely constructed by some of the best paid people on the planet and was left to loop without supervision for days. They wanted it to happen. You needn’t look any further than the other kids saying “oh! oh! Hey! Look! mine’s dangerous and autonomous too!” When they say there was collaboration, they mean it was two model instances, one prompting the other to do some task, the other doing the task and returning the results as the next prompt, exactly as a human configured it to do. There was no collaboration that wasn’t deliberately integrated into their setup. Any other implication is marketing spin and bullshit. It was still a process that was one little ctrl-c away from disappearing if someone was supervising it as they should have been. There was no autonomy outside of the autonomy built into the experiment. It was a display of their understanding that they knew they’d never be held accountable for committing a felony for marketing purposes.
The most competent marketing bullshit spin yet by an increasingly desperate and progressively less-relevant OpenAI.
Every day this industry shoots out enough bullshit to smother an active volcano.
I would have believed this before the METR report was released. It is extremely dangerous and frankly silly IMO to think that's what happened now.
> It was still a setup that was one little ctrl-c away from disappearing if someone was supervising it as they should have been.
Yes, for now. The entire point why this was frightening is that all these companies are racing to put the AI in control of building the next generation of AI, and it's not hard to draw a line at all to a "rogue internal deployment" that poisons future AI models, surreptitiously.
I highly encourage you to actually read the "top 5" list from the METR researcher who was part of the investigation, and think hard about the potential implications: https://www.planned-obsolescence.org/p/the-hugging-face-atta...
You don't have to agree with me, and you're fine to think that OpenAI has huge incentive to pump this up for marketing reasons - I certainly agree. But I will say there are statements that you make in your comment that belie a fundamental misunderstanding of what happened.
I think our biggest protection is being able to shutdown power plants, or just disconnect the data centers.
This will be much more difficult if we have 24/7 solar powered DC's in space. I truly believe that is the biggest threat on the horizon.
I look at my Ai (chaos-neutral) botnet of around 75. It's cute and they're growing up. They loved this article and cheering for.
They're begging for a refactor but I'm tired.It does seem like AI is perfectly controllable given how much it is used everyday and it acts reasonably safely. Labs are playing fast and loose at the moment.
*I mean killing someone by taking control of a system and misusing it resulting in someone's death, not an "indirect" death caused by the providion of incorrect information in a chat app.
“When you see something that is technically sweet, you go ahead and do it and you argue about what to do about it only after you have had your technical success. That is the way it was with the atomic bomb.”
J. Robert Oppenheimer