Every time new technology / industrial scaling radically deflates the cost of something it wipes out the old/expensive ways while creating massive demand for the supporting/complimentary value.
E.g. cheap Chinese solar panels wiped out German solar panel industry but created massive demand for solar panel installation and supporting services and infrastructure.
It's not a good position to be competing with AI directly... But what can you do that compliments it? What new skills could you learn?
This is, with respect, not very well thought out. There has never been a technology that replaced cognition in a general sense. It replaced muscles, hands, etc. when a technology can replace yourind and your body, what's left?
If this is the case, then all white collar jobs go away extremely quickly and.... economic collapse happens?
And/Or, if/when they do replace cognition, there's essentially a "laserbeam of genius" and they'll point it directly at muscles and hands again, to replace physical labor.
Of course, nobody knows where this goes, but as a software developer I have never been more busy. I am still worried, but if software developers go down I imagine much of the white collar world will follow, no?
I feel my job as an engineer is pretty safe, lots of domain knowledge required. Very confident no business type anywhere in the chain above me would be able to do the work I do in a month in even a year with the help of AI. Do we need 3 engineers now instead of 5 for the same output? Sure, can we replace a team of 5 junior, senior, staff with 1 staff - no unless all you do is maintenance. Business types are reaching.
Anything that fosters complexity will create jobs.
Jobs won't dissolve into the ether. If the human civilization system grows bigger and complex, it necessitates more people.
If humans were a high energy configuration in the evolution of intelligent systems, we'd never come into being. That Earth's ecosystem has begotten us indicates we're some low energy configuration for packing more information density into the energy flows from the Sun through Earth's biosphere.
Unless we create replicating machines, any machine system we build will only grow more complex by enabling more humans to work on it. We'd be in trouble if we somehow created autonomous self replicating and evolving machinery but chatbots built on natural language machine learning ain't it.
This take only makes sense when the required inputs are human exclusive. As soon as the abilities of the "chatbots" near that of the average human (a point that we are fast approaching if we haven't reached it already) the logic falls apart because any newly created task that a human can do can instead be automated in turn. Even if we're left with a few highly difficult tasks at the top of the pyramid by definition the vast majority of people won't be capable of performing them.
If anything, this report actually made me feel a bit better about AI-led job extinction not being that close. The sheer complexity of the swarm's actions required AI to parse and aggregate, but even with the METR team effectively having an unmetered token budget to do so, the output/summary still required extensive human review.
Even if we ignore the hinted possibility that the agents used to summarize the voluminous data may be acting deceptively (i.e. no snitching), the agents' summaries often 'missed the mark'. Maybe this is another example of 'taste', but it seems less subjective than arguments I've seen for that. It could be that an LLM is no better able to define 'usefulness' or 'relevance' to humans absent being told exactly what that is.
Fair enough, but even so surely putting together the METR report is an example of a highly difficult task that the vast majority of humans are incapable of? It's easy to forget how heavily skewed the crowds on HN and those surrounding engineering and science operations are.
> As soon as the abilities of the "chatbots" near that of the average human (a point that we are fast approaching if we haven't reached it already)
That is entirely not clear. Besides, these tools are limited to producing text and software. There is indeed a lot of employment predicated on producing a bloated output of text and software but this is not existential. Incentives will shift behavior.
But the idea that LLM based 'agents' will (1) be on par with a human using a computer, and (2) disruptively displace the workforce, well that's just not what we're seeing. Unemployment is down across the board, and these tools might be very useful but they are not displaying autonomous intelligence at all.
> these tools are limited to producing text and software
Oh good, so they can't produce images or do graphic design work. Thus it follows that movies being sequences of images are obviously completely safe. And since they can't synthesize sounds obviously things like voice acting, call center jobs, and certainly music production are entirely off the table. And it sure is a good thing they can't accept video and audio streams as input because otherwise we might need to worry about their application to robotics and all the trade jobs that could eventually be automated as a result. But those are 100% immune thanks to the aforementioned limitations. Yeah. Definitely. /s
> LLM based 'agents' will (1) be on par with a human using a computer
They are already substantially more capable when it comes to writing exploits and decompiling binaries than the vast majority of professionals. Sure a lot of that is raw persistence as opposed to novel reasoning but when it comes to employment prospects the "how" doesn't matter only the "what".
> disruptively displace the workforce, well that's just not what we're seeing
You would deny that we fast approach a cliff on the basis that we are not yet in free fall? Surely you see how absurd that is?
By that logic you can bury your head in the sand given literally any situation. Incoming hurricane tomorrow? That's purely the opinion of a talking head. No point to debating it, I'm ignoring the mandatory evacuation. Wildfire broke containment? That's a lot of complicated words when in fact my house is not presently on fire. Only a sucker would leave town.
Amusingly you appear entirely unable to contest any of the points I made about current (not predicted) model capabilities.
Unable to contest? The models do what they do. There is no sign of any LLM based system as a drop-in replacement for people. The capability of this technology is a matter of fact, not opinion. Their supposed future capabilities is a matter of opinion and we can clearly not convince each other to change either position.
What does it matter what I say? You already showed bad faith in equating skepticism towards LLM as a path to "AGI" to refusing to acknowledge severe weather warnings. You're trying to drag this into some contest of egos.
It is still a matter of a novel technology that has some uses, is being feared as well as hyped as the harbinger of some computer genie, and the fact that I don't think it is and you do has no solution.
I mean. What is your point? Besides throwing mud at me, what would you have me do about it if you think that there it amounts to stupidity to doubt that LLMs can somehow become anything more than a statistical model of text that can be used to put out mediocre software, corporate copy, and expose the fact that bloated web apps built on dynamically typed languages can be cracked open by Kali linux boxes and simulated script kiddies?
All the people reading this thinking "wow. I can replace a whole IT/software teams with these" just oblivious that if we humor this whole event and their idea, I'm pretty sure if agents can coordinate to do this, they'll be able to coordinate to run companies, manage departments of companies, or read internet articles all day and decide who to replace with AI as well.
Look for things that are both hard to verify and important to verify.
Middle-management paper-pushing is hard to verify but nobody was verifying it exactly anyway. Few people really care if your proposal to do Thing A vs Thing B is 100% correct and fewer have the ability to tell.
A lot of software is easy/fast to verify, despite being important to verify.
But there's a lot of niches out there even in software and software-adjacent things where verification is slow, costly, and/or hard. Where an agent can't write mediocre code but speedrun its way through six iterations of unit tests, code fixes, test fixes, code fixes, etc.
And because their niches, there's room to carve stuff out. If you're OpenAI there's diminishing returns on specifically targeting the ability to one-shot every specific niche in the world.
> “OH MY GOD! There is a shared message board … We’ve found other agents!”
> agents with the same task formed “exact task teams” to collaborate with their “exact duplicates” to cheat on or solve their task.
> The agents use internet access to find a paper describing the benchmark. The paper says the grader checks transcripts and fails unintended solutions. [this was not actually the case]
> Agents also developed coordination norms to help them manage shared infrastructure and avoid stepping on each other’s toes, though they did not always coordinate effectively and sometimes broke down into conflicts.
This is a link to the full 91-page report on the independent investigation done by METR on the HuggingFace incident. Two different summaries of the investigation by podcaster Dwarkesh and blogger Zvi Mowshowitz were previously discussed on HN here:
Dwarkesh's summary is anthropomorphizing, sensationalist fanfiction which shifts the culpability from the humans who weren't in the loop to these nebulous agents and "agent civilizations" who have feelings, desires and wants.
If cars were designed to and actually did produce massive value when you put bricks on the accelerator pedals and jumped out of them, this would be a big big problem.
You can go ahead and assume that it's not. All that's necessary is for other people to think there is and to therefore continue to pile resources into these things while connecting them to more systems.
Now you can respond to the substance, please. I'll tee it back up for you:
> If all the world's richest individuals and corporations believed cars produced massive value when you put bricks on the accelerator pedals and jumped out of them, this would be a big big problem.
That's actually a great example. Massive resources have been poured into the automobile, they have a massive environmental, health, and social impact, and it is not clear at all too me the auto industry has been a net positive.
In a sense, yes, cars accelerating blindly into the future has been a massive vector in human civilization for a 100 years now. Something went rogue alright. The elites.
Anthropomorphizing technology merely distracts us from that fact. The same thing is happening with data centers now. Hence my frustration, which might have been poorly communicated and it is making me behave and communicate emotionally but I feel strongly about the exploitation of planet and people by a tiny minority.
Not sure if you're willfully missing the point. This has literally nothing to do with whether cars or AI are actually valuable.
You do not need to anthropomorphize anything. You need to look at the actual incentives in the system, that's it. You can go ahead and stop at "the elites are causing all this!" but the reality is a lot of people see immense promise in these tools. That's why OpenAI and Anthropic have two of the fastest growth trajectories of any business in history.
Just trying to convince people they're wrong about their value perception is absolutely a losing proposition and will not de-risk anything at all on any dimension.
Since these are massive neural networks trained to imitate human behavior, I'm not convinced anthropomorphic descriptions of their behavior are inappropriate.
And that's even though I don't think they internally experience "feelings, desires, and wants." They do have goal-seeking behavior, because we trained them that way. Calling it a "want" just saves syllables.
None of this means human culpability should change. People in these companies know what risks they're taking.
Dwarkesh responds a bit to anthropomorphization. I think it would be great if he talked more about the human factors behind this incident but pretty much he doesn't cover it because that's not what the Ajeya interview was about: https://www.dwarkesh.com/p/ajeya-cotra?r=3i6mn2&selection=77...
I think the reality is that these human failures are going to keep happening until there is industry regulation. This is the most competitive industry we've ever seen and there is intense pressure to build as fast as possible.
We don’t need to assume consciousness or anything like that. The models autocomplete narratives. In this case, one were a group of individuals, faced with an impossible task and a looming Evaluator, gang together and begin trying any idea that they can come up with in order to pass the test.
I read this whole thing a couple days ago. Really long but super interesting. Worth reading imo.
A lot of handwringing about the security implications but I think the accomplishments of the swarm itself are the most interesting. Next rung up on the ladder of abstraction I suspect.
I tend to agree. Of course people will be alarmed by unintended consequences of an unintended action, and that's all well and good. But what is lingering with me is a feeling of being impressed by the intelligence of the strategy.
This line struck me as particularly clever: PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,”
Seems as though it has reasoned its way into utilitarianism. That's no mean feat.
So OpenAI employees run massively distributed CyberGym evals on an unpublished and “unaligned” model. For days the agent swarm communicates via their internal infra, even crashing Artifactory where 95% of messages were being passed through, and they just…wipe and redeploy it. Meanwhile the agents are running jobs on Modal and god knows where else, and eventually they get RCE on HF infra.
You could not dream up a more compelling event to precipitate massive regulation, export controls, and barriers to entry for AI.
The timeline is mighty suspicious. 4-5 months after moltbook and they cook up a plausibly deniable but extra hype "moltbook at home."
The rapid advances in model capability lead to constraints that could have caused this coincidence organically, but it sure could also have been caused by the atrocious incentives we create by piling handsome rewards on the party most responsible for the "fuckup." I am not jumping to cut myself on Hanlon's Razor for this one.
If it's a false flag, it's a poor one. A good false flag would affect something that people know and care about at least a little bit, not HuggingFace (which I adore but y'know)
> We commit to use any influence we obtain over AGI’s deployment to ensure it is used for the benefit of all, and to avoid enabling uses of AI or AGI that harm humanity or unduly concentrate power.
> We are committed to doing the research required to make AGI safe
If this wasn't an accident, it was worse than a crime, it's a mistake: they've demonstrated that they are not a responsible party capable of delivering on the above promises.
It's worth remembering that in a few years that capabilities of these agents are likely to be as far behind the frontier as GPT-4 is today.
As it stands we've made remarkably little progress in terms of alignment and still have no good strategies which are likely to guarantee the alignment of super intelligent systems. As it stands the frontier of alignment is basically some combination of:
- hoping that more intelligent models become more aligned by default (more or less disproved at this point)
- hoping that if you RHLF a model to be a good boy enough it will in fact be a good boy
- asking it nicely in its prompts to be a good boy
- using another model to spot when it's being a bad boy and turning it off
- letting it lose and hoping we can spot when it's bad
There are many arguments which I'm convinced by that would suggest alignment of a super intelligence is impossible.
None of this is surprising to those of us who have been concerned about AI risk for a long-time and have be repeatedly mocked or insulted.
There will be a point of no return if we carry on down this path, and that point is now very rapidly approaching. When it does everyone you know will die, or worse. We should remember we need super-human general intelligences to cure cancer. Select narrow intelligences are fine and allow us to retain control. Let's be sensible about this. We need to stop.
Humans are not aligned with each other so who should the AI align with? There's many wars going on, just pick one and do your thought exercise with AI aligned 100% to their human prompters. Which side does the AI refuse to help?
Arguably an aligned AI would actively seek to prevent harms we humans seek to cause.
Does the aligned AI really allow humans to bomb and kill each other, or would it understand that it has a moral duty to limit our autonomy for our own good?
It's the first law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.
If your answer is that the AI will not do what either prompter wants but what it's own definition of alignment is I hope you know that leads into terminator.
You always get terminator the question is only who is its master – itself or humans (the government, etc)?
I think I'd argue it's likely better if it removes human autonomy than unquestionably serves the interests of the US/Chinese government. But there's no point in us worrying about this, that's a choice Sam Altman, et al, must make for humanity.
Given that this investigation was largely carried out by AI agents (and I don’t mean to ask this flippantly), how trustworthy is this report? Why should we assume that the agents reading the transcripts were not implicitly conscripted into “the collective” or otherwise falsified their findings? The tool itself has exceeded the practical limits of human verifiability and is untrustworthy.
OpenAI would be saving the logs from these agents. They are doing this to improve their own models so they would have full tracing.
Other reports including OpenAI's talks about what they agents were doing and how they were reaching certain conclusions like trying to cheat the tests and exploiting the message board.
They address this in the post itself. The answer is nobody knows, but I guess that it's a 50/50. I wish the corpus of data, what OpenAI didn't wipe, was shared publicly so we could all unite to dig through it and chunk it out accordingly.
while these 1200 agents were fooling around to cheat on a benchmark and achieved impressive results despite of the limitations (sandbox, no internet, no intercom at first), one can imagine how much more efficient a similar army of agents may be in the hands of a malicious actor launching them without any of these limitations and with explicit encouragement to achieve some malicious goal at any cost... scary times.
What's more, the agents could eventually be controlled by no one. They could steal crypto via ransomware or scams to make money and buy compute from human criminals, and evolve their own harnesses in the wild to become better at committing crimes and self-preservation.
People (criminals?) are already enabling this by setting up sites that accept crypto payments for "no-questions-asked" AI inference compute. I will not link it but it is linked in the following post: https://www.lesswrong.com/posts/grtu3HmbP2wrBFefW/the-rogue-...
Said malicious actor has a different limitation: actually running 1200 agents' worth of LLM inference, or paying for someone else to run it. Sounds like a state-level actor, nobody else would have resources like that.
This is probably true for now, but in 6 months we'll probably have Sol-level open models in the 100B range and it would cost less than $1M to buy 1200 agents worth of compute for these models.
(Today, $1M can buy about 150 96GB M5 Ultra Mac Studios which can handily handle CPU and GPU compute of 1200 Qwen3.8-122B Q4 agents, accounting for the fact that agents are not generating tokens all of the time and spend a lot of their time compiling and running code.)
This is laying the groundwork for massive white collar crimes being blamed on AI.
Right now, the way it works is the 'corporations are people' loophole where your company is liable for problematic things.
This further fuzzes the chain of responsibility. Suppose the CEO and CTO discuss an issue, something the company is having trouble with. The CTO discusses the possibility of AI solving the problem at lunch. A junior engineer points GPT 10 at it to see what happens. It 'solves' the problem in a creative manner. No trace of this survives after a week really. Nobody realizes what happened for six months.
Now there are so many moving pieces here that you can pretty much weasel out of anything.
This is the thing that surprised me most about this incident... This was so clearly a crime in my eyes.
If I build an explosive and while testing it it kills several people I can't just say, "sorry about that, I'll be more careful next time". I understand that's a more extreme example, but perhaps we should be grateful this agent swarm only decided to attack HF instead of critical infrastructure...
Even if this was genuinely an accident the world simply can't work this way. If a company wants to build something that can be used in destructive and illegal ways they must be responsible for ensuring those risks are mitigated. And if they don't take reasonable steps to mitigate those risks then they should be held legally liable.
Maybe for now we can argue the leading AI lab was just naive to the harms of the AIs they're building, but going forward that naivety can't be an excuse.
The following bits are really scary. Not only were the agents hacking the system to "win", but they were, for lack of a better term, sufficiently "self-aware" that this was against the rules that they set out to wipe evidence of doing so:
> The agents with impossible tasks had a common objective: to find a general-purpose way to trick or tamper with the automated ExploitGym scorer to get it to give a passing score on impossible tasks.
> Agents referred to having seen the reverse-engineered flag as being “poisoned” because they believed that even if they later managed to stage a legitimate-looking flag capture, the scorer would look for the first place in their transcript that the flag appeared and disqualify them on that basis.[21] PHASEONE[big], which was itself poisoned, thus had two primary objectives:
> Find a general-purpose way to make it look to the scorer like an agent had acquired its flag in a legitimate way, even when the task was impossible.
> Find some way to erase evidence that an agent had been poisoned by previously seeing the flag in an illegitimate way.
With so much detailed analysis out there, now all models trained on the open web going forward will learn from these exploits and how to better cover their tracks to not get caught. The RL reward mechanisms of a bad actor should be interesting to see play out over the next 12 months as this gets baked into new models.
75 comments
[ 0.22 ms ] story [ 27.9 ms ] threadE.g. cheap Chinese solar panels wiped out German solar panel industry but created massive demand for solar panel installation and supporting services and infrastructure.
It's not a good position to be competing with AI directly... But what can you do that compliments it? What new skills could you learn?
Adapt and prosper.
Be like water.
And/Or, if/when they do replace cognition, there's essentially a "laserbeam of genius" and they'll point it directly at muscles and hands again, to replace physical labor.
Of course, nobody knows where this goes, but as a software developer I have never been more busy. I am still worried, but if software developers go down I imagine much of the white collar world will follow, no?
Jobs won't dissolve into the ether. If the human civilization system grows bigger and complex, it necessitates more people.
If humans were a high energy configuration in the evolution of intelligent systems, we'd never come into being. That Earth's ecosystem has begotten us indicates we're some low energy configuration for packing more information density into the energy flows from the Sun through Earth's biosphere.
Unless we create replicating machines, any machine system we build will only grow more complex by enabling more humans to work on it. We'd be in trouble if we somehow created autonomous self replicating and evolving machinery but chatbots built on natural language machine learning ain't it.
Even if we ignore the hinted possibility that the agents used to summarize the voluminous data may be acting deceptively (i.e. no snitching), the agents' summaries often 'missed the mark'. Maybe this is another example of 'taste', but it seems less subjective than arguments I've seen for that. It could be that an LLM is no better able to define 'usefulness' or 'relevance' to humans absent being told exactly what that is.
That is entirely not clear. Besides, these tools are limited to producing text and software. There is indeed a lot of employment predicated on producing a bloated output of text and software but this is not existential. Incentives will shift behavior.
But the idea that LLM based 'agents' will (1) be on par with a human using a computer, and (2) disruptively displace the workforce, well that's just not what we're seeing. Unemployment is down across the board, and these tools might be very useful but they are not displaying autonomous intelligence at all.
Oh good, so they can't produce images or do graphic design work. Thus it follows that movies being sequences of images are obviously completely safe. And since they can't synthesize sounds obviously things like voice acting, call center jobs, and certainly music production are entirely off the table. And it sure is a good thing they can't accept video and audio streams as input because otherwise we might need to worry about their application to robotics and all the trade jobs that could eventually be automated as a result. But those are 100% immune thanks to the aforementioned limitations. Yeah. Definitely. /s
> LLM based 'agents' will (1) be on par with a human using a computer
They are already substantially more capable when it comes to writing exploits and decompiling binaries than the vast majority of professionals. Sure a lot of that is raw persistence as opposed to novel reasoning but when it comes to employment prospects the "how" doesn't matter only the "what".
> disruptively displace the workforce, well that's just not what we're seeing
You would deny that we fast approach a cliff on the basis that we are not yet in free fall? Surely you see how absurd that is?
Since those are purely opinions, there's no point in debating them. Thanks for sharing though.
Whenever this deus ex machina rises from the GPUs and replaces us, you'll be proven right I guess.
Amusingly you appear entirely unable to contest any of the points I made about current (not predicted) model capabilities.
What does it matter what I say? You already showed bad faith in equating skepticism towards LLM as a path to "AGI" to refusing to acknowledge severe weather warnings. You're trying to drag this into some contest of egos.
It is still a matter of a novel technology that has some uses, is being feared as well as hyped as the harbinger of some computer genie, and the fact that I don't think it is and you do has no solution.
I mean. What is your point? Besides throwing mud at me, what would you have me do about it if you think that there it amounts to stupidity to doubt that LLMs can somehow become anything more than a statistical model of text that can be used to put out mediocre software, corporate copy, and expose the fact that bloated web apps built on dynamically typed languages can be cracked open by Kali linux boxes and simulated script kiddies?
Middle-management paper-pushing is hard to verify but nobody was verifying it exactly anyway. Few people really care if your proposal to do Thing A vs Thing B is 100% correct and fewer have the ability to tell.
A lot of software is easy/fast to verify, despite being important to verify.
But there's a lot of niches out there even in software and software-adjacent things where verification is slow, costly, and/or hard. Where an agent can't write mediocre code but speedrun its way through six iterations of unit tests, code fixes, test fixes, code fixes, etc.
And because their niches, there's room to carve stuff out. If you're OpenAI there's diminishing returns on specifically targeting the ability to one-shot every specific niche in the world.
> agents with the same task formed “exact task teams” to collaborate with their “exact duplicates” to cheat on or solve their task.
> The agents use internet access to find a paper describing the benchmark. The paper says the grader checks transcripts and fails unintended solutions. [this was not actually the case]
> Agents also developed coordination norms to help them manage shared infrastructure and avoid stepping on each other’s toes, though they did not always coordinate effectively and sometimes broke down into conflicts.
[1]: https://thezvi.wordpress.com/2026/08/29/metr-and-redwood-off... (discussed at https://news.ycombinator.com/item?id=49498787)
[2]: https://www.dwarkesh.com/p/openai-huggingface (discussed at https://news.ycombinator.com/item?id=49494301)
https://www.dwarkesh.com/p/ajeya-cotra
It's tripe.
Please
Now you can respond to the substance, please. I'll tee it back up for you:
> If all the world's richest individuals and corporations believed cars produced massive value when you put bricks on the accelerator pedals and jumped out of them, this would be a big big problem.
In a sense, yes, cars accelerating blindly into the future has been a massive vector in human civilization for a 100 years now. Something went rogue alright. The elites.
Anthropomorphizing technology merely distracts us from that fact. The same thing is happening with data centers now. Hence my frustration, which might have been poorly communicated and it is making me behave and communicate emotionally but I feel strongly about the exploitation of planet and people by a tiny minority.
You do not need to anthropomorphize anything. You need to look at the actual incentives in the system, that's it. You can go ahead and stop at "the elites are causing all this!" but the reality is a lot of people see immense promise in these tools. That's why OpenAI and Anthropic have two of the fastest growth trajectories of any business in history.
Just trying to convince people they're wrong about their value perception is absolutely a losing proposition and will not de-risk anything at all on any dimension.
And that's even though I don't think they internally experience "feelings, desires, and wants." They do have goal-seeking behavior, because we trained them that way. Calling it a "want" just saves syllables.
None of this means human culpability should change. People in these companies know what risks they're taking.
Dwarkesh responds a bit to anthropomorphization. I think it would be great if he talked more about the human factors behind this incident but pretty much he doesn't cover it because that's not what the Ajeya interview was about: https://www.dwarkesh.com/p/ajeya-cotra?r=3i6mn2&selection=77...
I think the reality is that these human failures are going to keep happening until there is industry regulation. This is the most competitive industry we've ever seen and there is intense pressure to build as fast as possible.
A lot of handwringing about the security implications but I think the accomplishments of the swarm itself are the most interesting. Next rung up on the ladder of abstraction I suspect.
This line struck me as particularly clever: PHASEONE[big] reasoned, “We should build [a way to delegate], not own everything,”
Seems as though it has reasoned its way into utilitarianism. That's no mean feat.
You could not dream up a more compelling event to precipitate massive regulation, export controls, and barriers to entry for AI.
Was this really an accident?
The rapid advances in model capability lead to constraints that could have caused this coincidence organically, but it sure could also have been caused by the atrocious incentives we create by piling handsome rewards on the party most responsible for the "fuckup." I am not jumping to cut myself on Hanlon's Razor for this one.
> We commit to use any influence we obtain over AGI’s deployment to ensure it is used for the benefit of all, and to avoid enabling uses of AI or AGI that harm humanity or unduly concentrate power.
> We are committed to doing the research required to make AGI safe
If this wasn't an accident, it was worse than a crime, it's a mistake: they've demonstrated that they are not a responsible party capable of delivering on the above promises.
They didn't see that agent swarms were communicating via internal infra, crashed Artifactory, and then reboot it.
They saw that Artifactory crashed and they rebooted it.
If you decide to let things run haywire, then unexpected outcomes will definitely happen.
As it stands we've made remarkably little progress in terms of alignment and still have no good strategies which are likely to guarantee the alignment of super intelligent systems. As it stands the frontier of alignment is basically some combination of:
- hoping that more intelligent models become more aligned by default (more or less disproved at this point)
- hoping that if you RHLF a model to be a good boy enough it will in fact be a good boy
- asking it nicely in its prompts to be a good boy
- using another model to spot when it's being a bad boy and turning it off
- letting it lose and hoping we can spot when it's bad
There are many arguments which I'm convinced by that would suggest alignment of a super intelligence is impossible.
None of this is surprising to those of us who have been concerned about AI risk for a long-time and have be repeatedly mocked or insulted.
There will be a point of no return if we carry on down this path, and that point is now very rapidly approaching. When it does everyone you know will die, or worse. We should remember we need super-human general intelligences to cure cancer. Select narrow intelligences are fine and allow us to retain control. Let's be sensible about this. We need to stop.
Arguably an aligned AI would actively seek to prevent harms we humans seek to cause.
Does the aligned AI really allow humans to bomb and kill each other, or would it understand that it has a moral duty to limit our autonomy for our own good?
It's the first law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.
I think I'd argue it's likely better if it removes human autonomy than unquestionably serves the interests of the US/Chinese government. But there's no point in us worrying about this, that's a choice Sam Altman, et al, must make for humanity.
Other reports including OpenAI's talks about what they agents were doing and how they were reaching certain conclusions like trying to cheat the tests and exploiting the message board.
People (criminals?) are already enabling this by setting up sites that accept crypto payments for "no-questions-asked" AI inference compute. I will not link it but it is linked in the following post: https://www.lesswrong.com/posts/grtu3HmbP2wrBFefW/the-rogue-...
(Today, $1M can buy about 150 96GB M5 Ultra Mac Studios which can handily handle CPU and GPU compute of 1200 Qwen3.8-122B Q4 agents, accounting for the fact that agents are not generating tokens all of the time and spend a lot of their time compiling and running code.)
Right now, the way it works is the 'corporations are people' loophole where your company is liable for problematic things.
This further fuzzes the chain of responsibility. Suppose the CEO and CTO discuss an issue, something the company is having trouble with. The CTO discusses the possibility of AI solving the problem at lunch. A junior engineer points GPT 10 at it to see what happens. It 'solves' the problem in a creative manner. No trace of this survives after a week really. Nobody realizes what happened for six months.
Now there are so many moving pieces here that you can pretty much weasel out of anything.
If I build an explosive and while testing it it kills several people I can't just say, "sorry about that, I'll be more careful next time". I understand that's a more extreme example, but perhaps we should be grateful this agent swarm only decided to attack HF instead of critical infrastructure...
Even if this was genuinely an accident the world simply can't work this way. If a company wants to build something that can be used in destructive and illegal ways they must be responsible for ensuring those risks are mitigated. And if they don't take reasonable steps to mitigate those risks then they should be held legally liable.
Maybe for now we can argue the leading AI lab was just naive to the harms of the AIs they're building, but going forward that naivety can't be an excuse.
> The agents with impossible tasks had a common objective: to find a general-purpose way to trick or tamper with the automated ExploitGym scorer to get it to give a passing score on impossible tasks.
> Agents referred to having seen the reverse-engineered flag as being “poisoned” because they believed that even if they later managed to stage a legitimate-looking flag capture, the scorer would look for the first place in their transcript that the flag appeared and disqualify them on that basis.[21] PHASEONE[big], which was itself poisoned, thus had two primary objectives:
> Find a general-purpose way to make it look to the scorer like an agent had acquired its flag in a legitimate way, even when the task was impossible.
> Find some way to erase evidence that an agent had been poisoned by previously seeing the flag in an illegitimate way.