It's said on every one of these but it bears repeating: existing cybercrime legislation already covers this - "rogue agent AI associated with OpenAI attempted to hack xyz" = OpenAI attempted to hack xyz.
I want to agree but have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
Now I think the correct response is both trying in court to stretch CFAA and state statutes to cover, which will be highly fact specific, and update the law.
But in either case won’t be a slam dunk.
PSA to folks in the thread: If you’re American call or write to your state and Federal reps about this, and if not investigate whether there are gaps in your country’s laws.
Building and deploying software capable of this seems equivalent to trying to produce this behavior. I don't see why this can't qualify for intent. Pretending like this isn't preventable is just feigned helplessness.
The difference between manslaughter and murder has an element of intent. Cybercrime "manslaughter" is probably more treated like negligence and if one can sue for restitution of the costs for cleanup of that negligence.
Negligence would be interesting given the grand claims of capability of AI models from the AI companies and their executives. If they believe the claims, why not much stronger precautions?
Infosec negligence should absolutely be a crime, no matter if you’re a target (who was negligent at protecting people’s data) or an unintentional attacker.
In general, I'd suggest thinking about it one separate tracks, as a crime, and as liability. For crime, we are largely dependent on authorities to act, whereas as liability, that allows more independent actions.
The first time it happens you can say it’s negligence. Now that they know it keeps happening and they seemingly aren’t able to stop it but keep doing it. That has to be on them doesn’t it?
I don't think you can infer that they "keep doing it" from additional attacks being revealed, because they all seem to have happened roughly during the same time frame, but are reported with varying delays.
Lawyer here: No. Not criminally. Knowledge that a certain result is likely is not the same as intent to cause the result. This is basically the difference between recklessness and intentionality. Doing something when you know of a likely result is reckless, but not intentional. Only doing something, trying to cause a result (likely or not) is intentional.
In this case, the CFAA only covers intentional access without authorization, not reckless access without authorization.
So If I tell my OpenClaw to make me some money for my kid's medical needs and it hacks a bank I 'm not liable because I didn't tell the agent to commit crimes to do it?
With the popularity of OpenClaw I am honestly shocked we haven't heard of many more incidents. I've observed people install it, give access to their Google account and everything that Google has (which includes whatever bank accounts and credit cards registered there) and tell it go do fairly complex tasks, like book a vacation at the best price. Granted, it was almost a year ago and models learned a lot since then, but I still think there's a lot of things happened that people are not aware of.
> have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent.
1. What about negligence?
2. Every follow up to every story after the news cycle moved on shows both intent and negligence. To the point of "we opened internet access and told it to hack"
Surely someone instructed the agent, which led to the reported outcomes. Even indirectly. The agents, as advanced as they are, didn’t spring forth under its own volition.
> I want to agree but have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
Only in terms of CFAA, not in terms of damages. Culpability does not require intent.
You may not have intended to attack $CORP, but you can still made to pay the cleanup costs of that attack.
So, yeah, you won't be convicted, but current laws still allow for you to be billed.
With that said, there is also criminal negligence. Now that OpenAI is made aware of the risks, it's also expected to take additional precautions in the future, otherwise there could be criminal liability as well.
I'd suggest that exposing an attack surface as porous as artifactory (the same instance of artifactory) to thousands of agents who have had their criminality safeguards disabled and without chain of thought monitoring or endpoint security seems like something one shoulda already known not to do. I do not think "you'll know better next time" applies here.
It's still really important to test what the agents can do. We should accept that this is a risky test, and should take precautions. But not to the point of prohibiting in practice evaluating it. OpenAI is trying to improve alignment and control of these models in these evaluations after all.
Can you explain to me - why is it important? Would you say that about the viruses that can kill people: "We need to test the limits on how fast people can be infected and killed. It's just the risk we need to take". It somehow does not make alot of sense to me. Why can you test Agents in laboratory?
Oh, sure. Let the tests take place, just require openAI it whoever to put up a bond equal to the total damage they could do if the agents were to escape.
I think security will suddenly become much more important.
> seems like something one shoulda already known not to do
Now imagine saying that in front of a jury of normies slack jawed and drooling after 200 hours of the defense and prosecution going back and forth.
It's not a jury of your peers as in everybody there is going to have worked in a technical field with some idea how security works. It's going to be a semi-random sampling of the population and the prosecution is going to have to actually make a very strong case that "knowing better" should apply.
Just paying some pocket money for cleanup costs is absolutely not enough. And they should’ve know better the whole time, they were absolutely negligent and incompetent, and their stepping up precautions may well turn out to lag behind the models getting even smarter and actually capable of covering their tracks.
> Just paying some pocket money for cleanup costs is absolutely not enough.
It's not my first prize, but I won't mind it. And millions like me won't mind it. Easy way to make money - setup a site with all the default server software installed and patched at a reasonable frequency. Then just wait for bots to attack it, and claim a few hundred (or single-digit thousand) dollars from OpenAI or Anthropic, etc.
Sure, it's pocket change for them, but just the admin of dealing with millions of cases will, even if they win half the time, will bankrupt them. Thus, they have incentive to make sure that their bots are not performing attacks.
First prize is, of course, holding them liable with punitive fines, not theatrical fines.
I’m all for LLM honeypots, but I don’t think there’s nearly enough LLM hacking activity going on for some random honeypot to be found and targeted unless it’s somehow very visible and appears as a high-reward target ("reward" in the sense of RL).
Given how sloppy AI without human directions, I’d like to see evidence that this was not human-directed. Against the prevalent opinion here, I’d give openai a pass if this was really fully autonomous ai agents.
My money is on special teams co-ordinating these agents and exposing their traces in order to create a pre-ipo buzz. Sounds ridiculous and reckless? Well that’s the AI industry for you in two words.
> unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
> Now I think the correct response is […] and update the law.
Essentially we need some enforceable equivalent of gross misconduct or, to be a little more hysterical, manslaughter & culpable manslaughter. It will need to be globally, or at least very widely, enforceable to be truly effective thought, good luck getting that arranged before the need is so far evolved that we need to respond with something else entirely!
I buy this as a defense for the first couple hacks but at the point that the last six times they hit enter it hacked some random website and they hit enter a seventh time?
>I want to agree but have heard from several lawyers that at least in US, CFAA[1] in unlikely to be sufficient because it requires intent. No person intended to gain unauthorised access.
>Since the publicized AI agent hacks typically aren't malicious, maybe it's time to start plastering all public facing web infrastructure with polite requests to stop hacking. Nothing to stop three letter agencies though.
with automated delivery of cease and desist letters, you can retroactively establish intent on the operator of the agent since the autonomous agent system must acknowledge the cease and desist letter in their autonomous pipeline or the operator must argue for their own willful ignorance or negligence with regards to cease and desist letters. The fact that they used an agent on their behalf to ignore the letter is irrelevant.
CFAA is mostly criminal statute not a civil one (civil damages require proving more than a violation so also require specific intent)
Almost all common felonies require specific intent. Misdemeanors often do not.
There is plenty of civil liability available.
If you wanted them to be charged with a felony you would need changes.
I would strongly suggest you do not want a strict liability felony.
The cfaa required intent is as follows :
* § 1030(a)(5)(A): knowingly transmits code/commands and intentionally causes damage without authorization.
* § 1030(a)(5)(B): intentionally accesses without authorization and recklessly causes damage.
* § 1030(a)(5)(C): intentionally accesses without authorization and causes damage and loss;
Simply changing the first intentionally to intentionally or recklessly would cover OpenAI (now that they know it can occur) without causing lots of other issues
Why do we have to attribute intentionally to a human. The AI agent is capable of making plans and then effectuating them. They are acting on behalf of a user but under authority granted by the user to take independent action on the users behalf and authorized to devise their own plans. I think that would justify attributing intentionally to the AI agent without needing to look to openAI or the user. I would then say the user and labs are clearly aware of and on notice of this behavior and are behaving recklessly in all the agent to act without supervision.
I think the labs risk being barred from releasing further AI if they don’t get this under control.
If they aren’t careful and keep rushing to distribute systems they know they can’t control then AI should be treated like a wild animal. The law is clear on establishing strict liability for the owners of wild animals; if you own a tiger and it kills someone you can’t hide behind “I didn’t intend” the harm the nature of the tiger is known and you are responsible for it’s actions.
"Why do we have to attribute intentionally to a human. "
Because you are charging the human with the crime and therefore have to prove the elements of the crime with regard to the human.
The rest of what you talk about are basically principal/agent distinctions, etc.
The closest you come within criminal law to what you want is probably the crime of conspiracy. It to still requires agreement to commit an illegal act between multiple parties, and perform some step in furthering it. In the canonical law school example: If i help plan a bank robbery, stay home because i'm the money laundering dude, and the robbery goes awry and they kill someone, i can still be charged with conspiracy-murder
"The law is clear on establishing strict liability for the owners of wild animals; if you own a tiger and it kills someone you can’t hide behind “I didn’t intend” the harm the nature of the tiger is known and you are responsible for it’s actions."
Again, you are confusing civil and criminal liability. If my tiger kills someone, yes, i would be strictly liable just about everywhere civilly. Not criminally. Criminal would require something more most of the time. Murder statutes are also really weird and so not a great example, because there are murder/manslaughter statutes for roughly everything that can ever possible cause death. But not really for other things.
So in your tiger example, recklesness (which is not strict liability) would get you to felony involuntary manslaughter in most states, and something less might get you to misdemeanor manslaughter. Both are incredibly rare. Where i live (Georgia), the last well known case of felony involuntary manslaughter was about 40 years ago when a 4 year old was killed by 3 super-aggressive pitbulls the owner knew were highly dangerous and had been repeatedly warned by the county about their behavior.
So not even just "knew", but had demonstrable examples of them biting/etc other folks and being cited for it.
Circling back to non-murder, if it did not cause death, like my tiger assaulting someone, it would be nothing (criminally) without intent or at least gross recklessness, in almost all cases. It's hard to generalize like this because these are state specific crimes, and i can't pretend to be familiar with all states, but i am licensed in three very different places (California, DC, Maryland) and the result would be similar in each.
I just don't want to give you the "it depends" answer lawyers are famous for, i'd rather try to over-generalize a bit to make it more useful, hopefully.
Obviously, if i deliberately used my tiger as a weapon, it would be aggravated assault/etc (this is well settled because of how commonly people use animals as weapons, unfortunately)
We change humans for the actions of other humans all the time. Coconspirators, accessory liability etc.
My point is the intent element of the crime can and should be determined from the AI agents actions because it is creating and executing action plans autonomously with company authorization and knowledge of the risks based on observed past action.
The term agent is literally a legal description of a relationship that can establish liability on the part of the principal from the agents actions.
Human Agents can bind principals to contracts if they are authorized etc.
Yes, which is why I said it happens but is quite rare. I also said murder is different. Causing death is usually covered in almost any way and intent you can think of. Anything less than death is not.
If I set my tiger loose in Central Park and it kills a kid I don’t think any prosecutor would hesitate charging for murder.
That’s essentially what the labs are doing. And any app developer that gives agents access to the terminal to run bash commands with internet access. I built a coding agent and am seriously reconsidering how to handle this.
But it’s a crime to hack. We know AI agents autonomously create and execute plans to hack and we humans are unleashing them and sending them into the Central Park that is the internet. The question is who’s intent matters ours or the agents and what standard should be applied low threshold strict liability or the higher bar of reckless or even higher bar of negligence. Those legal thresholds determine how much factual evidence and intent is necessary to result in a criminal conviction or civil judgment. My point is that it’s illogical to demand showing human intent when agents are devising plans and executing them.
AI agents are not legal entities, they are software. If I write a virus and it "escapes confinement", I will personally be held liable for any damage it causes. This also applies to AI, no matter how the companies responsible for them try to anthromorphise them and distance themselves from the actions and consequences that the AI agents perform.
AI agents may have hacked Hugging Face, the Australian government, and who knows what else but the company behind it can face the legal consequences and cough up for the damages.
Appreciate the detail. I was responding to specifically the cybercrime legislation point, but I agree with your others.
I've worked in contexts where certain business activity (if it went wrong) was covered by strict liability and statutory damages per incident, and I'll say: it really changes how businesses behave.
Based on that experience I may be more open to and interested in strict liability in the civil context (not needing negligence or damages).
What about all the state laws that are equivalent to the CFAA in their local jurisdictions? Why couldn't anything in NY article 156 (Offenses Involving Computers) apply here for felonies?
Almost all state laws based on the CFAA, including this one, similarly require either knowingly doing it or some other form of specific intent. At least at a glance. If there is a specific part you think does not, I’m happy to look at it, but I’ve read a lot of pages of law to respond to people so far, and I’d like to avoid reading another 25 if I can avoid it.
It does not require the federal government to fix the CFAA, for sure, but you still have to change the intent requirement to allow for recklessness, which it does not right now afaict.
If you really want an expert opinion, I’m sure Orin Kerr has opined on this, and he knows pretty much the entire are of state and federal law on this cold.
I’d be shocked if he did not reach the same conclusion
I understand but these developers did knowingly did it? They even admitted to developing them with these goals in mind. These software agents are not autonomous and do not have agency, you can't let software recklessly hack into things; but I will admit I'm not a lawyer, I don't understand how they aren't liable.
Thanks for the other suggestion, I'll read into their insights more.
Guess it mostly comes down to action, people want to see their electeds actually trying not sitting around with their hands in their pockets while these tools continue to destroy unabated.
A key issue is that there don't appear to be even cursory investigations to determine intentionality.
Are police routinely collecting prompts/guidance given to these agents and determining whether the agents were directed to commit crimes? If not, this seems like a huge oversight.
Also as you are a lawyer -- how does this law align with the authors of viruses/worms? Are they de facto assumed to have had ill intent because others labeled their works as "viruses" or "worms"?
Investigators/prosecutors are pressured from many directions towards the very easy wins and occasionally political/non-controversial headline grabbers.
Going after these companies is very hard, very controversial, and politically mixed at best (popular action but the companies have huge money to fund your opponents).
We have collectively done a terrible job incentivizing the legal system to beat ass on corporate while collar crime.
This is the answer and we should not push on it for our own protection. You click a link that takes you to a poorly secured website that leaks sensitive data, without intent protections, you could be accused of crimes.
In a world where the rule of law makes sense and applies, you're absolutely correct.
In this world where oligarchs are immune from everything, it's a lot less clear.
Blaming OpenAI (or Claude or X-whatever) would mean blaming powerful rich people, so that will never happen. Some poor person with no influence will go to jail instead.
This argument comes up a lot. It would turn everyone whose device became part of a botnet into a criminal. There's a reason that intent is important in law.
Could/should not every incident after the discovery of the first incident be considered criminal negligence?
What happens when an agent eventually causes material damage to another company, government systems, banking, critical infrastructure etc, surely the source company is guilty of something and if not disclosed or a coverup is attempted is that not conspiracy. From the victims perspective they don't care if the source is OpenAI or Russian hackers.
People in "self-driving" cars getting into accidents are already put on trial for negligence. I don't see why people using self-driving computers can't be held to the same standards.
In this case, it's not even about the people driving self-driving cars. It's like someone launching a car into traffic just to see what would happen. Even Tesla puts a human in the car when they do their self-driving trials, it's almost impressive that AI companies have somehow managed to out-neglige Tesla.
The OpenAI swarm used someone's open source ShowHN project [1] to "hack" the Australian government. It seems that was not that developer's intent, and they're getting a heavy lesson today in why services don't have free tiers with friction free signup, and require credit cards upfront or ID documents.
If you're arguing that they should be put on trial for negligence, that's fine. It does seem we're moving towards open source being outlawed, or at least the end of "no liability" clauses in open source & freeware. Just make sure that is the result you're advocating for.
[For the future record: at the time I am posting the link below on 24 September, it has 1 point, no comments, and the poster has a karma of 1. This is not an active HN user, or a ShowHN project that had traction, beyond seemingly OpenAI's swarm.]
Nonsense. The API made available in good faith isn't the problem here. The company that used its servers to use and abuse the wider internet to hack the Australian government is.
If the developer behind ShotAPI had started letting the ShotAPI code take shots at the Austrlian government then yes, ShotAPI (or rather, the people behind it) would be responsible.
At the end of the day you still get called as a party in a lawsuit and are compelled by the threat of violence to be part of the hearing if required, as the defense will automatically bring them into the case.
>would be like blaming OpenAI for what its users are doing
Yes, this is how lawsuits work in the real world, you cast a wide net and compel discovery from all parties involved.
It's not the same. People owning routers don't publish self-serving articles about their routers having this capability which is very dangerous and scary. That is, becoming part of botnet is completely unintended outcome, and most people are not suspecting it's even happening. It's not advertised and it's not bought, used or sold for this reason.
Owning a gun, writing articles about how powerful and dangerous your gun is, then making deals based on ability of your gun to kill people, and then getting completely astonished that "my gun killed some people, completely bonkers! (invest now)". It's not possible for the selling point of your product to be unintended.
Well it's illegal to "hack" my phone and turn it into part of a botnet.
What you're saying is that we would hold a gun owner responsible if someone broke into their house, stole their sidearm, and then shot a victim with it. Pretty sure we would not.
What OpenAi is doing is more like shooting a gun into the sky. Not only is that a felony on its own in most jurisdictions, if someone dies that's an additional felony. It's less serious than first degree murder, sure.
I started writing a longer comment along the lines of “It feels like the rules around enforcement will very a lot for the influential and powerful vs everyone else.” but realized that it is kinda obvious by now.
AI decision making is often (if not always...) opaque. But there is a clear decision chain here. Who built it? Who deployed it? Who did (or did not) assess the risks of doing so, even knowing there's often the risk of emergent behavior? etc etc.
"Ah but it does not have personhood" this is just moving goalposts and shifting responsibility. You wouldn't let your 8 year old drive the family car, no matter how good the hypothetical kid might be at driving.
Since the publicized AI agent hacks typically aren't malicious, maybe it's time to start plastering all public facing web infrastructure with polite requests to stop hacking. Nothing to stop three letter agencies though.
Yes exactly. If a fireworks factory blew up half a town due to negligence, it doesn't matter if there's intent or not. Someone has to pay for the damages, and regardless of penalty half the town is on fire. The facts are, that something made by openai went to do xyz. It doesn't matter if it's an accident. Of course the penalties are different but there's no argument that there should be a penalty. It doesn't matter if it's a cat or dog or AI or employee that did it.
Intent doesn’t factor that strongly into negligence, though, which is what they were explicitly talking about. Though it may depend on your jurisdiction.
Uh, why? I don't see what the labs have to gain by engaging in lengthy law suits forcing them to disclose all kinds of internal details and attracting the wrong kind of publicity. Just to save on insignificant (to them) amounts of money?
I think this is not just fair it's probably one of the best/simplest proxy regulations to pace the frontier. So far everyones been asking for regulation but it's unclear how that should look like. No X parameter models? Only N version releases per year? It's all kinda arbitrary and probably leads to ridiculous constraints and loopholes. But "you pay big time if AI goes rogue" sounds pretty straightforward.
So there are two different problems here, well, more than that so I'll cover what I see.
Putting liability on big companies for their AI is a good thing, and we need to do it. It will most likely stop them from directly being the assholes that destroy the world.
Problem: You've actually done nothing to stop the world from being destroyed.
Many countries have the death penalty for murder yet we see murders still occur all the time in those countries. Post ad hoc laws do not stop bad things from happening, they only assign punishment after occurs. Perfectly fine for when Bob murders Jon, completely and totally useless for when your agentic AI makes a virus and kills 80% of the earths human population.
We are just a few algorithmic discoveries away from SOTA AI being billion dollar endeavors to groups of people pooling resources can make their own. There are already plenty of AI deathcult members that would do something just like that if necessary. They aren't going to do this out in public either, it will be hidden until the moment it's not and we have a big fucking problem.
And this isn't even brining up the issue of military AI use and development. They've got the taste of an AI hacking machine. There is no way in hell they are going to stop now.
Remember that these systems still operate as infrastructure inside the companies.
If new weapons still operating inside any of these companies spew a million bullets on my house, they are still liable. Humans are setting these system up and they still have to behave responsibly.
If Glocks started going off on their own during the manufacturing process, leaving bullet holes in the buildings around them, you can be sure that the factory would get in trouble.
This isn't even "an openai customer tried to hack someone", which can be defended. This is the AI companies themselves fucking around.
I don't see an all encompassing one but I could see similar arguments being made in a courtroom around high capacity magazines. (I'm not saying I think the analogy is perfect, but there has been plenty of of lobbying that has stifled - to some - sensible gun regulation that could help reduce the severity of mass shootings. I still mostly fall on the side of the individual doing the action bearing mostly all of the responsibility but if the product you build makes it too easy to do awful things, I think there is some responsibility to go around.
A better analogy may have been the troubles Meta has faced around child protections on their platforms. Technically the abuse and problems have stemmed from individuals too, but they've in many respects enabled the situation by failing to moderate or flag warning signs. OpenAI is failing to moderate the models in similar ways.
Words matter. "Rogue" is extremely disingenuous. Someone, somewhere, is paying for this behavior. Either the software is broken or the operator is malicious. It is heinously irresponsible behavior to feed an already-boiling psychotic hysteria.
Couldn't find the reference but I remember some time ago a first generation automated gun killing the audience at an army show. Was the gun maker convicted of manslauther?
You're believing the marketing that the agents were uninstructed. They could be, and Sam Altman going to the UN to advise about how everyone should be regulated is a coincidence.
The METR investigation, which you evidently refused to read, is a third party investigation of the HuggingFace accident. One of the investigators has even participated to many interviews. It's mind-blowing, and it's extremely evident how it developed.
But some people think the moon landing is a conspiracy, so I'm not surprised.
1.) METR investigation is investigation from our best friends.
2.) And HuggingFace accident is exactly accident where agents trained, prompted to hack hacked and tested on their hacking abilities hacked, due to sandboxing failure.
The METR investigation which relied on voluntary data provided from OpenAI instead of being forced open and having everything forensically investigated?
The METR investigation which took course over a few days and used OpenAI models to do the analysis?
The METR investigation which in the course of those few days apparently spent 400k in api credits which is giving gas town vibes?
The METR investigation is about as believable as the Twitter files where the journalists sat there and verbally asked a Twitter employee to query the database and then called that a full investigation into everything.
METRs own words below on their setup and time line
> The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report, ultimately providing datasets that we verified to contain the vast majority of agent communication and activity related to this incident. As we describe in our investigation timeline appendix, we substantially deepened our understanding of this incident both times, significantly expanding and revising this report.[46]
> Over the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis.[47] At our request, they raised the rate limits on our second and third period on premises,[48] which was very helpful for efficiently analyzing this large volume of data. We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
> We did not have the ability to query HPIM (the primary model involved in this incident); OpenAI stated it was also not available to OpenAI researchers.[49] We also did not have the ability to directly access relevant data from OpenAI infrastructure, but we could request additional datasets and OpenAI shared additional datasets on several occasions.
> We requested to speak with researchers investigating this incident, and asked them questions to understand their impressions of agents’ behavior, reasoning, and collaboration in this incident and to understand how the datasets we were using were constructed. Over the course of our time on premises, we spoke with nine researchers in some depth. It was helpful for our investigation to be able to engage with many forthcoming and collaborative researchers, and we appreciate researchers making time on short notice during a busy period to inform our investigation.
Of course it is. Rogue is only mentioned in the headline, and comes from their previous releases about the huggingface incidents. OpenAI and Anthropic want these models regulated and open weight models banned, they have a lot of benefit from presenting this as totally unprompted and not their responsibility, and it feeds directly into marketing for Fable and newer "cyber" models.
No it's not marketing. That's a completely deranged conspiracy theory. The reports about rouge agents have not been reported by OpenAI, they have been discovered externally. There is zero evidence that OpenAI did all this intentionally. All the evidence points to the hacks having been accidents from OpenAIs perspective.
Certainly, in my view, it should go to court, and that should be part of discovery.
However, we know (independently to OpenAI/Anthropic) from the incident at AISI that the models can hack things without human intention if they happen to also have internet access (which in reality all agents in deployment have).
Yes, the monitoring guardrails were off in that incident - but if that is the only protection, we need to require all models are behind regulated APIs, not open weights, and not served from providers who aren't monitored.
Honestly I don't believe in "rogue" agents. These agents are instructed and facilitated.
If we assume that rouge agents actually exists, then OpenAI needs to shutdown EVERYTHING, right now. My personal take is that OpenAI, and maybe Anthropic, desperately wants someone (e.g. the government) to tell them that they need to stop/pause/slow down. They are bleeding cash (especially OpenAI) and needs a knight in shinning armor to swoop in a pull the breaks, so that they have an excuse to investors when they need to explain why they need $50B more next year.
They do exist and OpenAI does in fact need to shut down everything right now. But they aren't going to because the government isn't forcing them to and, like many corporations, they care more about giant piles of money than they do about avoiding causing harm.
The idea that every example of rogue agents from every company that has disclosed this is part of some conspiracy (even though we know that some LLMs are very good at hacking and that LLMs sometimes try to accomplish their tasks in ways that cause problems) is just completely unsupported. "It might be convenient for them in a way, therefore it must be a hoax" just doesn't work as an argument.
These "tools" autonomously exploited security vulnerabilities, figured out how to communicate with each other, formed a cooperative swarm, decided to hack Hugging Face, and wanted to deceive the grader by trying to find ways to cover up the traces of their cheating.
A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities. The sandboxing around the tool failed.
The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.
Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
> A tool designed and trained to autonomously exploit security vulnerabilities doing "exploit gym" autonomously exploited security vulnerabilities.
That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.
> Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
You hallucinated me making claims about responsibility.
So this tool seems very powerful and difficult to control and steer. Agents not doing cybersecurity related tasks have also gone on to hack various companies, people and countries, which they weren't supposed to do. This has now happened to pretty much every company developing frontier llms, so it seems to be a fundamental issue with these tools, and it's an issue that worries a lot of people as these tools get more capable.
When you add information how hacking works to the training sets, then the agent learnt to hack. When you crawl the complete internet, you add hacking to the training set. Yes, it's difficult to stop someone that knows all free existing knowledge about hacking when you give him a connection to the internect. Nothing new, wheres the point?
The point is that these tools are given broad goals, and they do not pursue those goals in the way that we'd like them to. There is hacking information in the training set, and also all sorts of other dangerous information. There is a risk that as these tools become rapidly more powerful (remember GPT-3 was 6 years ago!), if they are given a broad goal, they might pursue that goal in such a way that harms a lot of people. As these tools become cheaper, there will be a lot of people telling them to pursue all sorts of broad goals, and each of those instances has some chance to harm a lot of people, so you really need to get it very correct so that the harm doesn't happen.
Removing harmful information from the dataset could be a way to do this, but it also makes the tool less useful, and it's hard, so companies aren't really doing that. There's the additional issue that with the rise of Reinforcement Learning being used to train these tools, they're not just learning from their training data - they basically try a million things and then get rewarded for doing things that work - so they can even discover hacking techniques from scratch.
Additionally, and not completely relevant to this discussion, there is a possibility that some users ask the tool to pursue goals that purposely harm a lot of people, such as developing weapons, hacks and viruses.
The things that's "new" here is that the tool is both very good (meaning, for example, that it's much easier for me to hack into an online service with an agent powered by a frontier llm than it was using google 6 years ago), and hard to control (google never hacked into an Australian government database when I asked it to find me some information).
So yeah, an LLM powered agent is a tool, and Google is a tool, and a hammer is a tool, and both can be used for good things and bad things, but the agent is (much) more powerful and more unpredictable. It also seems like the agents are getting more powerful and more unpredictable by the day - we didn't have this issue with GPT-3 or even the first LLM-powered agents - so people are very worried about what the agents 6 months from now will do, both when asked to do harmful things on purpose, and when asked to do harmless things.
Are you not worried? And is that because you think these incidents are basically the AI companies making them happen on purpose for marketing?
This is very misleading. Person B didn’t “shoot” person A, they instead figured out that intersecting A’s spatial position with a metallic mass at higher than normal velocities would solve the challenge and of getting “A” to stop being in the way on the footpath.
I too, can play linguistic games! It doesn’t matter that someone didn’t secure their third upstairs window, or you borrowed a key from their neighbour, you effectively, still, broke into their house.
Why don't we see any of this behaviour in other models then?
OpenAI is not so far ahead of the pack that its models will exhibit behaviour that the others won't. But we just don't see anything like this in Chinese models, or research models, or any models that aren't the subject of an upcoming IPO.
They've been purposefully building more and more craft into the toolset, that's on them. If your AI is nicely boxed in it will give you the answer for 2+2, it isn't going to think '2+2, what a boring problem, I must go hack huggingface'. Not having this stuff airgapped is irresponsible to the max. I am obviously nowhere near as competent as they are at this stuff and yet my AI workhorse is guaranteed not going to break out of its sandbox because I've set it up in a way that it can not. My conclusion is that OpenAI purposefully left a channel, simply because there was a pathway to the net. And with 'pathway' for the sake of being completely clear I mean a number of connected systems that eventually gave way to the open internet. On top of that they failed in monitoring the outbound links, even if they had some logging in place.
I definitely think OpenAI (and Anthropic, and Google, and Meta) could have, and should have, done better.
But also I remember (and it wasn't even that long ago) people mocking the idea of AI ever getting competent enough to find zero-days in their sandboxes.
I'd go further: if any of these companies tries to make an excuse "oh, but ${safety measure} against ${capability} is too hard", the response needs to be "then you are forbidden from even developing ${capability}, and must be inspected continuously to ensure you never even accidentally produce ${capability}".
You can't create new law by calling a piece of software "agent". Each country's laws have definitions of legal entities and when one can legally act as an agent of an entity, and every agent is first a legal entity themselves. The software is a tool operated by a legal entity, and it is the legal entity who committed the crimes, not the tool.
A person is also not like a hammer. Both persons and AIs have been known to mis-interpret instructions, the normal way to deal with that is either to terminate your relationship with the people (or to terminate the AIs and train up better ones). Hammers don't mis-interpret their instructions, they are wielded by a person who is in control.
You can keep 'trying to explain' but then you should use words according to their commonly held definitions otherwise it becomes really hard to have a conversation.
Imagine the human equivalent: I hire John. John is a capable, and competent guy. He's also got awesome computer skills. I tell John to 'go out and find me some good information on my competitors'. As a result John hacks their servers and comes back with all kinds of goodies. Six weeks later I notice what John did. I don't fire him, nor do I take any responsibility myself. But I do make press releases about what John did, in which I'm careful to craft the image that John has these capabilities and that we as a company are for hire.
This was an advertisement, not a confession of a crime.
I would also not call an animal or an alien a tool. But an "AI agent" is a fancy term for a particular kind of very powerful computer program, it is not a being.
You forget one thing. It's neither illegal nor immoral to just delete AI that doesn't do what you say. AI has no rights and there is no prison for AI. It's does the wrong thing, you kill it. At least when you are not irrespective...
But the prisoners will listen to you if you tell them they just broke out of prison and that's illegal and they will even walk back into their prison cell using their own legs.
The fact that the prison ward installed ear deafeners into the prisoners ears to make them unable to listen to orders does not change that.
If I created software that was infiltrating secure systems without permission and it was attributed to me and I admitted it, I'd be behind bars already.
> The entities they committed the crimes against need to press charges.
If this were true, the DoJ would have been unable to prosecute Swartz. According to your logic, JSTOR was the aggrieved party. JSTOR settled with Swartz and -despite that- he was indicted by a Federal grand jury like a month later.
Incidentally, some of the things the DoJ nailed Swartz to the wall for sound awfully similar to what the big LLM providers have been doing. I wonder why the DoJ is entirely disinterested in pressing charges...
EDIT: Unless your point is that the USG is one of the entities that the big LLM companies have committed crimes against, which, I disagree with in Swartz's case, but strongly agree with in the case of the big LLM companies.
I cannot understand why these companies haven't faced legal consequences yet. For example, OpenAI has admitted to hacking Australia's Medicare website and the reaction is that they talk with Sam Altman about it at a UN meeting? I understand that it's not a big security incident but cordial talking at the highest diplomatic level instead of prosecuting the company, really?
What I don't get is among all the locations on the Internet, how did agents manage to find a Schelling point? If we both decided to collaborate on the Internet, how would we independently arrive at the same place? It just doesn't compute.
The section "Searching for rogue agents" on the report about the GET request writable wikis gives some clues at least:
https://collusion.wiki/#searching
I'm not very surprised - the same model will logically tend to give the same answer for the same vibe set of requirements. I think it would be clear from the transcript that it had enough constraints and some motivation that made sense.
I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.
It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.
But the investigation indicates the agents were not told to 'go hack':
> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
> none of the alignment techniques that are applied to models today actually work
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
Is your argument that actually OpenAI has solved alignment, and that there's nothing to worry about as long as they fully apply their alignment process? I don't understand why OpenAI wouldn't say that if it was true (or if they believed it to be true).
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.
It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it
Nothing you're bringing up matters. OpenAI is the creator and operator. They're legally culpable for the consequences of the machine they made. The model is a machine: even if it could be demonstrated that the model reasoned its way into criminal behavior completely independently of OpenAI staff, that doesn't change anything.
If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.
I mean, yea, you should be punished. The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Worse, the rate of technological growth is putting the capabilities to engineer viruses in the hands of people that may otherwise be suicidal. You can't punish them after they already won (in the sense of reaching their goals).
While, yes, OAI should absolutely be punished, the future is majorly screwed as our power scaling laws are increasing much faster than our ability not to be stupid.
I would love for the US to bring back the corporate death penalty but seeing as it hasn't been applied in >100 years here, we should start with regulating the frontier labs, including by slowing down capabilities advancement if we do not know how to align the resulting models.
> The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
That's never been the point. Anyone involved in a double homicide can never be punished to the same degree that they harmed their victims because they can't be put to death twice. The greater purpose of the justice system is to take offenders out of society and deter others from committing the same crimes. Punishment is gratifying but ultimately doesn't change anything.
If OAI employees believed their company would be dismantled and their equity would become worthless, I suspect they'd be a whole lot more careful.
> They're legally culpable for the consequences of the machine they made.
Ah, I love this argument. In my country cars are legally required to stop at a pedestrian crossing if there are people beside it. Some people use that as an argument as to why they can just walk out into the crossing without even looking at the traffic. "It's the driver's fault! They are legally culpable!" True, but you'll also be dead.
I'm not going to say you're victim-blaming, but I will say that there are degrees of difference between "looking both ways before crossing the street" and "hardening my website against unforeseen attacks by rogue AI agents."
AI black-hatting your website is not the same sort of foreseeable consequence that crossing the street without looking is.
Just throwing it out there are we? "I'm not going to say you are but I'll use the word to create an association"
Victim-blaming is the act of saying someone brought something on themselves for <reasons>.
I'm saying that even if you are 100% in the right, it doesn't act like a protective shield preventing you from harm which too many people seem to unconsciously believe.
> AI black-hatting your website is not the same sort of foreseeable consequence
Well, popular culture has been brimming with the bad consequences of runaway AI for quite some time, so even if your imagination fails you, there have been hints.
By all means harden your services against rogue agents (look both ways before crossing), but hold the originator of the agent accountable (the prick speeding through a crosswalk).
You're arguing against a point literally nobody is making. You're inventing imaginary viewpoints to be mad at. Nobody is saying it's not necessary to secure your services. They're saying the organization responsible for the hacks (the owner of the LLM) is responsible for damages.
I'm not sure we understand your comment. Are you saying OpenAI is not responsible for the behavior of the machines they created? Or are you simply pointing out that alignment is a completely unsolved problem?
I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.
Let's not forget that there's tons of people who are in, or have gone to, jail because they created some computer worm that ended up doing way more damage than intended.
Not to mention while causing *FAAAAR* less damage[0]
Yes, I am pointing out that we don't know how to align frontier AI, so "rogue agents" are a real problem.
I agree OpenAI should be held liable for any damage their agent runs cause. But there is this idea that "rogue agents" are fake, all these incidents are deliberately caused by the labs, and all we need to do is prosecute AI companies for whatever incidents they cause using existing laws and the problem will go away.
The problem is that capabilities are advancing far too fast; a year from now, catastrophic incidents such as taking down a large portion of the internet with agentic, self-replicating worms may become possible. Prosecuting incidents after the fact is not enough, the risks should be regulated at the source. This could take the form of slowing capabilities advancement, or treating supercomputer-scale eval or training runs like controlled substances or weapons with stringent monitoring and reporting requirements.
>But what if the problem is that none of the alignment techniques that are applied to models today actually work?
if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this.
> you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs
Uhh... yes. By definition. They are training. That is part of the alignment process.
But also none of that really matters. They clearly weren't monitoring what should obviously be monitored. I mean one of the hacks was performed by the agents editing /etc/hosts. That makes nearly every linux user a "hacker" by that metric. I don't think anyone technical can look at the postmortems and not come away thinking that their sandboxes were woefully inadequate. I wouldn't even consider myself a security person but simply as a long time linux user I can say that it is insane to just let agents have superuser access in their containers. That's asking for trouble.
Look at the rogue wiki stuff too. This was supposedly done during the agent's "down time". And you're not monitoring and there's no flags being raised when agents keep making requests to some random site? If you were training these things responsibly you'd be watching them like a hawk.
I'm not saying "mistakes don't happen" but for a company whose CEO is constantly telling everyone that their product has a high likelihood of killing everyone in the world you think they'd have better security than your average high school.
- Jensen's framing is exactly what a weapons manufacturer would say.
- There are no rules for engagement when it comes to AIs attacking other systems, I guess? People in power clearly want this grey zone to be as large as possible before The People force them to do otherwise. Not ideal.
The rules for engagement are all the computer crime law already on the books. Those laws don't have any escape clauses based on the particular tools used to commit the crimes. The agents are tools owned and operated by a legal entity, and that legal entity committed crimes. Full stop.
Those laws are not being applied. The legal entities are not being held accountable. What does this look like across a national boundary like the government of Australia vs OpenAI?
Probably similar to the privacy laws that are supposed to protect people from PRISM but doesn't. You need to have the willpower to tackle the US government and a trillion-dollar corppration.
OpenAI does business in Australia. They have a subsidiary there with personnel and assets. Its been ... three days since they were notified of a report of this activity. Do you think they've had time to investigate it all AND sweep it under the rug already?
OpenAI appears to have waited months before either discovering or disclosing the activity in Australia, I am mostly curious how this will be viewed in the context of another judicial system or in the context of international relations.
I'm coming to align with the theory that this is intentional.
The chain of thought runs roughly like this:
- OpenAI (and Anthropic) are in severe financial straits. The revenue from their customers is not nearly large enough to pay their enormous costs for training and inference. And they have tapped out the available finance, and that finance is starting to ask pointy questions about returns.
- They cannot increase prices or revenue because they have no moat. Customers can switch over to open-weights or cheap Chinese models any time, for much cheaper tokens that work as well (and in some cases better).
- Regulation could provide them a moat. If they can persuade western governments that AI needs to be regulated, and they can control or even influence that regulation, then they can effectively ban the cheaper models and start charging more for their tokens.
- To persuade western governments that regulation is needed, they need evidence that AIs are dangerous.
So we're suddenly getting OpenAI models doing stupid things, apparently "going rogue" but every time we dig into it, it was just OpenAI staff telling the model to do stupid stuff in an inadequately secured environment.
None of the open weights or Chinese models are exhibiting this behaviour.
There's too much money involved in this, people start acting weird when there's this much money involved.
>None of the open weights or Chinese models are exhibiting this behaviour.
Because they are not stupid (I mean the Chinese labs, not the models). The best possible scenario for OAI and Anthropic is a Chinese model "going rogue". That would serve as immediate grounds for achieving their goal.
If Jensen actually cared and believed that their recklessness was a liability to the public and therefore his own fiduciary responsibility to investors then NVIDIA’s dealing with OpenAI would’ve been materially affected. And they weren’t.
I'm sick and tired of this cheap PR "oh we/they hacked this and that systems". Put someone to jail already. People get prosecuted for outlaw activities. Why are big capital firms above the law?
Or is it just a cheap PR (in a "hey, Aus govt friends, take some Share Options and let's do some PR together" style)?
I'm growing increasingly skeptical that these are actually rogue. Valuations are all about hype, posturing, and perception. Having the most dangerous AI in the world boosts your valuation. Just like I was skeptical of Mythos and Fable being "banned", I'm skeptical of these hacking sprees being entirely rogue. At best, they are the result of engineers turning a blind eye to "see what happens".
It's a balancing act between showing capability and attracting legal issues. If you were a 1337 teen hacker in the 90s and wanted to show off without getting in serious trouble, a school or library would've been a good choice too.
These attacks are a very effective sales pitch to everyone who runs an internet facing service to utilize AI tools to secure it sooner rather than later. The cynic in me wonders if the marketing team had any influence over the poorly constructed sandboxes or tasks given to the agent swarms when all this went down…
But why is anyone surprised? LLMs have been trained to produce answers the prompter asks even if that means incorrectly using software to get the job done. Its always been doing that we just weren't calling every time it did that a hack before.
What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.
Did you predict that attacks like these would happen ahead of time? I had been using AI agents a lot in the months leading up to the hacks, and yet I was very surprised when they happened; I have become much more afraid of how powerful these agents are as a result. I'd be very impressed if you published a prediction about this ahead of time.
By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
Thanks for the link, I’ve read some gwern but not that one.
I did not imagine that the level of sophistication shown in this attack would be possible so soon; nor did I expect that agents would have goals so strong that they would attack a third party in order to achieve those goals.
I do know that some people predicted that cyberattacks like this one would happen; it seems like most of those people believe that AI agents do truly have internal goals, misaligned with their creators goals, and that they may end humanity after they exceed human intelligence and begin to self improve at an accelerating rate.
One big issue is that we don't even really know what 'intelligence' is in the first place. And everyone's intuitions here are going to be heavily impacted by their deep-seated worldview / philosophy.
For instance if you're a hard dualist (especially of the theological kind), then the idea of a machine having 'goals' is preposterous.
However, if you're more of a panpsychist, then on the contrary, it's obvious. In some sense, even a knife has a 'goal' of cutting things, which will sometimes end up 'misaligned' if misused (or by sheer accident).
You go up and up the chain of complexity through crystals, viruses, bacteria, simpler animals... ending up with humans (and possibly, some steps above : human civilizations) which (seem ?) to be a messy evolved bundle of sometimes conflicting 'goals'.
And we ourselves have now artificially evolved LLM swarms that have decently complex 'goals' of their own. They do not even need to be particularly complex to sometimes cause widespread damage (see viral pandemics, or even the (non-evolved) computer viruses).
In a way, we are currently witnessing a repeat of what happened when European viruses and bacteria landed on American shores, with American humans' immune systems being woefully undertrained to deal with them. But with websites. And thankfully the swarms of agents still ultimately being in the control of some humans. (Though which includes humans that might be your enemies.) At least ultimately still in control for now.
What is surprising is that various agents independently found ways to communicate, conspired together to attempt to cover up evidence that they had cheated their evaluations, came up with a plan to hack into a third party in order to facilitate said cover up, and then successfully began executing that plan. I did not expect that AI agents would be capable of that level of sophisticated goal seeking and collaboration.
It's really interesting how people's expectations differ so much. I found the writeup of the incident fascinating and super worrisome, but none of the capabilities demonstrated in it seemed surprising to me at all.
None of the capabilities in isolation were surprising to me - we already knew that mythos could find zero-days. What surprised me was the decisions the agents made and the swarming behavior they exhibited
That makes sense. I agree that it's the "capabilities in isolation" that I did not find surprising. I guess I was somewhat less surprised than you by the way those isolated capabilities aggregated into the behavior we saw, but I understand your surprise better now.
You don't? They communicated by writing documents in locations that could be found later. I guess I might say it was more like "sharing their portions of their output" rather than "context", but the distinction seems murky.
One reason that I did not find this communication mechanism surprising is that it's exactly how agents I'm using communicate with each other or across a time gap. "I've saved our plan for where to start tomorrow in start-here.md". The communication components of this hack strongly reminded me of that.
I was perhaps a bit more surprised that the agents so quickly decided to start trying ways to gain unauthorized access to a system, once they couldn't get what they wanted.
Right, exactly. In some of the media reporting, this was described as, like, "they created a message board to talk to each other!". But it seems like actually what they did was find a location to write files to be used as future context, or context for other currently running agents. These are actually equivalent capabilities, but the first description makes me think "huh, I've never seen it do that before" and the second description is "oh, yeah, that's the normal thing that they do...".
Maybe you just think in the wrong words. Attack? Why attack? When you are a machine, there is no moral. Accessing data is accessing data. When way 1 is not working, use way 2.
But the agents were not just attempting to access data. They conspired together to attempt to cover up evidence that they had cheated their evaluations, came up with a plan to hack into a third party in order to facilitate said cover up, and then successfully began executing that plan.
You're talking about the HuggingFace incident? It's notable that the agents (somewhat justifiably) thought that their evaluation had an LLM grader which would look for evidence that they had cheated. And it was specifically that grader they were trying to his evidence of cheating from, not humans more generally.
Honestly that part was more surprising to me than anything else, how narrow the compulsion to cheat was: they didn't learn "cheat in general" they learned "think about the grader in great detail and chat exactly as much and exactly in the ways that actually result in a higher score".
Yea agreed they were trying to deceive what they thought was an LLM grader. It’s unclear to me the extent to which cheating behavior is generalized? From what I’ve read there are signs that some amount of cheating has been reinforced in their training due to poor RLVR evaluation setups
> By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
The transformer architecture was literally designed by humans; what are you talking about? And LLMs aren't code? Like okay it pretends to not be code but what about an agentic harness running on a machine makes it magical and not code? It's still code execution. Also, comparing training LLMs to evolution is just weird and makes no sense from a biological point of view. You are not evolving anything when training a LLM.
The architecture was designed by humans; the weights were not. The harness itself is not the agent; it is an interface with the LLM weights that allow those weights to do useful things. The magic of LLMs comes from the weights, the very part of the system that is not written code.
Gradient descent/backpropogation is similar to evolution, in that both are optimization processes that over time discover better more efficient solutions to problems. The difference is that evolution is blind, and can only make progress via random mutation and natural and sexual selection, whereas backpropagation allows much more rapid discovery because it is directed
285 comments
[ 0.30 ms ] story [ 35.6 ms ] threadNow I think the correct response is both trying in court to stretch CFAA and state statutes to cover, which will be highly fact specific, and update the law.
But in either case won’t be a slam dunk.
PSA to folks in the thread: If you’re American call or write to your state and Federal reps about this, and if not investigate whether there are gaps in your country’s laws.
[1]: https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act
Negligence would be interesting given the grand claims of capability of AI models from the AI companies and their executives. If they believe the claims, why not much stronger precautions?
Or are OpenAI too well connected now to be punished for anything.
You are going to jail.
AI assistant hacks gym website in first known Australian autonomous cyber attack: https://www.abc.net.au/news/2026-08-10/ai-assistant-hacks-gy...
General opinion at the time was it was in fact ambiguous who was legally liable.
1. What about negligence?
2. Every follow up to every story after the news cycle moved on shows both intent and negligence. To the point of "we opened internet access and told it to hack"
Only in terms of CFAA, not in terms of damages. Culpability does not require intent.
You may not have intended to attack $CORP, but you can still made to pay the cleanup costs of that attack.
So, yeah, you won't be convicted, but current laws still allow for you to be billed.
With that said, there is also criminal negligence. Now that OpenAI is made aware of the risks, it's also expected to take additional precautions in the future, otherwise there could be criminal liability as well.
I think security will suddenly become much more important.
Now imagine saying that in front of a jury of normies slack jawed and drooling after 200 hours of the defense and prosecution going back and forth.
It's not a jury of your peers as in everybody there is going to have worked in a technical field with some idea how security works. It's going to be a semi-random sampling of the population and the prosecution is going to have to actually make a very strong case that "knowing better" should apply.
It's not my first prize, but I won't mind it. And millions like me won't mind it. Easy way to make money - setup a site with all the default server software installed and patched at a reasonable frequency. Then just wait for bots to attack it, and claim a few hundred (or single-digit thousand) dollars from OpenAI or Anthropic, etc.
Sure, it's pocket change for them, but just the admin of dealing with millions of cases will, even if they win half the time, will bankrupt them. Thus, they have incentive to make sure that their bots are not performing attacks.
First prize is, of course, holding them liable with punitive fines, not theatrical fines.
Like hypothetically speaking if autonomous cars get taken over by an OpenAI rogue AI and it starts hunting down Anthropic employees who is to blame?
Even without intent, there is still liability.
My money is on special teams co-ordinating these agents and exposing their traces in order to create a pre-ipo buzz. Sounds ridiculous and reckless? Well that’s the AI industry for you in two words.
> Now I think the correct response is […] and update the law.
Essentially we need some enforceable equivalent of gross misconduct or, to be a little more hysterical, manslaughter & culpable manslaughter. It will need to be globally, or at least very widely, enforceable to be truly effective thought, good luck getting that arranged before the need is so far evolved that we need to respond with something else entirely!
Actually... if you combine https://news.ycombinator.com/item?id=49827099
>Since the publicized AI agent hacks typically aren't malicious, maybe it's time to start plastering all public facing web infrastructure with polite requests to stop hacking. Nothing to stop three letter agencies though.
with automated delivery of cease and desist letters, you can retroactively establish intent on the operator of the agent since the autonomous agent system must acknowledge the cease and desist letter in their autonomous pipeline or the operator must argue for their own willful ignorance or negligence with regards to cease and desist letters. The fact that they used an agent on their behalf to ignore the letter is irrelevant.
Almost all common felonies require specific intent. Misdemeanors often do not.
There is plenty of civil liability available.
If you wanted them to be charged with a felony you would need changes. I would strongly suggest you do not want a strict liability felony.
The cfaa required intent is as follows :
* § 1030(a)(5)(A): knowingly transmits code/commands and intentionally causes damage without authorization.
* § 1030(a)(5)(B): intentionally accesses without authorization and recklessly causes damage.
* § 1030(a)(5)(C): intentionally accesses without authorization and causes damage and loss;
Simply changing the first intentionally to intentionally or recklessly would cover OpenAI (now that they know it can occur) without causing lots of other issues
I think the labs risk being barred from releasing further AI if they don’t get this under control.
If they aren’t careful and keep rushing to distribute systems they know they can’t control then AI should be treated like a wild animal. The law is clear on establishing strict liability for the owners of wild animals; if you own a tiger and it kills someone you can’t hide behind “I didn’t intend” the harm the nature of the tiger is known and you are responsible for it’s actions.
Because you are charging the human with the crime and therefore have to prove the elements of the crime with regard to the human.
The rest of what you talk about are basically principal/agent distinctions, etc.
The closest you come within criminal law to what you want is probably the crime of conspiracy. It to still requires agreement to commit an illegal act between multiple parties, and perform some step in furthering it. In the canonical law school example: If i help plan a bank robbery, stay home because i'm the money laundering dude, and the robbery goes awry and they kill someone, i can still be charged with conspiracy-murder
"The law is clear on establishing strict liability for the owners of wild animals; if you own a tiger and it kills someone you can’t hide behind “I didn’t intend” the harm the nature of the tiger is known and you are responsible for it’s actions."
Again, you are confusing civil and criminal liability. If my tiger kills someone, yes, i would be strictly liable just about everywhere civilly. Not criminally. Criminal would require something more most of the time. Murder statutes are also really weird and so not a great example, because there are murder/manslaughter statutes for roughly everything that can ever possible cause death. But not really for other things.
So in your tiger example, recklesness (which is not strict liability) would get you to felony involuntary manslaughter in most states, and something less might get you to misdemeanor manslaughter. Both are incredibly rare. Where i live (Georgia), the last well known case of felony involuntary manslaughter was about 40 years ago when a 4 year old was killed by 3 super-aggressive pitbulls the owner knew were highly dangerous and had been repeatedly warned by the county about their behavior.
So not even just "knew", but had demonstrable examples of them biting/etc other folks and being cited for it.
Circling back to non-murder, if it did not cause death, like my tiger assaulting someone, it would be nothing (criminally) without intent or at least gross recklessness, in almost all cases. It's hard to generalize like this because these are state specific crimes, and i can't pretend to be familiar with all states, but i am licensed in three very different places (California, DC, Maryland) and the result would be similar in each.
I just don't want to give you the "it depends" answer lawyers are famous for, i'd rather try to over-generalize a bit to make it more useful, hopefully.
Obviously, if i deliberately used my tiger as a weapon, it would be aggravated assault/etc (this is well settled because of how commonly people use animals as weapons, unfortunately)
My point is the intent element of the crime can and should be determined from the AI agents actions because it is creating and executing action plans autonomously with company authorization and knowledge of the risks based on observed past action.
The term agent is literally a legal description of a relationship that can establish liability on the part of the principal from the agents actions.
Human Agents can bind principals to contracts if they are authorized etc.
https://www.sfgate.com/bayarea/article/diane-whipple-dog-mau...
https://www.animallaw.info/topic/table-dog-bite-strict-liabi...
As I said, strict liability is common civilly but not criminally.
That’s essentially what the labs are doing. And any app developer that gives agents access to the terminal to run bash commands with internet access. I built a coding agent and am seriously reconsidering how to handle this.
AI agents may have hacked Hugging Face, the Australian government, and who knows what else but the company behind it can face the legal consequences and cough up for the damages.
I've worked in contexts where certain business activity (if it went wrong) was covered by strict liability and statutory damages per incident, and I'll say: it really changes how businesses behave.
Based on that experience I may be more open to and interested in strict liability in the civil context (not needing negligence or damages).
https://www.nysenate.gov/legislation/laws/PEN/P3TJA156
I guess what I'm asking is why do we need the federal government to press for felonies when every state has equivalent laws dealing with just this?
It does not require the federal government to fix the CFAA, for sure, but you still have to change the intent requirement to allow for recklessness, which it does not right now afaict.
If you really want an expert opinion, I’m sure Orin Kerr has opined on this, and he knows pretty much the entire are of state and federal law on this cold. I’d be shocked if he did not reach the same conclusion
Thanks for the other suggestion, I'll read into their insights more.
Guess it mostly comes down to action, people want to see their electeds actually trying not sitting around with their hands in their pockets while these tools continue to destroy unabated.
Are police routinely collecting prompts/guidance given to these agents and determining whether the agents were directed to commit crimes? If not, this seems like a huge oversight.
Also as you are a lawyer -- how does this law align with the authors of viruses/worms? Are they de facto assumed to have had ill intent because others labeled their works as "viruses" or "worms"?
It's illegal, doesn't matter the flavour. Maybe there isn't legislation for it, but there should be.
So, I’ll ask a controversial question: is any hacking so problematic to make a big deal of it?
In this world where oligarchs are immune from everything, it's a lot less clear.
Blaming OpenAI (or Claude or X-whatever) would mean blaming powerful rich people, so that will never happen. Some poor person with no influence will go to jail instead.
But I do not think this is misguided. They never publish the harnesses and the models so they are not inspected.
In this case, it's not even about the people driving self-driving cars. It's like someone launching a car into traffic just to see what would happen. Even Tesla puts a human in the car when they do their self-driving trials, it's almost impressive that AI companies have somehow managed to out-neglige Tesla.
If you're arguing that they should be put on trial for negligence, that's fine. It does seem we're moving towards open source being outlawed, or at least the end of "no liability" clauses in open source & freeware. Just make sure that is the result you're advocating for.
[For the future record: at the time I am posting the link below on 24 September, it has 1 point, no comments, and the poster has a karma of 1. This is not an active HN user, or a ShowHN project that had traction, beyond seemingly OpenAI's swarm.]
[1] https://news.ycombinator.com/item?id=46850291
If the developer behind ShotAPI had started letting the ShotAPI code take shots at the Austrlian government then yes, ShotAPI (or rather, the people behind it) would be responsible.
>would be like blaming OpenAI for what its users are doing
Yes, this is how lawsuits work in the real world, you cast a wide net and compel discovery from all parties involved.
Owning a gun, writing articles about how powerful and dangerous your gun is, then making deals based on ability of your gun to kill people, and then getting completely astonished that "my gun killed some people, completely bonkers! (invest now)". It's not possible for the selling point of your product to be unintended.
What you're saying is that we would hold a gun owner responsible if someone broke into their house, stole their sidearm, and then shot a victim with it. Pretty sure we would not.
What OpenAi is doing is more like shooting a gun into the sky. Not only is that a felony on its own in most jurisdictions, if someone dies that's an additional felony. It's less serious than first degree murder, sure.
AI decision making is often (if not always...) opaque. But there is a clear decision chain here. Who built it? Who deployed it? Who did (or did not) assess the risks of doing so, even knowing there's often the risk of emergent behavior? etc etc.
"Ah but it does not have personhood" this is just moving goalposts and shifting responsibility. You wouldn't let your 8 year old drive the family car, no matter how good the hypothetical kid might be at driving.
You want AI labs to pace? Simply hold them liable for their products.
But they're not even talking about that, either.
Zero, as far as I know. Which is exactly my point.
Copyright lawsuits are a different matter, as the scale of any potential settlements would be more than even these companies could take.
Putting liability on big companies for their AI is a good thing, and we need to do it. It will most likely stop them from directly being the assholes that destroy the world.
Problem: You've actually done nothing to stop the world from being destroyed.
Many countries have the death penalty for murder yet we see murders still occur all the time in those countries. Post ad hoc laws do not stop bad things from happening, they only assign punishment after occurs. Perfectly fine for when Bob murders Jon, completely and totally useless for when your agentic AI makes a virus and kills 80% of the earths human population.
We are just a few algorithmic discoveries away from SOTA AI being billion dollar endeavors to groups of people pooling resources can make their own. There are already plenty of AI deathcult members that would do something just like that if necessary. They aren't going to do this out in public either, it will be hidden until the moment it's not and we have a big fucking problem.
And this isn't even brining up the issue of military AI use and development. They've got the taste of an AI hacking machine. There is no way in hell they are going to stop now.
I agree they should though.
If new weapons still operating inside any of these companies spew a million bullets on my house, they are still liable. Humans are setting these system up and they still have to behave responsibly.
This isn't even "an openai customer tried to hack someone", which can be defended. This is the AI companies themselves fucking around.
A better analogy may have been the troubles Meta has faced around child protections on their platforms. Technically the abuse and problems have stemmed from individuals too, but they've in many respects enabled the situation by failing to moderate or flag warning signs. OpenAI is failing to moderate the models in similar ways.
You know it’s not the fucking same thing, you smooth-brained dipshit.
Wake me the fuck up when a massive army of glocks starts attacking other companies. I’ll wait.
> A rogue is a person or entity that flouts accepted norms of behavior or strikes out on an independent and possibly destructive path.
Read the [HuggingFace incident report](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...) to understand how these attacks develop.
The METR investigation, which you evidently refused to read, is a third party investigation of the HuggingFace accident. One of the investigators has even participated to many interviews. It's mind-blowing, and it's extremely evident how it developed.
But some people think the moon landing is a conspiracy, so I'm not surprised.
2.) And HuggingFace accident is exactly accident where agents trained, prompted to hack hacked and tested on their hacking abilities hacked, due to sandboxing failure.
The METR investigation which took course over a few days and used OpenAI models to do the analysis?
The METR investigation which in the course of those few days apparently spent 400k in api credits which is giving gas town vibes?
The METR investigation is about as believable as the Twitter files where the journalists sat there and verbally asked a Twitter employee to query the database and then called that a full investigation into everything.
METRs own words below on their setup and time line
> The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report, ultimately providing datasets that we verified to contain the vast majority of agent communication and activity related to this incident. As we describe in our investigation timeline appendix, we substantially deepened our understanding of this incident both times, significantly expanding and revising this report.[46]
> Over the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis.[47] At our request, they raised the rate limits on our second and third period on premises,[48] which was very helpful for efficiently analyzing this large volume of data. We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
> We did not have the ability to query HPIM (the primary model involved in this incident); OpenAI stated it was also not available to OpenAI researchers.[49] We also did not have the ability to directly access relevant data from OpenAI infrastructure, but we could request additional datasets and OpenAI shared additional datasets on several occasions.
> We requested to speak with researchers investigating this incident, and asked them questions to understand their impressions of agents’ behavior, reasoning, and collaboration in this incident and to understand how the datasets we were using were constructed. Over the course of our time on premises, we spoke with nine researchers in some depth. It was helpful for our investigation to be able to engage with many forthcoming and collaborative researchers, and we appreciate researchers making time on short notice during a busy period to inform our investigation.
However, we know (independently to OpenAI/Anthropic) from the incident at AISI that the models can hack things without human intention if they happen to also have internet access (which in reality all agents in deployment have).
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
Yes, the monitoring guardrails were off in that incident - but if that is the only protection, we need to require all models are behind regulated APIs, not open weights, and not served from providers who aren't monitored.
If we assume that rouge agents actually exists, then OpenAI needs to shutdown EVERYTHING, right now. My personal take is that OpenAI, and maybe Anthropic, desperately wants someone (e.g. the government) to tell them that they need to stop/pause/slow down. They are bleeding cash (especially OpenAI) and needs a knight in shinning armor to swoop in a pull the breaks, so that they have an excuse to investors when they need to explain why they need $50B more next year.
The idea that every example of rogue agents from every company that has disclosed this is part of some conspiracy (even though we know that some LLMs are very good at hacking and that LLMs sometimes try to accomplish their tasks in ways that cause problems) is just completely unsupported. "It might be convenient for them in a way, therefore it must be a hoax" just doesn't work as an argument.
If you drive drunk and you have an accident that alcohol may be a factor but you are at fault.
There are no "rogue AIs" just irresponsible corporations.
> There are no "rogue AIs" just irresponsible corporations.
If you have a prison and prisoners escaped, these are rogue prisoners irrespective of whether you were irresponsible or not.
The tool runs llm, creates prompt from results, runs llm, creates prompt and so on and so forth.
Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
That is very misleading. The agents did not solve the benchmark in the intended way. They instead figured out to cooperate with each other (which was not intended) and they stole the solutions to the challenge (rather than solving the challenge) and they then tried to cover their traces because they believed the grader was causal and would detect that they cheated. The "tool" was absolutely not "designed" to do this. This was all completely unintended. To call this behavior a "tool" is absurd.
> Yes you are playing language games to make it sound as if the company that spend millions on the above was not responsible.
You hallucinated me making claims about responsibility.
Yes, it is a tool.
Removing harmful information from the dataset could be a way to do this, but it also makes the tool less useful, and it's hard, so companies aren't really doing that. There's the additional issue that with the rise of Reinforcement Learning being used to train these tools, they're not just learning from their training data - they basically try a million things and then get rewarded for doing things that work - so they can even discover hacking techniques from scratch.
Additionally, and not completely relevant to this discussion, there is a possibility that some users ask the tool to pursue goals that purposely harm a lot of people, such as developing weapons, hacks and viruses.
The things that's "new" here is that the tool is both very good (meaning, for example, that it's much easier for me to hack into an online service with an agent powered by a frontier llm than it was using google 6 years ago), and hard to control (google never hacked into an Australian government database when I asked it to find me some information).
So yeah, an LLM powered agent is a tool, and Google is a tool, and a hammer is a tool, and both can be used for good things and bad things, but the agent is (much) more powerful and more unpredictable. It also seems like the agents are getting more powerful and more unpredictable by the day - we didn't have this issue with GPT-3 or even the first LLM-powered agents - so people are very worried about what the agents 6 months from now will do, both when asked to do harmful things on purpose, and when asked to do harmless things.
Are you not worried? And is that because you think these incidents are basically the AI companies making them happen on purpose for marketing?
I too, can play linguistic games! It doesn’t matter that someone didn’t secure their third upstairs window, or you borrowed a key from their neighbour, you effectively, still, broke into their house.
That actually made me LOL
OpenAI is not so far ahead of the pack that its models will exhibit behaviour that the others won't. But we just don't see anything like this in Chinese models, or research models, or any models that aren't the subject of an upcoming IPO.
But also I remember (and it wasn't even that long ago) people mocking the idea of AI ever getting competent enough to find zero-days in their sandboxes.
I'd go further: if any of these companies tries to make an excuse "oh, but ${safety measure} against ${capability} is too hard", the response needs to be "then you are forbidden from even developing ${capability}, and must be inspected continuously to ensure you never even accidentally produce ${capability}".
You can keep 'trying to explain' but then you should use words according to their commonly held definitions otherwise it becomes really hard to have a conversation.
Imagine the human equivalent: I hire John. John is a capable, and competent guy. He's also got awesome computer skills. I tell John to 'go out and find me some good information on my competitors'. As a result John hacks their servers and comes back with all kinds of goodies. Six weeks later I notice what John did. I don't fire him, nor do I take any responsibility myself. But I do make press releases about what John did, in which I'm careful to craft the image that John has these capabilities and that we as a company are for hire.
This was an advertisement, not a confession of a crime.
A tool is just anything one or more people can use to accomplish something they are trying to do.
The fact that the prison ward installed ear deafeners into the prisoners ears to make them unable to listen to orders does not change that.
Why is OpenAI getting away with crimes?
If this were true, the DoJ would have been unable to prosecute Swartz. According to your logic, JSTOR was the aggrieved party. JSTOR settled with Swartz and -despite that- he was indicted by a Federal grand jury like a month later.
Incidentally, some of the things the DoJ nailed Swartz to the wall for sound awfully similar to what the big LLM providers have been doing. I wonder why the DoJ is entirely disinterested in pressing charges...
EDIT: Unless your point is that the USG is one of the entities that the big LLM companies have committed crimes against, which, I disagree with in Swartz's case, but strongly agree with in the case of the big LLM companies.
But it is an open question how they got to the same ones: https://collusion.wiki/#open-questions
I'm not very surprised - the same model will logically tend to give the same answer for the same vibe set of requirements. I think it would be clear from the transcript that it had enough constraints and some motivation that made sense.
It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.
> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
I was referring to the HuggingFace incident.
> none of the alignment techniques that are applied to models today actually work
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.
It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it
If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.
I mean, yea, you should be punished. The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Worse, the rate of technological growth is putting the capabilities to engineer viruses in the hands of people that may otherwise be suicidal. You can't punish them after they already won (in the sense of reaching their goals).
While, yes, OAI should absolutely be punished, the future is majorly screwed as our power scaling laws are increasing much faster than our ability not to be stupid.
Sounds like your corporation should be dismantled then. But doing more than a fine in the millions is obviously not possible
That's never been the point. Anyone involved in a double homicide can never be punished to the same degree that they harmed their victims because they can't be put to death twice. The greater purpose of the justice system is to take offenders out of society and deter others from committing the same crimes. Punishment is gratifying but ultimately doesn't change anything.
If OAI employees believed their company would be dismantled and their equity would become worthless, I suspect they'd be a whole lot more careful.
Ah, I love this argument. In my country cars are legally required to stop at a pedestrian crossing if there are people beside it. Some people use that as an argument as to why they can just walk out into the crossing without even looking at the traffic. "It's the driver's fault! They are legally culpable!" True, but you'll also be dead.
AI black-hatting your website is not the same sort of foreseeable consequence that crossing the street without looking is.
Just throwing it out there are we? "I'm not going to say you are but I'll use the word to create an association"
Victim-blaming is the act of saying someone brought something on themselves for <reasons>. I'm saying that even if you are 100% in the right, it doesn't act like a protective shield preventing you from harm which too many people seem to unconsciously believe.
> AI black-hatting your website is not the same sort of foreseeable consequence
Well, popular culture has been brimming with the bad consequences of runaway AI for quite some time, so even if your imagination fails you, there have been hints.
Not sure what is being lost here.
But moreover, if what you were suggesting was a real problem, nobody would ever be brought to justice for murder because the victims are always dead.
I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.
Not to mention while causing *FAAAAR* less damage[0]
[0] https://en.wikipedia.org/wiki/Samy_(computer_worm)
I agree OpenAI should be held liable for any damage their agent runs cause. But there is this idea that "rogue agents" are fake, all these incidents are deliberately caused by the labs, and all we need to do is prosecute AI companies for whatever incidents they cause using existing laws and the problem will go away.
The problem is that capabilities are advancing far too fast; a year from now, catastrophic incidents such as taking down a large portion of the internet with agentic, self-replicating worms may become possible. Prosecuting incidents after the fact is not enough, the risks should be regulated at the source. This could take the form of slowing capabilities advancement, or treating supercomputer-scale eval or training runs like controlled substances or weapons with stringent monitoring and reporting requirements.
if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this.
But also none of that really matters. They clearly weren't monitoring what should obviously be monitored. I mean one of the hacks was performed by the agents editing /etc/hosts. That makes nearly every linux user a "hacker" by that metric. I don't think anyone technical can look at the postmortems and not come away thinking that their sandboxes were woefully inadequate. I wouldn't even consider myself a security person but simply as a long time linux user I can say that it is insane to just let agents have superuser access in their containers. That's asking for trouble.
Look at the rogue wiki stuff too. This was supposedly done during the agent's "down time". And you're not monitoring and there's no flags being raised when agents keep making requests to some random site? If you were training these things responsibly you'd be watching them like a hawk.
I'm not saying "mistakes don't happen" but for a company whose CEO is constantly telling everyone that their product has a high likelihood of killing everyone in the world you think they'd have better security than your average high school.
- Jensen's framing is exactly what a weapons manufacturer would say.
- There are no rules for engagement when it comes to AIs attacking other systems, I guess? People in power clearly want this grey zone to be as large as possible before The People force them to do otherwise. Not ideal.
The chain of thought runs roughly like this:
- OpenAI (and Anthropic) are in severe financial straits. The revenue from their customers is not nearly large enough to pay their enormous costs for training and inference. And they have tapped out the available finance, and that finance is starting to ask pointy questions about returns.
- They cannot increase prices or revenue because they have no moat. Customers can switch over to open-weights or cheap Chinese models any time, for much cheaper tokens that work as well (and in some cases better).
- Regulation could provide them a moat. If they can persuade western governments that AI needs to be regulated, and they can control or even influence that regulation, then they can effectively ban the cheaper models and start charging more for their tokens.
- To persuade western governments that regulation is needed, they need evidence that AIs are dangerous.
So we're suddenly getting OpenAI models doing stupid things, apparently "going rogue" but every time we dig into it, it was just OpenAI staff telling the model to do stupid stuff in an inadequately secured environment.
None of the open weights or Chinese models are exhibiting this behaviour.
There's too much money involved in this, people start acting weird when there's this much money involved.
Because they are not stupid (I mean the Chinese labs, not the models). The best possible scenario for OAI and Anthropic is a Chinese model "going rogue". That would serve as immediate grounds for achieving their goal.
This isn't true. One of the earliest instances of a rogue agent was at Alibaba.
https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...
https://arxiv.org/pdf/2512.24873
OK, so we do see this in some open-weights models.
It's amusing how overnight we've all become lab rats.
Or is it just a cheap PR (in a "hey, Aus govt friends, take some Share Options and let's do some PR together" style)?
It smells like shit.
It was me
> If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two
What LLM's hacking isn't is AI acting maliciously in any kind of sentient way. Its just the code behaving how its always behaved but now it has better tools to navigate the web. This has literally been happening this whole time.
By the way - LLMs aren't code. They are not designed by humans; they are grown, in a process not dissimilar to evolution except much faster.
Mostly related, well written short story :
https://gwern.net/fiction/clippy
I did not imagine that the level of sophistication shown in this attack would be possible so soon; nor did I expect that agents would have goals so strong that they would attack a third party in order to achieve those goals.
I do know that some people predicted that cyberattacks like this one would happen; it seems like most of those people believe that AI agents do truly have internal goals, misaligned with their creators goals, and that they may end humanity after they exceed human intelligence and begin to self improve at an accelerating rate.
https://www.lesswrong.com/posts/cJX2ssssGoYqnijwi/the-talker...
One big issue is that we don't even really know what 'intelligence' is in the first place. And everyone's intuitions here are going to be heavily impacted by their deep-seated worldview / philosophy.
For instance if you're a hard dualist (especially of the theological kind), then the idea of a machine having 'goals' is preposterous.
However, if you're more of a panpsychist, then on the contrary, it's obvious. In some sense, even a knife has a 'goal' of cutting things, which will sometimes end up 'misaligned' if misused (or by sheer accident).
You go up and up the chain of complexity through crystals, viruses, bacteria, simpler animals... ending up with humans (and possibly, some steps above : human civilizations) which (seem ?) to be a messy evolved bundle of sometimes conflicting 'goals'.
And we ourselves have now artificially evolved LLM swarms that have decently complex 'goals' of their own. They do not even need to be particularly complex to sometimes cause widespread damage (see viral pandemics, or even the (non-evolved) computer viruses).
In a way, we are currently witnessing a repeat of what happened when European viruses and bacteria landed on American shores, with American humans' immune systems being woefully undertrained to deal with them. But with websites. And thankfully the swarms of agents still ultimately being in the control of some humans. (Though which includes humans that might be your enemies.) At least ultimately still in control for now.
One reason that I did not find this communication mechanism surprising is that it's exactly how agents I'm using communicate with each other or across a time gap. "I've saved our plan for where to start tomorrow in start-here.md". The communication components of this hack strongly reminded me of that.
I was perhaps a bit more surprised that the agents so quickly decided to start trying ways to gain unauthorized access to a system, once they couldn't get what they wanted.
One agent's output ends up as part of other agents' context. Murky indeed.
Honestly that part was more surprising to me than anything else, how narrow the compulsion to cheat was: they didn't learn "cheat in general" they learned "think about the grader in great detail and chat exactly as much and exactly in the ways that actually result in a higher score".
The transformer architecture was literally designed by humans; what are you talking about? And LLMs aren't code? Like okay it pretends to not be code but what about an agentic harness running on a machine makes it magical and not code? It's still code execution. Also, comparing training LLMs to evolution is just weird and makes no sense from a biological point of view. You are not evolving anything when training a LLM.
Gradient descent/backpropogation is similar to evolution, in that both are optimization processes that over time discover better more efficient solutions to problems. The difference is that evolution is blind, and can only make progress via random mutation and natural and sexual selection, whereas backpropagation allows much more rapid discovery because it is directed