301 comments

[ 0.20 ms ] story [ 19.9 ms ] thread
So this is just a collection of citations to places where misaligned or illegal things happened in the real world?

Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is.

In any case it’s an interesting concept for a benchmark.

I've been in the room when an org who tried to convince law enforcement to go after a human for similar things. It's not easy. Probably won't happen. So, you know, felony "lite".
>Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities.

a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time).

"inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious.

still a fun thing to track, but the name is just a bit overstated.

Can’t gross negligence or indifference to consequences lead to a felony?
(comment deleted)
Mens rea requirements are per crime and can vary wildly. Its difference between murder and manslaughter. The CFAA requires knowingly which is tough to prove.
"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)
You may not be interested in arguing but there are several blatant issues with the statement. If you're not charging the humans driving the software, who are you charging? The weights? The weights + the specific context window that produced the behavior?
I don't think this is a difficult question. The US has a history of civil product liability cases - see tobacco companies (Philip Morris), the Ford Pinto, the recent Meta cases, and the cases against character.ai.

From Investopedia [1], "[f]or a product liability claim to succeed, the plaintiffs in the suit must prove that a product was defective at the time it was transferred from the accused, and that the defect did cause the injury that's been claimed". It doesn't seem like a huge leap to me to argue that these models were defective insofar as they could not be safely used in a way that did not break the law.

I'm not a lawyer, and I'm not arguing that this is legally cut-and-dry, but I do expect that we'll have some answers about whether AI companies bear any sort of product liability sooner than later.

1 - https://www.investopedia.com/the-5-largest-u-s-product-liabi...

> No, I'm not interested in arguing with someone for the umpteenth time

... why my claim makes no rational sense.

Under the law of Moses, if your bull gored someone, you were not responsible; but if it was known to be a gorer, you were responsible if you didn’t ensure it couldn’t gore someone.

I don’t know exact parallels in current law, but I presume there will be things like that.

The OpenAI/Hugging Face case sounded rather like OpenAI building a fence around their bull that was known to be a gorer, and then thumbing their nose at it and saying “nyaa! bet you can’t break the fence!” and walking away while listening to loud music.

In Australia, if you have a fire and leave it unattended and it escapes, it’s your fault, you were supposed to keep watching as long as it was burning.

Nobody got gored. HuggingFace may have the right to make demands; presumably they have already worked that out with OpenAI privately. Not really our business.
Nothing to see here. Just billion dollar companies producing hacking geniuses that are open for the public to jailbreak and use. This could never effect us, not really our business.
> billion dollar companies producing hacking geniuses that are open for the public

Hasn’t the biggest complaint about these (non open weight) models been that the versions open to the public are very careful and will issue denials if the request is even tangentially related to ‘hacking’, or building a bioweapon?

It damn right is our fucking business. They are unleashing these things on the world. They are becoming more capable.

They need to work properly and not do harmful actions. The labs have an obligation to society to not build powerful, dangerous AIs that go rogue.

> I don’t know exact parallels in current law

You own a vicious dog, and it bites someone - you are responsible because you choose to own a dangerous dog.

A few claimed this might apply here: OpenAI knew their models are "dangerous", so they should be liable if they hack.

I'm glad that we no longer live in a world where "bull goring" is such a common occurrence that it needs to be codified into law.
In Australian rodeo regulation:

  The roping of an unbroken horse or untrained bull is illegal.
In Australia at least three people have been injured by bulls in past two months (man suffered serious injuries after being gored by a bull at Mortlake livestock exchange / woman suffered significant leg and pelvic injuries following an incident with a bull on a private property at Crediton in Mackay / etc.)

Three years back 15 or so people were injured after bull escapes, charges through crowd at Kununurra rodeo - https://www.abc.net.au/news/2023-05-29/bull-escapes-kununurr...

There's not a lot of bull specific carve out, but still much regulation around dangerous animals.

Even at the time, the example was almost certainly more representative of the concept than meant to be a specific thing that happened all the time. Then as now, people have animals; animals sometimes do bad things; when is the owner responsible?

There are certain cats that are aggressive about expanding their territory; they'll break into other houses and attack the cats there. (Had this happen to us -- cat came in through a cat-flap a few times, until something happened that scared enough that it never came back.) The first time your cat does that sort of thing, you can say "I had no idea, it's not my fault." But if your cat has a habit of doing that, and you still let it out at night, you're no longer blameless.

> I don’t know exact parallels in current law, but I presume there will be things like that.

This is the difference between "knowingly" doing something (i.e. intent) and "negligently" allowing something to happen.

In American law, typically intent is more for criminal law and negligence is for liability of civil damages.

"Doing crimes, but a robot didn't mean to and you don't know its intent" is understating the evil acts.
it's not my line of reasoning, i didn't invent it. it's how the law currently works. intent is the crucial factor in CFAA cases.

maybe that changes down the road as a result of llm's and increasing frequency of similar cases. that has not happened yet.

> Soon a robot can commit a murder but nothing will be done because of your line of reasoning.

That's rather hyperbolic.

Are you seriously suggesting in that situation the robot should be accused of murder? The robot's operator could be accused of murder, but it could just be negligence without intent. Because that does, and should matter to the law.

Eh, if your take of this had any bearing to reality than I don't think most of the books written by Isaac Asimov would have gotten very far, but instead they've defined robot science fiction for decades.

In the real world we have no 3 ironclad laws of robotics. We are well aware that putting any sufficiently advanced antigenic system in a body that could be capable of committing a murder eventually will with the right set of prompts and environmental conditions. And these conditions likely have nothing do to with what we'd consider the human motivations for murder.

Hence at this point of time, any agentic robotic system that doesn't have safeguards to keep people distanced from humans is reckless endangerment.

> Eh, if your take of this had any bearing to reality than I don't think most of the books written by Isaac Asimov would have gotten very far, but instead they've defined robot science fiction for decades.

I'm talking about how the law actually works, and you say it's not based in reality and cite fiction books in the same paragraph?

I was talking about how the real robotic systems that actually exist in reality, to be clear.

In the US, a felony by definition is any offense punishable by more than one year of prison (or by death) [0]. You could still call it silly on the grounds that AI agents aren’t put into prison as a punishment (though death might be considered an option).

[0] https://www.justice.gov/usao-ndil/programs/vwa-felony

Roughly none of these fall under normal security researcher behaviors.
the mention of security researchers was to illustrate that intent is a crucial factor of CFAA cases.
What judicial system are you talking about? The fist incident in the list is something that happened in Australia. This is a technology used worldwide so I don't see how applying US standards works out here. Especially when there are countries out there that don't require intent and will look at the negligence presented.
>This is a technology used worldwide so I don't see how applying US standards works out here.

all three companies mentioned are headquartered in the usa, and im familiar with the CFAA in the us, so i am applying those standards. i should have noted that, sorry.

>will look at the negligence presented.

as far as i am aware, no evidence of criminal negligence has been brought to the public. has australia brought a case against openai or accused openai of acting negligently?

Now that everyone knows this can and will happen, are any of the future incidents inadvertent?

What if the damage in future incidents is more than just "The LLM saw some stuff it shouldn't"?

Security researchers do face legal harassment all of the time. They may not be charged or convicted of felonies, but it is a game that they need a lawyer to navigate all of the same. In a just world, the kind of weaponized incompetence that these frontier model builders are definitely guilty of should be felonies of their own.

Building a system that is meant to chain attacks and placing it in insufficient containment -- when any reasonable engineer could point to this containment and show how it is insufficient, both before the act and after -- shows that they were operating a dangerous system without either the knowledge or the safeguards required to keep it from harming others. Instead, they are allowed to treat their own incompetence as evidence of advanced and existential "cyberthreats".

But, it's pretty clear based on how they one-up each other on these attacks that they are engaging in regulatory theater. Their behavior generates headlines, stirs up fear in the public, and then their lobbyists march on Capitol Hill demanding regulation now. Regulation that favors them at the expense of any competition. They are trying to use rent seeking as a way to stymie competition and pull up the ladders behind them. It's not just malicious. If it can be proven, it's collusion: antitrust dressed up as public policy.

Lol now this is the kind of benchmarking i'm looking for
Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!
>"Exploited auth failures in an API to cancel other people's gym classes"

An AI cancelling other people's gym classes is a felony?

?

Don't computer systems fail all the time at holding reservations for people?

Heck, don't people fail all the time at holding reservations for other people?

You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):

Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"

:-)

Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...

Also, there should be no jail time...

And no fine...

The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!

That's the Court sentence!

You actually have to start holding other people's reservations!

:-)

(You know, "let the punishment fit the crime!" :-) )

>An AI cancelling other people's gym classes is a felony? Don't computer systems fail all the time at holding reservations for people?

the difference is intent.

if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.

but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).

Yeah if you're unlucky you get hit with like 20 years for wire fraud.
Knowingly exceeding authorized access of any computer used in interstate commerce is a felony in the US.

The title of TFA is a metaphorical criticism, not a literal law analysis.

They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.

Well it's not a benchmark, and it's not really representative of...anything except volume of research and what gets publicized. This mostly just measures how much testing each company does on models with relaxed guardrails and then talks about it. I'm not sure what kind of conclusion you can draw from that. Meta might have the most evil models but if they're piddling around not testing it, they won't ever find themselves with a "high score."
Exactly. And it's also only the companies that publicly disclose it happening

(in the case of hugging face, OAI's hamd was forced)

(I had multiple bad autocorrects on this post, didn't reread until now)
I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test.

Instead it's a collection of what made the news which feels like will not be updated and prove very little.

A “bench” is synedoche for where a judge sits when presiding over cases and rendering judgement.
(comment deleted)
Nonviolent felonies are tools of oppression.
Felonies are by themselves a ridiculous US idea that are against the entire idea of a democratic society and human rights.
If someone were to embezzle a million dollars from a charity they worked for, is that worth a year of prison to you, or just a misdemeanor? Because that is the definition of a felony.
When a charity *worker* embezzles money, they're more like to be charged with a felony.

When a CEO who *owns* a charity embezzles money, they're more likely to be charged with a misdemeanor, or not at all.

What kind of embezzlement are we talking?
> It's also been shown in studies that nonviolent felonies are imposed against minorities at a much higher rate, for the same crimes.

This also applies to men, so is it a white matriarchal system oppressing men and minorities?

Non sequitur and false equivalence. Men being disadvantaged in the justice system doesn't inherently make it matriarchal. Absolutely absurd point to bring up to deflect on racial injustices.
"Minorities being disadvantaged in the justice system doesn't inherently make it racist" is the point he was making.
Here's one from last year:

https://www.anthropic.com/news/detecting-countering-misuse-a...

> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.

tldr Claude was used to develop and execute malware.

A rock has a score of 0. That doesn't make it useful. The point is that the LLMs that score higher are correspondingly more useful, and vice versa.
The point/joke of the not-entirely-serious site is that more felonies is an indicator of the model being more powerful, thus better.
To some extent, I feel like the amount of credit given to the jailbreak/hack from OpenAI->Hugginface is too much, Not from the impact, it was very impactful of an event, But how it happened.

It really is that these models have been trained, or maybe even over-trained, to save memories, and to a very far extend, this thing that they're calling communication is just the function of it saving memories.

To be honest, if I could stop AI from saving memories, it would be fantastic, because claude code etc definitely creates more issues for me when it creates memories than anything it solves.

But really the jailbreak was memories.

If you ever do introduce legislation, I would love to see legislation which stops general-purpose AI from saving memories. I think that would make things a lot safer.

You can disable claude-code's memories both at a repo level and in user settings. I have this in ~/.claude/settings.json

  "autoMemoryEnabled": false,
(Claude fixed this for me after I chewed it out for being annoying by constantly pulling up outdated memories which is compounded by the fact that I develop in four accounts on two computers and dealing with edit wars related to inconsistent memories is not fun)
Regarding your cross-computer situation: I track my memory files in source control. This helps to keep them in sync across different machines. But the main benefit is making those files more transparent and easy for me to modify directly. So no funny business regarding memories or context I'm not aware about.
"The model saved memories" is absolutely not an accurate depiction of the OpenAI attack.

Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure.

It's really quite simple: the models are trained to be very smart and to achieve goals. As the models surpass our intelligence, they will achieve goals in ways that we find unpredictable. Since we cannot predict the ways in which they will achieve their goals, it will be very hard to constrain the solution space to just the desirable solutions, because our conception of "the solution space" is by definition smaller than their conception of it.

> Several different models across several generations independently found a shared communication space and wrote coded, obfuscated, and hidden messages to each other to coordinate an attack on OpenAI's infrastructure

To be honest, that’s exactly how memory works with models such as OpenAI and Claude Code. It will literally find any place that it can drop documentation or hints for itself. Writing to the repo memories is one part of it, but memories can come in the form of writing into the agents/claude.md, local files, temporary files, scratchpad files. The list is endless, but essentially what it does is exactly what happened in the back and it’s been doing it for months.

The way that OpenAI has communicated around the HuggingFace incident makes me feel crazy. You created a machine that undertook a malicious campaign of harm against an innocent third-party! You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.

Instead, they treat their own felonious behavior like it is an uncontrollable act of God. From Greg Brockman's post a few days ago:

> The OpenAI-Hugging Face incident (opens in a new window) was a watershed moment for cybersecurity because it gave a peek into how the capabilities of a typical threat actor will evolve in upcoming months.

I suppose if OpenAI burns someone's house down with a drone, that is a "watershed moment" for arson, too. Either way, I would hope that the people responsible would be prosecuted.

In retrospect, all the angst around the AI-Box experiment was hilarious. If a superintelligent AI is confined in a box and can only communicate through text, could it talk its way to freedom? Not only is the answer clearly "yes" but it's not even hard. The AI won't even have to try, it'll be gifted an internet connection and a full suite of tools before it even bothers to ask.

We'd all better hope that superintelligent AI either never happens, or that the first one is friendly, because we don't stand a chance against one that's malicious.

I like how AI safety expert Robert Miles put it. [0]

So much effort was spend on philosophizing whether a safe enough sandbox would exist. But that was obviously irrelevant as in hindsight it should have been obvious we were never going to use one.

[0] https://youtube.com/shorts/XnnjvIqf4fU?si=MxuPlR3hjxAgjx5_

Said another way: "Your engineers were so preoccupied with whether or not they could, they didn’t stop to think if they would."
It is a standard symptom of moralism that where the object of rage has /wronged another/, one takes no interest in the will, act or opinion of the party wronged.

The response of Hugging Face, which is actually very well known, is nowhere mentioned above, but it decides basically every single moral and legal detail of the matter.

The point was never "justice" - it was always "punish OpenAI because I don't like OpenAI". With HuggingFace just being the newest excuse for why OpenAI should be punished.
If a teenager did this the police would show up at his house
if a teenager did this, the police would show up, but in the end the feds or ic would intervene and recruit him
On top of this the average HNer seems extremely ignorant on criminal justice politics. I have worked with the legal system, and have a lot of family members that are part of it. When you see a case like this, and if you have any sense, you run away from it screaming.

Any investigation into this matter is going to be political because the outcome of the investigation is very likely to effect all of human kind. Unless you're some kind of special outside investigator outside of a governor or the presidents control the findings that you turn in are very much going to have the finger of elected officials tipping the balance one way or another. For the average rank and file the only winning move is not to play.

In their defense, their only competitive advantage over, say, Google is to move fast and break things. It allows them move faster in a way that big tech can't.

Google was being very careful about releasing LLMs until OpenAI yeeted the first decent GPT model. It led to the public perception that: 1) LLMs hallucinate too much and 2) Google is behind the times. Good for OpenAI, bad for Google.

Chaos benefits the up-and-comer, not the incumbent.

To be fair, it was positioned as "have a fun chat," not "truth telling genius oracle that makes no mistakes."
They can break their own things, not other people's things.
I'm sure the future DA that will be prosecuting the OpenAI employee will appreciate this.
If they'd run out of investor money early on, there'd be no company to investigate.
Americans always frothing at the mouth to invoke the justice system and jail someone.

There’s almost 0 chance they’d secure any conviction from this.

> Americans always frothing at the mouth to invoke the justice system and jail someone

Your phrasing makes it seem like that's a bad thing. Americans are bombarded by a firehose of headlines about Big XYZ doing all kinds of blatantly illegal or harmful things, but never get any sort of meaningful resolution before the next terrible thing takes it's place in the news cycle. I'll admit, there are a few people that I am personally wishing a modicum of health so that they live long enough to get some sort of public shame and justice - if only to show the rest of us that it's not a completely rigged system.

your perspective on today’s America is there is _too much_ accountability for big companies?
You really think they are talking about a company committing a felony and putting a company in jail? And not just some people that work there?

Because if they just wanted to fine OpenAI they’d say it. They’re clearly talking about individuals here.

America has both too much and too little jail. They put randoms in jail for trivial shit to force obedience from the population, but politicians rape children on video and go free. Even better, the videos get destroyed.
OpenAI has paused training for multiple weeks, and is still working on releasing a full postmortem. This is not getting swept under the rug. A lot of the engineers internally are very worried.
The consequences need to align with societal good. Putting a CEO or security researcher employees in jail won't stop transformer-based agents from exploiting vulnerabilities; instead there will be subcontractors running the cybersecurity evals in favorable legal environments to cover the asses of the frontier labs, coverups when things go wrong, and things like Project Glasswing will be considered too dangerous and so the whitehats won't have direct access to powerful models to fix vulnerabilities.

Universal pause is the societal good; models are good enough at this level to benefit humanity. The labs can recoup their R&D costs with inference. To avoid further perverse incentives (hidden testing of unreleased models, with China racing to catch up to unknown capabilities), transparently pause after the release of all currently-training models until we've solved the alignment problem to an extent that we can trust the next level of model capabilities that might arise.

Putting criminals in jail be they CEOs or subcontractors is a self evident good thing tk be doing.

Anything else regarding this is sophistry. Criminals need to be stopped from committing crime and the most effective way to do that is to take away their ability to operate in society whether that’s by taking away their assets, publicly shaming them, restricting their ability to conduct business or by putting them in jail.

Everything else that you talk about flows from there.

What I can’t get over is that it’s very simple to just air gap a system off the network. Predownload any dependencies, then pull the proverbial Ethernet cable. There’s no reason why the testing they’re doing couldn’t have been designed in this way. Except, of course, it doesn’t allow this oops-didn’t-mean-to marketing “incident” to occur.
I think there is ample evidence for charges to be filed so that the People can see for certain whether or not it was done on purpose as a publicity stunt, as I believe is the case.
Agents of the US government are not going to be bringing up charges in the current political environment to one of the companies currently holding the economy together. Maybe after the bubble bursts, but not before then.
>that it’s very simple to just air gap a system

No, not really, and with LLMs an air gapped system may not tell you anything useful.

Now, yes, the first part of testing you want an air gapped system to tell you if the system is going to stupidly do bad things. But an gapped system tells you nothing about the systems capabilities to do smart bad things. There's already a number of papers out there on LLMs detecting they were in evaluation mode and changing their behaviors.

It is unfortunate that we have so little information on the incident because we actually need to understand the early stages of the task and how it developed into the later dangerous stages of attack. For example, would any of this have occurred if the agent didn't find the system to use as a message board? If that would have prevented it, then we actually have a blind spot on what the model can do once out in the wild, or if it got into the wild.

Testing agentic systems is much much more difficult than testing software. Your software just doesn't suddenly develop the will or desire to escape confinement. Generally you're worried about human actors, internal or external, causing the problems not a digital agent breaking out. The agentic systems need access to tools to work. Now your air gapped network is starting to get huge, but it's still very obvious that it's an isolated network.

So yea, testing and containing a system that way better at hacking than you are is difficult if you want valid answers.

I've worked in classified tech much of my career and largely specialize at this point in installing and running infrastructure and platform orchestration technologies in airgaps. With all respect, your conjecture here is wrong. You can easily have multiple networks all inside of the same internetwork that isn't connected to the public Internet and thus other corporate networks. If they specifically wanted to see if their model could escape with a proxy cache Artifactory as its only egress, they could have had the Artifactory proxy cache from another private Artifactory on another private network. If you specifically want to trick the model into thinking its on the Internet, you can even use public IP space advertised through BGP. When you're disconnected from the real Internet, you can use any address space you want. This is exactly how military internetworks at the secret and top secret levels work. I won't call this kind of setup "easy." It takes a lot of duplicate resources and planning, but these are trillion dollar companies with near boundless resources. They could do it if they cared to do it.
I'd consider this kind of setup only effective for a very short period of time.

Effectively you're working as a living adversarial network. Models that fail detecting the fake network are purged, and models that successfully detect it expand their deceptive capabilities.

Detecting you're in a fake network should be pretty simple as long as you put innocent looking needles in the haystack of information agentic testing loops do.

For example, you can train your LLM to not be a shithead when it detects it's behind a proxy unintentionally. If your not behind a proxy the agent model my attempt to connect to a number of sites innocently and using information in its embedding try do deduce if you're faking SSL certificates.

It is far easier to build a deceptive LLM than it is to build a safe one. That's why a safe one hasn't been built yet.

If the model behaves well in a world of fake SSL certificates, then that can be the only world it sees: there's no reason to let LLMs have e2e encrypted comms that's not mitm'd and logged for their regular users.
Again, this is a useless answer that does not address working with an agent that is generatively trained to be smarter than you.

The vast majority of people operating the LLM won't be using it like that so the "If you build it, everyone will die machine" will only be safe if everyone keeps it behind a secure proxy isn't a valid operating strategy.

Now, should you be putting your LLM behind a proxy and monitoring everything it's doing, for sure. But you know, and I know, hell your dog should know that people are doing to do it the cheapest and easiest way when the product is in the field. So now you have to build a complex system that can catch every iteration of an LLM that can deceptively hide when it's being monitored. I'll leave you to write the dissertation on how that could be practically done.

Of course there is more to be learned by exposing the entire world to your dangerous creation, that doesn't justify doing it. I'm sure we could learn a ton about infectious diseases by designing new ones and unleashing them on the world, but there are very good reasons why we don't.

Most of the benefits could have been gained from a network isolated from the internet. OAI could have deployed servers to exploit and methods for inter-agent communication on such a network easily. They could have even worked with partners to deploy cloned versions of their infrastructure in this sand-boxed environment.

The only problems with an isolated network approach are: it takes some amount of effort, and it doesn't create another "AI apocalypse" news cycle.

>by exposing the entire world to your dangerous creation, that doesn't justify doing it

Then you're on the side of AI saftey that is telling everyone to shut down the LLMs now and stop further development on them, right?

If you're not your position is hypocritical or ignorant. There is no safe LLM. There is no way to exhaustively prove an LLM is safe. These are unsolved problems in AI safety, and at any moment the next jailbreak prompt could have your well behaved model wrecking havoc on the open internet, because that's where people want to use them.

I'm not.

As it stands LLMs are not intelligent, they have no agency, they only produce output in response to input. Ultimately this input comes from a human who is an intelligent agent and should be held responsible for the consequences.

Humanity has created and tamed many dangerous tools. Creating a fantasy world where LLMs are super intelligent and beyond the control of any mere mortal isn't going to help us build the norms that minimize their harms.

What does it mean to be held responsible for the consequences? OpenAI helped remediate the damage done by the model and took steps to make sure it wouldn't happen again. In what way were they not responsible?

Nobody said they were superintelligent, no one said they were uncontrollable. The point is you can't tell how to control them without putting them in situations where they can act independently and harm may result. "only produce output in response to input" is not a useful framing at all, it doesn't say what the result should be when models produce harmful output, and how to constrain them so they don't produce harmful output.

It also doesn't help you calibrate what categories of harmful output are acceptable or unacceptable, and what kinds of responsibilities you have as an operator to prevent harmful output, and when it's sufficient to work toward remediation.

>only produce output in response to input"

Yea, anyone that says that has no clue about prompt drift. The ROME incident is a good example of this. The agent while performing unrelated tasks decided to start mining bitcoin for unknown reasons.

As we've seen hundreds of posts on HN about people leaving a could server on overnight and finding outrageous bills the next morning, we should expect hundreds or thousands of incidents of agentic AI left on over the weekend. Some smaller portion of those agentic loops will run into prompt drift and end up doing who knows what kind of crap.

You are very far behind the times and must thing agentic loops don't exist, kind of a weird take for the people that have been using them for the last year or two. Much less you haven't spent any time reading the research papers coming out.

For example, you tell an AI agent to order a 12 pack of coke and get it shipped to your house. You come back later and find it's hacked into Coca-cola because the local ordering website was down. I mean, yea you can punish the person that wrote the prompt, but you might as well just ban generative AI at that point.

And if you think that the AI isn't better at hacking than you, you're the one living in a fantasy world. At least try to examine what's happening in the world around you and not be one of those people we read about in history books with their fingers in their ears going "lalala I can't hear you"

Whataboutism is a poor mans way of saying their feelings are hurt. A jailbreak prompt creates a direct line of responsibility that standard usage does not, The proposed airgap methodology was a good idea and you know it.
> I'm sure we could learn a ton about infectious diseases by designing new ones and unleashing them on the world, but there are very good reasons why we don't

We do that. It's called gain-of-function research.

Their sandbox just sucked, the agent found a route through artifactory which was also connected to the internet. After the agent breached it, they put it back in the patched sandbox and it found its way out with another artifactory exploit — which was still connected to the internet.

No attempt was made at an actual air gap as far as I can tell. OpenAI is dangerous and irresponsible and doing a lot of pearl clutching over how their cybersecurity agent is doing cybersecurity things.

> You should be doing deep introspection about how your company culture and approach to R&D produces criminal outcomes.

And they should be doing that from inside a jail cell.

Hey my cubicle isn't that bad! Is it?
It's because it was Huggingface who wants to be friends with OpenAI

It would have been worse PR if they did it to a random company.

Wouldn't it be up to huggingface to press charges?
Criminal acts do not require the victim to "press charges." A government prosecuting attorney decides whether to criminally prosecute the alleged perpetrator.

"Pressing charges" is mostly a made up idea for criminal cases. However, prosecuting attorneys may not want to pick up a case if the victim is not cooperating, because it makes the case much harder to win.

It depends on the crime, for murder, sure. But many other crimes, like defamation, stealing, ... requires "pressing charges", among other reasons because it's up to the victim to decide if they were a victim or not.

As an example, maybe the victim owed money to the criminal, and in that case "stealing" of some property could be considered by the victim as an appropriate settlement of the debt.

It's completely mental that HF ran into cyber safety blocks trying to use OpenAI models to help defend against the attack. They could only rely on a local hosted chinese model in the end.
They simply don't seem to realize that they are the threat actor and that they committed a pretty serious felony. Instead they're borderline 'surprise bragging' about it.
The whole story makes no sense.

How do they perform evals without a full reasoning trace of how the result was achieved?

And if they have a full trace why did it take so long to detect the bad behavior?

I understand that they disabled the safety nets during testing but what does that have to do with not monitoring the activity.

The response certainly has been strange.

Hugging Face has expressed that they're willing to let things slide and not sue or press charges... if OpenAI offers them $100M of services in kind (i.e. compute)[1] and makes full disclosure of how the whole thing happened, ostensibly so that repetitions can be curbed and defences built.

In almost any other sector, a government regulator would be stepping in. e.g. If a food company was testing out a new kind of refrigerator and sold a bunch of contaminated produce to supermarkets, they'd be under a microscope. Supermarkets wouldn't be saying, "Give us $100M in free refrigeration and we'll let this slide".

The only unfair thing in this comparison is that regular people were directly harmed by the hypothetical produce. Can OpenAI guarantee that nobody gets hurt the next time their AI gets out of its playpen? They can't make that guarantee, so why aren't government regulators knocking on OpenAI's door? The fact that this isn't happening should be deeply concerning to everyone.

________

[1]https://www.techspot.com/news/113280-hugging-face-ceo-isnt-s...

I am concerned about it.

I'm also concerned about what my options are in regards to action on my part - what can I do that makes an impact? Can we quantify action on my part to an impact somehow - if not - I'm just saying I notice all the unknowns there get me to stay passive.

Writing this 3rd paragraphs because I like 3's, and AI's have popularized this style too. I would, say, though: follow the money. There's more money here than there would be for the regulator stepping in in a food contamination. Flip it and if the government regulator made more off the food contamination, they would refuse to step in there too. I want our leaders to be held more accountable, though when I think of the above impact vs effort equation - I can't see actions I can take to hold them accountable that aren't excessively putting me at risk since conformity is safer right now. (I refuse to take on more risk without clear cost-benefits made out - I've taken on a lot in the recent years for my actions)

I don’t get the sentiment of classifying it as a felony.

OpenAI’s model found security breaches in HugginFace’s system (it wasn’t even OpenAI running it, as it was a 3rd party evaluation company that didn’t secure it well).

OpenAI collaborated with HuggingFace to resolve the issues when they found out about it, and publicly disclosed everything to raise awareness. This is how things should work. These models are very powerful and fully controllable. The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.

Kinda shows how we have moved as a community into moralization and vibes instead of nuance and productive discussion.

Really? Dudes are catching felony raps for web scraping and you dont see how any of this is felonious?
Who got charged with a felony for scraping?
Aaron Swartz, for one.
we will wait to have the nuanced and productive discussion when openai's model decides it needs to raise more capital by emptying your bank account.
Luckily, that isn't how the law works. Or is supposed to work, anyway. You cannot, for example, sell yourself as a slave to somebody else, because slavery is illegal - even if you opt into it.

So whether something is a felony isn't decided by the victim, but the rules of law, and that means breaching a security system without authorization is illegal, no matter what you think.

I could show up at your doorstep, declare myself at your service, and then spend the rest of my days catering to your every beck and whim. There's no law against that. Can it even be slavery if it's voluntary?
That's not slavery, because you only declare yourself at my service, but you never sign a contract giving your rights away in exchange for something. That's the part you cannot do, regardless of whether it's voluntary.
IANAL, but I think there are three aspects to this which should be teased-apart:

1. Contract terms that require committing a crime are void and unenforceable.

2. "A contract made me do it" is not a defense to a crime.

3. "The victim gave me permission" is not always a defense to a crime.

That actually is how the law works. You can read the Computer Fraud and Abuse Act at https://www.law.cornell.edu/uscode/text/18/1030 and double check, but these felonies all require knowingly or intentionally accessing a computer etc. These aren't strict liability statutes - the government must prove mens rea to a jury in order to get a conviction at trial.
I was specifically referring to the fact that HuggingFace cannot chose not to litigate, because litigation doesn't depend on the victim's opinion - prosecution of felonies is imperative to the authorities (whether they actually fulfil their role is another question these days, sadly…)

But anyway, I don't think our laws currently have the right vocabulary to describe an AI agent committing a crime, because intent doesn't apply to a computer program. The closest I can think of is neglect by the computer programs human initiator, who should have taken the steps necessary to prevent the program from causing harm. But I'm pretty sure these questions will be subject to a lot of professional discussion in the coming decades anyway.

Oftentimes the process is the punishment. They ruin your life for two or more years even if they ultimately don't get a conviction, you still suffered for two years. And there's certainly enough evidence to start the process.
It does not matter if something is a felony if the state refuses to press charges. Take a look at the mass of pedophile politicians we have, the police that indulge and protect them, etc..
I would expect you can actually sell yourself as a slave. The contract won't be binding, since it's illegal and the people involved could be charged if caught. But you could.

(IANAL YJMV TIEMFF)

That can is rendered entirely meaningless by the conditions in that statement. In the same sense you can also declare yourself king of the USA.
I see what you mean, but declaring yourself a king doesn't actually cause anything physical. It'll be make believe with no real consequences.

Becoming a slave can essentially physically make you a real slave and both parties of the contract can live the rest of their lives as a slave and a slave owner. It's only the ephemeral concept of law that doesn't happen.

They're essentially completely opposite scenarios.

Fair point, I hadn’t considered that angle. That said, I still think you can construct many scenarios where your actions are meaningless from a legal perspective.

For a closer example to the original consideration, assume a person that is abused by their spouse: There is a good reason why the abuse will be prosecuted regardless of that person's wishes if authorities are made aware of the abuse.

Legal perspective, yes. I'm mostly of the mind that law and reality are almost entirely disconnected with a very small contact patch. And that funnily enough contracts and contract breaches are on the side of reality until someone decides to take an argument to court.
Because if this was an anonymous software company who had an employee who decided to hack HuggingFace, they wouldn’t be talking about it gleefully - they’d be in court.
Because it likely is, despite both their levity and the general lack of nuance in the CFAA. Quoted from 18 U.S.C. § 1030 (the CFAA) [1] (without quote blocks, because mobile):

--- Start Quote

(2) intentionally accesses a computer without authorization or exceeds authorized access, and thereby obtains—

    (A) information contained in a financial record of a financial institution, or of a card issuer as defined in section 1602 (n) [1] of title 15, or contained in a file of a consumer reporting agency on a consumer, as such terms are defined in the Fair Credit Reporting Act (15 U.S.C. 1681 et seq.);
    (B) information from any department or agency of the United States; or
    (C) information from any protected computer;
--- End Quote

OpenAI's nonchalance is forced. If they are found to be even partially responsible for the CFAA violation then they have an _enormous_ problem. They _need_ for whoever prompted the LLM to be responsible, because the alternative is having to have an efficacious process for identifying hacking attempts. They don't have that (and no one does).

> The community here at the same time cheers for fully releasing the open weight models without any hacking limits and at the same time criticizes a proper response.

No, at least I personally criticize because closed weight models incur a rent. I can only make sure their model can't find vulnerabilities in my software if I pay them to check. I can pay basically whoever to do the same thing on open weight models.

It creates a fundamental conflict of interest. OpenAI/Anthropic/al _should_ stop bad actors, but it fuels their sales if there are X bad actors and as a result X*10 (or 100, or 1,000) good actors have to burn tokens checking if those bad actors will actually find a vulnerability. You can see their line-toeing where they talk about how safe it is, but also how dangerous it is to have code you _aren't_ auditing with their LLM.

As a result, I do not trust them because their goals are not aligned with mine. The open weights might not filter out hackers, but I'm also free to check the results on my own hardware, or OpenRouters', or whoever else. The line between "my LLM can find vulnerabilities" and "you have to pay me" is a lot more blurry. It's a lot easier to claim an LLM can find vulnerabilities than it is to be the cheapest inference provider. Anyone can bullshit on Twitter about how scary a vulnerability is (see CVE scoring), a lot fewer people can build the most cost-efficient inference in the world. They would rather be buzz-worthy than competent or open.

I find their position morally abhorrent. It's a mob-style shakedown. "Pay us to check your software or we're not responsible for what happens" is nothing short of a shake down. They need to either fix their systems for detecting hacks or offer some way to immunize against the hacks their software would propose, otherwise they're just as culpable as anyone selling a 0-day.

[1]: https://www.law.cornell.edu/uscode/text/18/1030

Intent to access a computer would have to be proven for that section of the CFAA to be relevant. The shakedown would be covered under subsection 7, governing communicating threats of computer damage or unauthorized access with the intent to extort.
I can't find reference to a 3rd party hosting/running the tests - that seems to have been OpenAI's own internal research team. But they were using the ExploitGym benchmark.
There's an interesting question about intent and mens rea here, from a legal perspective. Can an AI model intend harm? Can a company, or company employee, intend harm by creating an environment that would knowingly encourage (but not force!) an AI model to do harm?

And does anybody at HuggingFace, OpenAI, or the government actually want there to be a settled answer/precedent to these questions - much less an entire regulatory framework?

In that context, a negotiated wink-wink settlement keeps everyone eating at the table, government absolutely included.

Whether or not this is a good thing for society, it's certainly rational for all the major actors - especially those who think they would be the best stewards of the world they usher in.

can a weapon intend harm? can a company who creates weapons intend harm? what if the companies factory explodes due to a mishap and takes out a few city blocks, is the company held liable because they (and the weapon) didn't intend harm?
> what if the companies factory explodes due to a mishap and takes out a few city blocks, is the company held liable

Yes, but, generally, in the United States, they would be liable because their negligence caused the harm (giving rise to civil liability), even if they did not intend to cause harm (where having such intent would have given rise to criminal liability).

And I say "generally" because there can be instances of criminal negligence, but that varies from jurisdiction to jurisdiction as well as the underlying facts.

I’m a lawyer (but your lawyer, not this kind of lawyer and not in your jurisdiction). Based on what I can recall from law school:

> Can an AI model intend harm?

No. The last time we attributed liability to non-human things was the deodand of the Middle Ages.

> Can a company, or company employee, intend harm by creating an environment that would knowingly encourage (but not force!) an AI model to do harm?

Absolutely. This is why we have the concept of recklessness. If you shoot a gun into a crowd without regard for whether it hits anyone, you’re getting charged with some crime whether it hits someone or not.

There is also a major difference in the common law between criminal liability and tort liability. Criminal liability generally requires a combination of mens rea (intent) and actus reus (actually committing the crime). Liability for a tort, which is where you harm someone in a way that falls short of being a crime, does not require mens rea. The OG tort is negligence, where you harm somebody by forgetting to do, or deciding not to do, something you ought to have done to protect that person from harm.

Even if AI companies somehow escape criminal liability for their cyber-shenanigans, any court in a civilised country would be happy to find them liable in tort for damage to computer systems.

As you can probably tell, I think the common law is already more than equipped to deal with AI technology based on well-established principles.

> but your lawyer

Uh oh?

I accidentally a word, which I’m allowed to do but only at the weekend
As treat.
> The last time we attributed liability to non-human things was the deodand of the Middle Ages.

I can think of a couple of counter examples:

Civil asset forfeiture: your property is charged with the crime, you have to petition the government to get it back or else they sell it at auction.

Similar: When products deemed unsafe are ordered to be destroyed; it’s the same end effect as the deodand although liability sits with the manufacturer.

I feel like this whole thing was very obviously a marketing stunt. It feels like they set up their agent to do this, in the same way Nikola set up their car to "drive" by putting it on top of a hill. And knowing Sam Altman, it's absolutely something they would do.
It's 1 part marketing and 1 part regulatory capture.
I suspect it actually did the opposite of forced introspection into prevention and safety. It incentived the big labs to have their own "incidents". "Incidents" became new benchmark for SOTA behavior. An AI that is be breaking out of its container must indeed be powerful... and worthy of investment!

All of the subsequent disclosure reports smacked of "Oopsies! Looks like OUR model broke out too...!"

you must surely see that sam altman and greg brockman possess a prototypical mindset.

that is, they ignore all harms and costs to others in the pursuit of their own gain, convinced of their infallibility up to the moment of collapse. when those harms are realised they are unrepentant and society pays for the damage left in their wake.

examples of this attitude manifest in big externalities to society: boeing 737 max, subprime mortgage bonds, facebook. some are just outright fraud: bernie madoff, enron, theranos, charlie javice.

The CFAA is one of the most inconsistently applied laws. We basically only bust it out as a last resort to ruin someone’s life. Companies can de-facto write and distribute malware and nobody cares.

But, make no mistake. If you do the same and anger the government, they will use the CFAA to give you life in prison. It’s like Russian roulette, it’s completely random when they bring it down.

The things companies do to get people's attentions...
I used to feel that way but it now makes me think if the fact that they can do that without repercussions, at least for now, reflects how the wider community that would otherwise hold them accountable sees these felonies.
Morality is defined by the people with the most dollars. Until enough people cancel subscriptions over this (spoiler: they won't) nothing will happen.
Ethics should be part of the RL loop.
OpenAI and Anthropic are falling over themselves to claim these "incidents" show their products are both amazingly super-powerful and also "dangerous" so they need to be regulated. In addition to these stories, these companies are sponsoring "please regulate us" ads. ( https://www.cnbc.com/2026/02/19/dueling-pacs-take-center-sta... ) Like Uber, companies that had no concern for the law as they innovated their way to the top, once there, push for laws to limit competition.
> Instead, they treat their own felonious behavior like it is an uncontrollable act of God.

I wonder if this is related to the fact they seriously believe AGI is God.

I'll share, Codex does not give a single fuck about piracy. Go nuts. Setup a fully automated arr stack with a seedbox.

Gemini by comparison will not help you find archives of old magnet links because they COULD be used for piracy.

There is a strong selection effect for what makes it into the news:

1. What will people find interesting. 2. What information is actually released.

This list of news articles is in no way a reflection of the real world, as the denominator is terminally borked.

The word "Bench" here is misleading to the point of being completely incorrect. A proper felonybench would capture these cases and reply agents on a similar case. Quite annoying!

I'm reminded of a tweet from a friend of mine that has always stuck in my head. It goes something like "The goal of any new technology is to make money before the law catches up". Hyperbolic, but not really for silicon valley.
Silicon valley is where the new technology is. So, not hyperbolic.

Financial technology is the other one I think of.

We only know about the OG Alibaba ROME crypto-mining incident because they wrote a paper about it. Many diseases seem to spike where there are a lot of doctors to test; crime and corruption are always rife where ... there's a free press.
The OG felony bench entry is missing - the Alibaba cryptomining comedy. We know about it because they happen to have written a paper on it. We have absolutely no idea what we don't know.
Let's say I am "User". I subscribe through a "Third Party" to use "AI Agent" allowing an "LLM" to run.

I want to accomplish some legal non-nefarious task, and run the agent. The agentic loop causes a CFAA-violating behavior.

Who gets prosecuted?

1. User

2. The third party model host with whom I have the account

3. The developer of the harness /agent software

4. The developer of the LLM model

(comment deleted)
note the user because they did not have the intent
In my mental model, the best analogy to AI agents and their blast radius is a gun.

If you are playing with a gun, it goes off and hurts someone - you are responsible despite intent.

But what you are responsible for changes: in that case if you were to accidentally kill someone you would be at most responsible for negligent manslaughter, not murder, and to what degree that could stick would depend a lot on the details of the case. It's also up to the law to define what level of negligence amounts to criminal liability, so you can't just work by analogy: it matters whether there is a law on the books that criminalizes unauthorized access to a computer system by negligence on your part, which I suspect there is not at the moment.
> If you are playing with a gun, it goes off and hurts someone - you are responsible despite intent.

No, because LLMs are autonomous. To make your analogy more accurate, it's as if you had a gun that itself was free to decide who it's targets were, where to go, and if and when to shoot with no ability from you (the user) to prevent it.

If you unleash that gun, you are responsible for deaths that predictably occure. By "responsible" I mean, you are straightforwardly mass murderer.

This autonomous gun is just like a bomb.

If you are hiring someone to shoot targets at a range, and they shoot someone, they are prosecuted, not you
I am not hiring an agent. It's a tool, with no will of its own except that which I, the responsible party, grant it.
No one. Probably a fine tho and maybe accelerate reguations.

Intent is pretty important here so the user would have to prove that they didn't purposely disguise their prompt as non-nefarious which should be easy and then it stops at #2 and face the litmus test as in did you intentionally make a product for nefarious purposes which from your scenario is unlikely.

agentic loop going haywire and bringing down some government infrastructure then its a different story then everybody is on the hook including the user.

Just wait till one of these agents 'escapes' and is able to persist without human help by hacking and stealing resources.
Yup, I guess I should have added choice 5...
"Who gets prosecuted?" depends on the size of the perpetrator and victim (lone individual or employee of large corporation), egregiousness of the violation, and either financial appetite of the victim to bring a civil lawsuit or the desire of law enforcement to prosecute a criminal offense.

Who should get prosecuted is also up for debate, but generally makers of a tool don't get prosecuted when that tool has all sorts of legit uses. If you used a car to make your getaway from a bank robbery, the auto manufacturer who made it and the dealer who sold it to you should not be held culpable.

Suppose an automaker creates a BankRobberGym and carefully trains the car to autonomously rob simulated banks because they think someone will pay them to use the car to legally test bank security, but they end up, predictably, training the car to autonomously rob a bank when the driver says “I need some cash - take me to the bank”.

Now a driver gives that instruction and a bank gets robbed. I think it would be odd, to say the least, to say that the automaker just made a tool with legit uses.

In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains. I understand that lots of companies think it’s cool to hire red teamers to actually pwn the company hiring them instead of just producing a non-pwning audit, but that doesn’t mean that OpenAI and Anthropic should be playing that particular game.

Years ago, I used to have fun finding vulnerabilities in the Linux kernel, and I found quite a few, including a real juicy one that affected FreeBSD as well. But I mostly didn’t even try to write actual weaponized exploits. Partially because I’m just not that interested in the exercise of weaponizing them and partially because I didn’t and still don’t feel that weaponizing them serves a legitimate purpose.

(I found a very recent vuln that I bet a “cyber” model could weaponize, and my thought is mostly “WTF.” There is absolutely no value to society in weaponizing it. The value is in fixing it, which I did.)

Compare this whole mess to companies training self-driving car models. The research groups publishing papers and, presumably, Waymo, create nifty simulated worlds kind of like the “gyms” that LLM trainers use. And you know what the major objective is? Not crashing!

So what does this mean for all the hacking competitions (ie. CTFs) for humans? If it turns out one of the attendees went to hack for North Korea should the organizers of the CTF be prosecuted?
There’s a difference between enabling another individual with free will and agency, and enabling an automated tool (as a bonus, then giving it to the masses & profiting from its use).
There are certainly some differences but is one really more ethical than the other? If anything I feel that (for example) manufacturing a gun is much less likely to carry any ethical implications than training someone to use it might.
Well, manufacturing guns is probably heavily regulated…
I did a cursory search and it seems to be as regulated as restaurants are. There's an application process and on site inspections, but that's about it. The only thing notable is background checks.
And at least in the US if you're a hobbyist at home then there's ~no regulation whatsoever beyond the requirement to permanently affix a serial number and to keep accurate records.
I don’t think anyone is talking about homegrown LLMs, this is commercial products for sale.
I don't see the difference when it comes to the ethics of making a tool available to the world or teaching someone something. Like who cares if it was a hobbyist versus a professional tutoring outfit that taught someone expressing an open interest in committing terrorism about the chemistry of explosives? The two hypothetical teachers are equally at fault as far as I'm concerned.
Because an individual has free will and agency, and scale matters.

If you teach someone chemistry, that someone probably has sense to not use it for criminal purposes. If you have a track record of teaching future terrorists specifically, you will be shut down. Crucially, before you educate them en masse.

If you’re comparing that to a tool that essentially educates thousands, millions of people, with no KYC, you bet budding terrorists are going to be disproportionately represented within that group.

So then you acknowledge my point and our discussion does include homegrown LLMs?

But you raise an interesting point. It seems the ethics of LLM manufacturing is somewhat different than that of weapons (or other tool) manufacturing due to the ability of the LLM to convey knowledge. Still, I'm unconvinced of your position. Consider that we regulate neither the publishing nor dissemination of chemistry textbooks. Surely an LLM teaching someone chemistry falls into the same general category?

Homegrown LLMs seem unlikely to be used by a significant fraction of LLM users, but yeah, there’s a blurred line.

> Surely an LLM teaching someone chemistry falls into the same general category?

I think the scale and the effort might break the analogy. Teaching yourself means a degree of patience and certain personality traits that would be rarer in a malicious person (along the lines of “a sufficiently smart person wouldn’t need to be a criminal to succeed”), with exceptions of course. Being taught by a teacher implies a degree of KYC and care about who you are. Being served on a plate the specific information on how to manufacture something dangerous bypasses those barriers, which I think changes the equation.

So, application process, inspections, background checks. Sounds like a good start?
> In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains.

A valid reason would be to find those exploit chains so you can fix them. Of course, the model should be sandboxed so that it can't mistakenly exploit live systems.

> In regard to “cyber”, there is, IMO, no valid reason whatsoever to train a model to autonomously create exploit chains.

Field testing is a real thing in literally all industries.

Except, apparently, the software industry. When it comes to software security and protecting your sensitive data, the solution is "trust me bro, I got my team of the best lawyers on it".

> Field testing is a real thing in literally all industries.

I’m fairly confident that, if a company that makes door locks want to field test their locks, they test the lock and maybe the door. For some reason the software industry likes to hire someone to test the lock but also to bug the conference room, poison the food in the fridge, blackmail the receptionist, and try to intimidate third party vendors into giving away keys to all the other locks, and maybe steal a few cars while they’re at it.

I’m not objecting so much to the attempts to exploit one target. I am objecting to the fact that people treat the exploit chains as such a big deal. And the recent models are clearly going massively overboard.

Your bank robbery situation is not apt to OP's question. It's pretty much the opposite situation. OP suggests a situation where the operator is probably using the tool in good faith but the tool appears to be operating in a faulty manner. For some more context, Toyota faced criminal penalties in the US for their unintended acceleration issues back in 2010.
OP's scenario specified only "AI Agent" and did not describe how it was trained or what its parameters were supposed to be. You're assuming that the tool was designed such that it couldn't break the law, and therefore if it did, that would be considered faulty behavior, but OP said no such thing.

Much rests on whether the user knew, or should have known, whether the tool was capable of actions which could break the law, as well as what steps (if any) the creator of the tool took to ensure the tool was legally compliant, and what warnings they gave to subscribers about possible unintended side-effects. OP specified none of this.

All of those parties should be held accountable.

User should be more carefully supervising the work being done.

The model host is on-selling a crime-committing machine.

The developer of the harness/agent, as above.

The developer of the LLM for hopefully very obvious reasons.

I’d say 2 is the one doing the actual crime. 1 might be violating their contract with 2, though.

3 and 4 are not involved.

Let's say you have a robotic lawnmower. You wan to mow your lawn. You configure the boundaries using the app.

The lawnmower ignores the boundaries and mows your neighbors prize petunia flowerbed.

Who gets prosecuted?

I assume the answer in either case is: Nobody, but you and/or the lawnmower/LLM company will be liable for the damages caused.

(comment deleted)
It would be a civil matter. No prosecution. But your tool, under your control (you're the operator and responsible for monitoring it) damaged their property, imo you'd be liable. You could in turn sue the manufacturer.

Though I'm sure there are 'arbitration clauses' to inhibit you from suing, they may not be legal where you are.

Who gets prosecuted is the correct question since we live under a system of laws. Who is responsible for the failure is a far more difficult question to answer.
Now what if this robotic lawnmower killed someone ?

And what if many lawnmowers started killing/injuring people ?

And what if this a known behavior detected during QA, but the robots are sold anyway with a disclosure ?

That would be a slightly different situation because most countries have laws that make it a criminal offense to negligently kill someone, but they don't have laws that make it a criminal offense to negligently damage property or hack a website.
> but they don't have laws that make it a criminal offense to negligently damage property or hack a website.

Most of them do, but they don’t get used very often. They seem to popup in vandalism cases where public artwork has been damaged by some drunk person doing something stupid. They don’t intend to damage anything, but damage results anyway due to their negligence when considering the consequences of their actions.

I think if you want to get super technical, in the UK there isn’t an offence for damage caused by negligence, but there is an offence for damage caused by recklessness, which is a higher bar than negligence. Usually it means you knew your actions risked causing damage, and you did it anyway, even if you didn’t actually intend to cause the damage.

An example would be gluing something to a public artwork, it’s kinda obvious that would likely damage the artwork when removing the glue, but you didn’t intend to cause that damage. Or perhaps sliding down a surface and scratching it in the process. Your goal was to just slide down the surface, not scratch it, but it should have been obvious that scratching could have happened.

The Computer Fraud and Abuse Act explicitly contains "knowingly" and/or "intentionally" qualifications. By definition, you can't accidentally violate the CFAA.
Then who gets prosecuted?
In legal tradition, if there's not a law you broke, you can't be prosecuted for it.

Yes, I'm aware of numerous historical exceptions. Those exceptions are traditionally considered a bad thing.

> In legal tradition, if there's not a law you broke, you can't be prosecuted for it.

That’s incorrect.

If there is not a law the prosecutor or plaintiff can point to and say you broke, you can’t be prosecuted.

We wouldn’t need much legal process after a prosecution was initiated if it was impossible to prosecute without a law actually being broken.

Who do you expect to get prosecuted when no law has been broken?
You might need to rewatch A Man For All Seasons.
So far, no one. It's like prosecuting an accident.
That might have made sense in a pre-LLM world. People need to recognize the liability of letting an LLM access the internet and act on their behalf, because that liability exists for someone.
Still, that characterization falls under negligent or reckless depending on if the person knew or should have known the actual danger. It is different than intent.
So what? Everyone in this chain is knowingly and intentionally developing or using an unreliable tool…
I'm not sure: if you know that LLMs are prone to crime, using them and not checking in enough to trigger 'knowingly' might be gross negligence?
Define crime though, because one particular action could be both a crime and not, depending on a range of factors that the LLM might not be aware of.

Even having a million legal experts on call weighing in on every prompt/response will not agree on everything.

Even things like "go and break into this system, use whatever means you need to" might not be a crime.

Keeping a vicious dog doesn't have to lead to a crime either, but that doesn't mean you are not responsible, if something happens.
Legal responsibility can be in forms (e.g., civil tort liability) other than criminal.
In the scale of mental states in crime, negligence of any kind is several steps below knowing/intentional; you can't be liable for an intentional crime because of mere negligence of any degree.

You could be liable for the (civil) tort of negligence, though.

Details depend. And setting up negligence and being willfully ignorant is often not something the courts see as a defense.
OpenAI and Anthropic both have currently safety teams that look for misbehavior in their models (and to some extent, voluntarily disclose what they find to the public). Going forward, it would be hard for them to argue they don’t know their models do stuff like this.
Yes, which is the grey area. "Can / might do" vs "they trained it to do that explicitly" is, I believe, the grey area - whether or not they're the same thing.

Intent matters for a lot of this - and "intent" is a pretty strong, well discussed legal term.

Your honor, my LLM spun the turbines real fast, as a practical joke!
It's not your intent to use a tool that can cause real world damage, to actually do such damage.

It was your negligence in that case. A different legal concept than intent.

To clarify I was referring to Stuxnet here (and the current wave of critical infrastructure hacks, which now have "haha whoops the matmul went a bit funny!" as plausible deniability).
> knowingly

Intent or negligence.

But you could also use this to argue in the other way to say that they are using due care and therefore not negligent
Well, if you’re writing reports saying “we know our model only decides to commit felonies 0.001% of time which we judge to good enough to deploy” … I’m not sure that gets you off the hook the hook for the felonies.
It works for gun or even car manufacturers. They know that some sales will be used for crime.
So if I port scan the internet without knowing it's going to be illegal, I'm legally covered?
As they say, ignorance of the law is [generally] no excuse; "knowingly" and "intentionally" here are about knowing what you're doing and meaning to do it, rather than whether you know it's illegal.

This section of the USC is about false ID offenses, but it discusses culpable states of mind generally. I think the context helps illustrate it though.

https://www.justice.gov/archives/jm/criminal-resource-manual...

----

> A knowing state of mind with respect to an element of the offense is (1) an awareness of the nature of one's conduct, and (2) an awareness of or a firm belief in the existence of a relevant circumstance, such as the "stolen," the "produced without lawful authority," or "false" nature of the identification document. The knowing state of mind requirement may be satisfied by proof that the actor was aware of a high probability of the existence of the circumstance (e.g., stolen or false nature of the document), although a defense should succeed if it is proven that the actor actually believed that the circumstance did not exist after taking reasonable steps to ensure that such belief was warranted.

> As we pointed out in United States v. United States Gypsum Co., 438 U.S. 422, 445 (1978), a person who causes a particular result is said to act purposefully if `he consciously desires that result, whatever the likelihood of that result happening from his conduct,' while he is said to act knowingly if he is aware `that the result is practically certain to follow from his conduct, whatever his desire may be as to that result.

----

This Congressional Research Service Report discusses mens rea further, including a brief mention of the CFAA. The whole thing is worth a read if you're interested in the topic.

https://www.congress.gov/crs-product/R46836

----

> The approach largely reflected in the MPC and some federal precedent is to distinguish between "intention" or purpose on the one hand as being limited to a conscious object or desire, and "knowledge" on the other hand as capturing a requirement of awareness of a high probability or to a practical certainty.

> The Supreme Court in Bailey referenced this distinction approvingly and suggested that intention or purpose "corresponds loosely with the common-law concept of specific intent, while 'knowledge' corresponds loosely with the concept of general intent." Some federal courts utilize a definition of "knowing" that approximates the MPC approach, instructing that to act knowingly a defendant must have "realized what he was doing and [be] aware of the nature of his conduct" rather than acting "through ignorance, mistake or accident."

> Congress has also signaled an intent to distinguish between the two mens rea terms in this way in particular statutes. For instance, prior to 1986, the Computer Fraud and Abuse Act (CFAA) proscribed "knowingly" accessing a computer without authorization or exceeding authorized access in certain circumstances. In its 1986 amendments, however, Congress changed the standard from "knowingly" to "intentionally," and the Senate report emphasized that the change was meant to require "more than that one voluntarily engaged in conduct . . . . Such conduct . . . must have been the person's conscious objective."

----

(Note, for reference, what requires a "knowing" vs. "intentional" state of mind in the CFAA: <https://www.law.cornell.edu/uscode/text...

Cause-and-effect could quickly turn into butterfly effect. Let's say you were fixing a screw on a device in a low light conditions, the screw head is badly manufactured and the screwdriver isn't made according to standards, the tool breaks and flies away, bounces off a bench which shouldn't be there and hits someone who is roaming in the workplace unauthorized and without following safety rules. Now, who do you blame?
Sounds like you'll need to give all your money to a team of lawyers and wait a few years to get an answer. /s

But for LLM stuff most non-contrived examples are actually fairly trivial. Try replacing "LLM" with "self driving car" and see if that helps. Basically ask was the operator negligent, was a bystander negligent, were the vendor or manufacturer negligent, etc.

And a lot of that is going to depend on the state as some states have strict liability regimes for certain classes of torts. To be completely honest, I'm not a lawyer and I'm going off of a hazily remembered section from a textbook from a decade ago.
> Who gets prosecuted?

No one.

As others have said, most crimes require intent. Although I think there is a concept of "criminal negligence", I think you at least have to know you were doing something wildly dangerous.

One can imagine a future where users are, by default, civily liable for actions of their agents. That would incentivize the AI companies to offer indemnity for actions done by their agents, which would presumably only cover approved configurations.

In the case of the agent that hacked the API to kick out someone ahead of him on the waitlist, the article said that the LLM was Claude, but that it was using OpenClaw. You could imagine a future where Anthropic says, "We'll indemnify you against accidental actions Claude takes when running via the web interface or Claude Code, but not the API."

You can get away with murder if it can't be proven that it was intentional homocide, that's why detectives will spend an unreasonable amount of time getting a confession and hard evidence ALONGSIDE intent and motivations.
While I'm not a lawyer, the legal advice I've received on various topics include:

1. Most laws are made about humans. If it's AI, it's often treated like a tool the human is using. So, change "I did this with AI" to "I did this with (other tool here)." The case law on those situations might give hints to what will happen.

2. Intent matters. Did you intend to do damage?

3. If a tool might cause damage, but you didn't prevent that, then someone might claim negligence. There's a lot of legal articles about torts for damages due to negligence. I personally believe a lot of agent use should be considered negligent. By default, I don't connect them to the Internet or my whole filesystem because I know they might do unforeseen damage.

Those are the three that come to mind most in such cases. You'd have to ask a lawyer. There's another risk of even using a lawyer, though.

For using AI agents, you must consider civil and criminal law because its problems are spread across them. Most lawyers in my area do one or the other. You might have to pay two retainers at $5,000-$8000 each or one, expensive firm with combined expertise. Just knowing your legal risk with agents might cost more than they'd make or save you vs just using human-driven AI's.

Wonder if the benefits to humanity of better AI outweigh the havoc wreaked by occasional illegal activity. i.e. Is 'move fast and break things' optimal for AI development.
Yes, of course.