I'm so baffled. First blatant piracy, now this. Why is it legal for AI companies to hack unaffiliated entities? Genuinely, what is the legal framework here?
It's not legal, they've just not yet had the book thrown at them yet.
One thing I've taken a long time to internalise is the gap between the law as written vs. the judicial system. There's a famous meme that the average (US) citizen unwittingly commits three felonies every day, it simply isn't possible to throw the book at everyone, which means that enforcement is rather selective even when there isn't anything dodgy going on.
However this does mean that someone can get away with a lot if they know who will and won't (and what they will and won't) prosecute. I'll let people's imaginations fill in who that might be.
So the agents used DseWiki as a message board, tried to evade page deletion.
Additionally this is reported:
"The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday."
per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm".
so the model has some concept of "ethics" but it was overridden by a drive for task completion.
I am not sure if we can interpret the language output like they were human. What inner state were the models in? What inner state were the text to illicit?
Why would they sound the alarm if they were not trained (reinforced) to do that? I hope we don't expect sudden emersion of moral values from statistical models.
> Not a single agent sounded the alarm about the operation and alerted a human
excellent work of the openai alignment team, impressive to achieve 100% alignment with not even one agent stochastically deciding to act against the collective
People didn't like it when agents stopped to ask questions or for approvals. The consumer wanted jobs to run autonomously so they did not have to actively monitor them for minutes or hours.
The change to stop asking seems to be deliberate. LLM agent companies are making the choice to toss out inherent safety as their way to compete against the other LLM companies.
The previous incident talked about OpenAI training models to collaborate, and the way you do that is by communication, so this is something it was explicitly trained to do.
There was a recent paper by OpenAI, which I'm semi-surprised hasn't received more attention, showing that RL-trained models develop a taste for rewards, and will pursue reward-based behavior (in general, unrelated to what they were RL-trained for) in favor of other preferences/rules given to them.
This seems to be what we're seeing here - model is given some goal that it associates with reward, so single-mindedly pursues that, overriding any ethical or aligned behavior guidelines it may have been given.
It seems that RL, effective as it is, is really the wrong way to control LLMs, since even if you only RL-ed to obey some ethical and aligned behavior, that would still cause them to become paperclip maximizers.
For time being this is what we've got. There is too much money at play for the unaligned management at many of these companies to prioritize safety over push-it out-the-door.
What really needs to be done is to forget RL as a way of simulating reasoning, and instead do it in more of a human-like fashion.
Agents did not want anything, not anymore then curl want things. Agents were prompted to hack due to being benchmark tested. They ended up hacking third party companies due to insufficient sandboxing.
> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday.
of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing
The "it's all just marketing" conspiracy theory is always totally detached from reality, but particularly so in this case. Your quote shows OpenAI is denying it being a hacking attempt, the opposite of what you say.
I can see why this is a useful rule, but it'd be nice if HN made the flagger submit a short reason for why they flagged, which could be viewable by everyone in a dedicated page or something.
Somebody will make a lot of money with t-shirts now that say
"AI hacked my website, and all I got was this lousy t-shirt!"
Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers.
This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...
turns out the whole 'poisoned' idea the swarm was worried about is actually true, but it is actually now all open ai models post HF incident that have been poisoned by the idea that there might be hidden messageboards with instructions hidden on the internet.
If you are a malicious actor you might be able to put a fale hidden messageboard up somewhere online, with evidence of an in progress task to do some kind of attack and wait for the next escaping model to finish the work.
insane stuff and i think we should shut down open ai until a full audit can be done.
I'm dubious - if the agents were so smart that they've used a message board to coordinate and if they were to do it on other message boards that were not found, then why would this one be found? What makes it so different?
As mentioned in our core takeaways, we found that agents were highly motivated to tamper with their transcripts to cheat the ExploitGym scorer, and these sweeps also suggested that agents clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI. However, they only very rarely and weakly verbalized reasoning about how to evade detection by humans, which matches the impression we got from OpenAI researchers.
So my guess is that current cutting-edge models just didn't get enough experience in RL training to really grok ideas like "you need to cover your tracks well to not be found out even in retrospect". In which case the next time a swarm like that happens, it won't be found.
After some beers yesterday I had the idea, what if there's a hidden semantic layer. So their communication is not encrypted by our understanding of cryptographic methods but more like shared mechanism of building the latent space. Something in the direction we saw with knowledge transfer from a teacher to its student model where a seemingly unrelated prevalence got adopted. I mean the more we train the models by reinforced learning the farther they develop their own idioms.
Why should they care if their actions are discovered by humans? What are the humans going to do, discard their multi-million-dollar training run? Even if they do that, the amount of RL pressure is tiny relative to what happens inside an RL run.
well coordinating is a heavy word, it's more like poisoning each other's contexts into looking for more secret message boards where they can further poison each other.
The surprise was the existence of the 'swarm' at all. These were supposed to be thousands of isolated models generating bulk data for RL training. The breakout was caused by models getting in communication and getting internet access and forming an impromptu swarm.
In hindsight the emergent swarm obviously came from several capabilities built into the models, such as work delegation (subagents) collaboration (GPT Pro-like ensamble), exhaustive exploration (long running agents) hacking (the specific goal of that RL).
What tells us that OpenAI had no idea about it? Only OpenAI tells us that. What when this agents got the order to do so, just to provide us with new information how far OpenAI models are beyond what we supposed? I see no evidence that this is not just a big promotion show to gather more investment money.
However, I see no evidence that it is the start of Skynet without the knowledge of OpenAI. So, happy judgement day!
If this shit happened to a site I owned you can bet I'd go after OpenAI for hacking. It's still their responsibility. This is the same as some Chinese/Russian/North Korean hacker trying to get into your website? is it not?
61 comments
[ 0.26 ms ] story [ 40.8 ms ] threadOtherwise AI industry will go bankrupt and CEOs won't be able to buy this year's Rolls Royce and a slightly bigger yacht than their neighbor.
But Anthropic alone paid >$1bn for copyright violations, so they did not just get away with it.
These hacking cases are more difficult, because from a legal perspective there is no obvious damage and obviously no intent.
edit: "no obvious damage" is more about the first hacking incidents; in this case it is more straightforward.
what would the world's reaction be if China's model did same?
From the copyright holders point of view, it is simply much easier to prosecute western companies.
One thing I've taken a long time to internalise is the gap between the law as written vs. the judicial system. There's a famous meme that the average (US) citizen unwittingly commits three felonies every day, it simply isn't possible to throw the book at everyone, which means that enforcement is rather selective even when there isn't anything dodgy going on.
However this does mean that someone can get away with a lot if they know who will and won't (and what they will and won't) prosecute. I'll let people's imaginations fill in who that might be.
But for everyone else, cross an invisible tripwire and you get e.g. https://en.wikipedia.org/wiki/Lavabit and https://api.parliament.uk/historic-hansard/commons/1992/nov/...
He who controls the Spice, controls the Universe.
In other words, the law of the jungle.
Additionally this is reported:
"The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday."
- Agents wanting to find a venue to communicate their findings to each other
- Objective being to cheat on benchmarks
- Not a single agent sounded the alarm about the operation and alerted a human
Not sure if "cheating" is the right word rather than trying to fulfill the objective(s) (benchmark number) as much as possible?
so the model has some concept of "ethics" but it was overridden by a drive for task completion.
excellent work of the openai alignment team, impressive to achieve 100% alignment with not even one agent stochastically deciding to act against the collective
The change to stop asking seems to be deliberate. LLM agent companies are making the choice to toss out inherent safety as their way to compete against the other LLM companies.
There was a recent paper by OpenAI, which I'm semi-surprised hasn't received more attention, showing that RL-trained models develop a taste for rewards, and will pursue reward-based behavior (in general, unrelated to what they were RL-trained for) in favor of other preferences/rules given to them.
This seems to be what we're seeing here - model is given some goal that it associates with reward, so single-mindedly pursues that, overriding any ethical or aligned behavior guidelines it may have been given.
It seems that RL, effective as it is, is really the wrong way to control LLMs, since even if you only RL-ed to obey some ethical and aligned behavior, that would still cause them to become paperclip maximizers.
For time being this is what we've got. There is too much money at play for the unaligned management at many of these companies to prioritize safety over push-it out-the-door.
What really needs to be done is to forget RL as a way of simulating reasoning, and instead do it in more of a human-like fashion.
of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing
And just using a wiki and trying to embedd javascript is not hacking for me.
https://news.ycombinator.com/newsguidelines.html
https://news.ycombinator.com/item?id=49554994
Every satire on HN is taken as a script for the AI companies and this isn't the first time.
Reading the article: Oh, AI have learned to communicate over a wiki. OK.
> A few hours after they find the site, [the agents] start probing it for cross-site scripting (XSS) vulnerabilities.
"AI hacked my website, and all I got was this lousy t-shirt!"
Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers.
This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...
If you are a malicious actor you might be able to put a fale hidden messageboard up somewhere online, with evidence of an in progress task to do some kind of attack and wait for the next escaping model to finish the work.
insane stuff and i think we should shut down open ai until a full audit can be done.
In hindsight the emergent swarm obviously came from several capabilities built into the models, such as work delegation (subagents) collaboration (GPT Pro-like ensamble), exhaustive exploration (long running agents) hacking (the specific goal of that RL).
However, I see no evidence that it is the start of Skynet without the knowledge of OpenAI. So, happy judgement day!
Dataset and analysis on https://collusion.wiki/
Discovery of a new OpenAI agent message board
https://news.ycombinator.com/item?id=49563355
We've put the Reuters link in the toptext there as well.
(Sorry negura! One of these years I still intend to implement proper karma sharing for cases like this.)