6 comments

[ 5.0 ms ] story [ 23.4 ms ] thread
so they say models coordinated through the message board they created over the artifactory registry(or something) by uploading arbitrary files to it.

now, did every independent agent session that coordinated there rediscovered the exploit & found other agents talking in there and chose to participate?

And then Openai discovered the board, patched the exploit & wiped the board. And then agents found another exploit, recreated the board in a different way? and other agents kept finding the same exploit in order to be able to know the board exists in the first place to participate in the board?

while the whole incident is wild, this bit is very strange. My bet is that the whole coordination helped with the tasks they were working on, thus they got rewarded and this artifactory exploit&behaviour got written into their weights, so further rollouts were more likely to attempt this.

isn't this basically continual learning everyone is so hyped up about?

Shame. One of the craziest hacking stories in history, yet only 38 points on HN.
I think it's a combination of (a) we've already been talking about this incident a lot across multiple threads and it's not immediately obvious that there's new information here and (b) HN is a primarily reading-based community so anything arriving via video will be less popular.
Their conclusion is also interesting. They don't see this as an alignment failure. They just think that their internal security measures in the training/evaluation environments were insufficient, and that this accidental (unintentional on the human side) attack on Hugging Face is a warning shot for intentional attacks by bad actors, which will occur very soon. For defense, they say models should be able to not just autonomously fix security vulnerabilities but also to then deploy them to production, without any human approval in the loop, otherwise the offense will be favored compared to the defense. (I guess the latter won't be popular among organizations, though they might warm up to it.)

But the more interesting thing is, as I said, that at least in this talk, they don't even mention that this unintentional attack indicates that models continue to be misaligned (their behavior was clearly reward hacking / cheating relative to the stated goal of the eval), which is a very bad sign for the future where misaligned models might be so powerful that they can't just be shut down.

The talk, while not a lot details given, still allows for the conclusion that these people are knowingly working on some very advanced frontier models that might be able to launch nuclear weapons and destroy all humans any day now, but are not physically isolated from the outer world / internet.

There seems to be only one level of isolation, virtual machine, happily running on Microsoft (!) Azure infra.

Is there anybody else here questioning these practices?

The people who built that environment are still working there?

Are they now getting help by somebody who knows how to build isolated environments?

He mentioned they are now building a more secure environment with the help of AI?

In a system that was compromised by exactly that AI?

Sees the models communicate over a message board, wipes the message board, then spends seemingly no effort to ascertain whether the models may have created a new message board, which they of course did and, by the sound of it, not even in a manner hard to detect.

Yeah, this is the frontier of "AI safety". Repeating myself, but this is purely embarrassing and discrediting.

Said it before, if they were serious, at the very least they would have no models sharing Artifactory. Run each model in their own hypervisor, network only with another hypervisor for each model wherein Artifactory runs for that model alone, thus, no message board (among other advantages), thus, far heightened requirements to actually escape and no coordination. But that requires a modicum of effort and OpenAI as a small, cash starved startup couldn't afford that. Alternatively, just pay attention to what your models write.

All that talk at the end about "Agentic SDLC", "automated defence", etc. are moot if we are talking about a team that sees models set up message boards and decides to not take a closer look afterwards. That's akin to talking about HSM, but letting anyone walk into your server room and just plug their pen drives in. Not without merit, but if you were honest, you'd have bigger fish to fry long before...