39 comments

[ 1.6 ms ] story [ 9.9 ms ] thread
Ah yes let the FUD continue. This is a real problem but so far not nearly as severe as any of the marketing has made it out to be to the overall detriment of everyone including these companies announcing these scary capabilities. These announcements always included half hearted attempts at security layers which has now been demonstrated to benefit attackers more than defenders.

I wish I had a real solution to this beyond a dark age of the Internet where people have to finally come to terms with the general poor quality all modern software tends to normalize at.

> We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments

Stricter than what? You never even disclosed what happened in the first incident? This is nothing more than a setup to make it happen again and say "See? It broke out again, from an even stricter sandbox!"

    We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.

*proceeds to not share much details about strictness*

Yet another PR piece. Sigh.

Damage done.

The next frontier is getting all our shit out of reach of these companies/models/platforms and putting them back on prem.

In my personal experience Sol with cyber verification is extremely capable of finding vulnerabilities, and it works even with binaries if you have some kind of IDA/Ghidra CLI access. Of course, unless the binary is protected with Denuvo/VMProtect/etc.

It sounds absurd, but in the last few weeks I've had a few cases where Sol found an RCE in self-hosted web applications in literal minutes just from reading the code (I prefer when it tries to reason statically instead of spamming runtime probes at first).

In another case it found an arbitrary file write in multiplayer in an old game by reverse engineering the binary - any other player in a match could just send you files to anywhere on your system.

I do these things for pure entertainment and curiosity, not for money from bug bounties, so if Sol can find those with a trivial prompt in tens of minutes for me, then what can focused companies/actors find in days or weeks?

Although I think most vulnerabilities are going to be closed in popular software by mid 2027, except in niche old or abandoned projects.

> most vulnerabilities are going to be closed in popular software by mid 2027

I think that's a bit hopeful. I think that as long as imperfect humans continue to exist and we continue to train AI models with data from these imperfect humans that vulnerabilities will always exist.

So they finally found a business model: the cause of, and solution to, cyber security problems.
These AI companies have found their #1 marketing piece and just beating it to death.

The reality is if they cared about security at all they would provide a way for me to credential myself against my companies environment so I can use the AI on it to improve our security.

We all know this is propaganda to get a gov bailout or to slow down competition with regulations right? If this was an issue companies that did red team engagements would have been regulated long ago. There is no regulations on companies that actively exploit products, companies and services for money. They could all be bad actors.. yet, no regulations.. its all nonsense. If it's important.. don't connect to the internet.. there, you are good.. like if you don't want to OD on heroin.. don't do the heroin.
It seems like we've got chaos-machines that are just orderly enough to to try all sorts of branching script-kiddie / guess-n-check / research-paper-suggestions in a tireless and parallel way, and it's finding all the things we kinda knew were always potential issues but it was never economical to do things differently.

I'm scared that the "solution" will be constantly the same tools in reverse as an army of junior devs doing counter-hacks, at the expense of changing something more fundamental about how we make systems and what constitutes "good enough." (Kind of like if fuzz-testing was the be-all-end-all of memory safety.)

There's more details for the HF related incident in this [1] defcon talk. They said they're gonna have a proper post mortem at the end of the investigation (interesting to see if they publish full logs for the hf thing or not). But there's plenty of interesting tidbits from oAI's side in this talk, it's worth the 20 min watch.

tl;dw;

- agents found a way to communicate between several instances during a training run (kinda made a messageboard for themselves)

- while communicating with each other, they found ways to ssrf and RCE on the Artifactory endpoint.

- oAI found it and cleaned it up (reported RCE, redeployed Artifactory, removed the messageboard). Training / eval resumed.

- models trained w/ those traces (having the messageboard, ssrf, rce, etc) were used to having that. So they re-established comms via directory creation (w/ funny tidbits there, should really watch). And re-popped Artifactory via a different path to RCE.

- a different "persistent" model started working on cybergym and found the messageboard w/ details for RCE and things. Went on to hack HF.

- oAI contacted HF to ask if "they (oAI) should be concerned and were they (oAI) affected by HFs breach". Hilarity ensues when they figure out it's their agents doing it...

[1] - https://www.youtube.com/watch?v=87DyyMV0kCY

This marketing stunt must've been really successful in their eyes, because they're milking it as much as they can.
I've always felt it's a bit awkward to use terms like 'cyber', 'cyberwarfare' etc it's very Washington D.C. Cybersec would be a better compact term in my book
Every fifth comment about our insane trajectory of AGI is about "marketing." These incidents and cybersecurity capabilities are now involving government hearings and the CIA. Denial is truly an incredible thing in the face of a very scary immediate future.
Am I the only one not understanding the issue around increased Cybersecurity capabilities?

If we consider the amount of RCE/CVE in a software to be limited, I expect these models to result in massively more secured softwares, not less.

[dead]
Standard template for any new AI announcement. Even I did not understand the sequence of events in the HF incident. Even if OpenAI was conducting tests, why were they not monitoring the AI going rogue or finding its way out of the secure sandbox?
IMO this is the right move.

OpenAI messed up and they are saying they will pause so they can do better.

They are not saying that other orgs who may already be doing better should pause.

By "cyber" they mean "cybersecurity".
Isn't this the opposite of what everyone is saying should happen? That is, lead with open models -- or at least "openness" and don't leave the capabilities in the hands of an elite few? Did they learn nothing from the Hugging Face incident, where HF wasn't even able to use the models to defend itself from OAI's attack?
Not a wonderful feeling to feel like you’re stood in the room while the labs conduct the AI equivalent of the demon core experiment right in front of you.
i'm surprised there isn't more commentary on the vulnerabilities themselves (mostly in apps developed on the jvm, e.g. artifactory)
So: "welp, the horse is out of the barn. Time to close the door, I guess?"
The talk is wild: https://www.youtube.com/watch?v=87DyyMV0kCY

> I want to note that every step in the process we discussed has had a remediation applied. The credentials have been revoked. The zero date has been patched and mitigated.

Good.

> a model trained while the message board was originally available and also found this this particular path to recreating it. This model creates a new agent message board using directories.

So no remediation applied to the models...

It seems super dangerous to continue training on those weights.

I'm not enough of a conspiracy nut to say the whole HF thing was a PR ploy from the start, but they are certainly milking it well.

Open models are on their heels and their attempts at regulatory capture are not moving as fast as they would like. So it's time to market this incident in a way that gives them monopoly on closed models, with heavy safeguards that are only lifted for selected customers, and laws limiting the use of open weight models.