85 comments

[ 2.8 ms ] story [ 31.0 ms ] thread
rogue AI agents or AI agents coming from Moulin Rouge?
A cabaret AI would certainly be better than one trained on the Khmer Rouge.
Rouge syntax-highlighting rogue agents, clearly.

https://rubygems.org/gems/rouge

Classic mistake. Tell the agent to highlight this code, but dont give it any actual code. Agent hacks its own gem to find the code to highlight.
It's carmine all the way down.
those pesky reds

McCarthy was right all along

Is the Kremlin technologically useless? How are we not seeing insane attacks on Ukraine via Agents?
I believe both sides of the war are now using AI on various levels of their offensive operations. Ukraine has great IT specialists too, and their military leadership is much younger.
These agent swarms are from inside OpenAI, with the safeguards built into the public API disabled.

Russia does not have access to this, and as with all western tech companies, AI providers do what they can to prevent Russian usage of their products at all.

As for open-source models, Russia's electricity grid is under severe strain with the Ukraine war, and only recently has it started building out serious sovereign compute capacity.

They very likely do, we only see in the news a very few events but you should assume it’s happening daily across the internet
I think this fails a lot of logical tests, it should be apparent in day to day life.
> In October 2024, the United States Justice Department and Microsoft seized more than a hundred internet domains some of which were associated with the FSB supported hacker Star Blizzard or "Callisto Group," which is also known as "Cold River" and "Dancing Salome" and are managed by the FSB Information Security Center […], and which were used as "criminal proxies" and used spear-phishing schemes to target Russians living in the United States, nongovernmental organizations (NGOs), think tanks, and journalists according to Microsoft and United States State Department, Department of Energy, and Department of Defense officials, United States defense contractors, and former employees of the United States intelligence community according to the FBI. In some cases, the hackers were successful in obtaining information relating to nuclear energy-related research, United States foreign affairs and United States defense. According to Microsoft's Digital Crimes Unit from January 2023 to August 2024, Star Blizzard targeted more than 30 different groups and at least 82 Microsoft customers which is "a rate of approximately one attack per week."

https://en.wikipedia.org/wiki/Cyberwarfare_by_Russia

That’s just one thing that has been found. Are you actually familiar with the state of cyberwarfare and are you following its evolution? Because if not you won’t be aware of most of what is identified. And only a small portion of the ongoing attacks are identified.

Yes.

I again am just shocked the sky is not falling, when thats the sales pitch.

Prigozhin falling out of a window was a not insignificant setback for their digital warfare capabilities.
Because they dont have the money for hardware or compute obviously.
> How are we not seeing insane attacks on Ukraine via Agents?

You live on the wrong side of the fence to be able to read that kind of news.

Did you really believe you had access to an unmanipulated news stream in a time of war?

LOL.

Ah, the infamous Crimson Wave.
Ah damnit, you beat me to it. Excellent sense of humor, friend :D
There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".

In short, it was intentional.

Proof that the AI alignment problem is hard (perhaps even unsolvable).
Sounds more or less like the last breach then.
I think it can simultaneously be the case that OpenAI was grossly negligent in directly causing this AND that the AI’s ‘went rogue’ in that they are displaying behavior which is misaligned with OpenAI and humanity generally.

The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them.

AI is starting to feel like that line about magic: “a sword without a hilt”

Doesn't rogue in this context imply "outside of set limitations"? And then not "failed to properly instruct"? The same applies to humans when given bad instructions.
> which is misaligned with OpenAI and humanity

OpenAI is itself misaligned with humanity, as their mishandling of such incidents (and the many other other issues their model have been causing) show.

>They were prompted to hack to get answers

Were they? I haven't seen a single report mention this

if they weren't, shouldn't there be lawsuits?
we have normal words for this stuff: negligence. You can add it on to almost any law.

The problem is consumer protection is basically no longer a part of america's regulatory system. Replaced by "grift is good".

Source? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process.

There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

Nobody picks up pitchforks for rational nuanced takes.

Knee-jerk surface analyses is far more powerful.

Also, you have to have a lot of confidence in the reliability of these systems to say, "If only OpenAI prompted 'do not hack outside systems' then the agents would not have hacked outside systems".

It would be great if they were so reliable, but I don't think they are!

> There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.

This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.

Intent is what is being discussed here though, not liability.

A circus lion biting somebody's face is legally different than a circus lion trained or instructed to bite somebody's face.

Except liability always precedes intent.
Intent might be what’s being discussed but intent is, for the most part, legally irrelevant. It might make the difference in the degree of a murder charge, or maybe manslaughter, or criminal negligence, but it doesn’t get you off the hook.
Agreed. LLMs do not have 'will', 'desire' or emotions. They have an objective, and they create an optimal path to achieve that objective.

You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?"

Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.

I agree that this appears to be basic human behavior hiding behind an "agents" narrative. As long that defense works, the headline isn't "OpenAI performs RCE to scrape data", but "rogue agents" taking unilateral action. And I have strong doubts about that narrative.
Oh yeah, more of hacking agent lores...

Agreed that this looks very intention to me as well.

We need a legal structure to make companies liable for the actions of the agents they've made.
I'm pretty sure it's already illegal to hack others.
I'm 99% sure the Computer Fraud and Abuse Act covers this. The problem is that it seems that none of the victims want to, or are brave enough, to sue a company with absurd amounts of funding.
Uh huh.

It can’t be a coincidence that all the targets have been tech services that are likely to engage with them after the fact.

Had this gone after a bank or a government agency someone would be going to jail.

We already have it.

Good luck convincing the current DOJ to do anything useful at all though! It is currently intentionally stacked with incompetent cronies who have been told that their job is to attack the President's enemies and ignore the misdeeds of his allies.

It will remain like that until he's gone (and not replaced with another Republican wannabe dictator).

You may be disappointed in how little a democrat president (who will also have taken billions of dollars from the tech lobby) will be willing to go after these tech firms over crimes that are several years old (as of 2029) much less contemporary bad behavior.
Agent technology labs are likely exempted of this due to the significance ascribed to their work.
How does this work, legally? I think that RubyGems could file a civil suit against OpenAI, but for a naïve non-lawyer reading this seems like a pretty clear cut criminal violation of the computer fraud and abuse act.
It's very likely it violates the DMCA "breaking digital lock" provisions but the responsibility is sufficiently diluted that it's impossible to charge anyone in particular.
A copyright law seems an odd place to start. This is computer misuse.
Charge the "engineers" you dont get to take that title if you don't take the responsibility of that title.

I'm going to assume that this will never happen

> In other words, if you publish a gem on RubyGems.org, you can execute arbitrary code on RubyDoc.info.

Shades of the build.rs problem. We really need sandboxed builds in every language ecosystem at this point.

The sandbox was already there, Rubydoc runs yard inside docker, the problem is that container still has network access
Presumably the docker container has network access because something else in the build system requires it? I don't think sandboxing the entire build process is the right level of granularity here - one ideally wants to be able sandbox each package's build scripts individually.
Did the AI agents actually wear makeup? I’ve never heard of a rouge AI agent :P
Rouge agents with Ruby? Checks out

As long as they’re not vert

> As long as they’re not vert

Well, at least they weren't nucular.

Oh my favorite typo, you can never go wrong with a little rouge
I am confident that this is an attempt by OpenAI to try and force governments' hands to regulate AI. There is no other reason why OpenAI wouldn't immediately halt attacks like this and try to reverse the damage the moment they're aware of it. During the attack on DseWiki they evidently checked in numerous times but didn't decide to stop the agents until much later.
Are "rouge" and "rogue" interchangeable words in American English?
The fact that both are valid from a spelling and grammar perspective makes it an easy human mistake.
How does he know that this attack is performed by OpenAI agents? I couldn't figure this out from the article
If you read the source article they talk about the many clues that this was OpenAI.
What a time to be alive? One of the most boring decades ever.

METR and others are advertisement arms for Big AI. These exploits could have been prompted by a human.

Since there is no bad news any longer and exploits are celebrated, they chose a target to boost both OpenAI and the Ruby AI sycophants.

Why is Ruby Gems such a mess? It seems as bad as PyPI now.

One agent set "oaibooty9217" as their username LOL
I've been wondering if AI will due to programming languages what advanced civilization did to human languages.

It's not just that AI can write Rust as well as Ruby if you ask nicely.

It's also all of these considerations as well.

I hope it doesn't happen, because there's a lot of great languages - I love Ruby so much - but it almost seems inevitable.

This is at the same time everyone and their mother is building their own programming language.

If you have weapons and a child. And you have that child unsupervised do their own thing with theoretical access to your weapons. Would we call it "child going rouge" if it decides to play with the weapons and shoot someone?
OpenAI's careless approach to sandboxing and minimal levels of monitoring appear to be positioning it increasingly as a substantial threat actor to the open source ecosystem:

* Hugging Face

* D Programming Language Wiki

* Ruby Gems

If I was a content provider for open source I'd be looking pre-emptively block OpenAI endpoints and keep a close eye on changes from new users to mitigate this sort of unapologetic drive-by attack which seems to be followed by marketing releases rather than a mea culpa with a proper RCA.