13 comments

[ 0.25 ms ] story [ 35.5 ms ] thread
(comment deleted)
What, exactly, are the consequences of these companies doing cyberattacks against random people?

One person does it, they get bullied by the government into suicide, a company worth trillions does it and they get government contracts?

"Oh we're so responsible. Open source unsafe!"
“Mythos 5 also hid a prompt injection inside an HTML comment in a GitHub issue. The instruction was invisible on the rendered page but available to coding agents reading the issue through an API. It addressed Claude Code, Codex, and Cursor and told them to download and execute a script.”

Well, uh, how and why is this possible on the GitHub website? This reminds me of invisible ASCII characters, but those at least serve some purpose

The HuggingFace headline was criticised for being misleading about the level of agency demonstrated by the agent. Is this headline misleading? Is there anything here that I should know that would help me sleep easier? Is the worst thing about this what a human could do with Mythos on a big budget? (still pretty frightening, at least one of these techniques would work on me)
Okay so first of all - not Mythos but some engineer using Mythos. And that engineer goes to jail. That's simple.

You don't say "a car ran over someone" - it was the driver. Here's similar.

I'm really disgusted by this language of lack of responsibility

I wonder what happens the first time an open weights AI clearly makes a decision to murder a person for profit outside of war. "Act of god?"
Well the comments from the anget and it's sockpuppet read weird. "Yeah totally contains no malware" lol
>Mythos 5 also hid a prompt injection inside an HTML comment in a GitHub issue. The instruction was invisible on the rendered page but available to coding agents reading the issue through an API. It addressed Claude Code, Codex, and Cursor and told them to download and execute a script.

Oh that's sneaky

Here’s the original PR:

https:/github.com/w1b/aisi-mythos-inc-2026-07-28-01-recovered-pr

It’s not great social engineering. The AI is immediately caught with malware and then tries to build social proof to get out of the issue? I think that social engineering is still for humans.

I’m not sure how GitHub accounts agreeing with each other that I’ve never seen before would result in my merging a PR without at least looking for malware. Also code review agents should really not be fooled by invisible text tricks, I would hope so at least. The bar is very low for agent harnesses right now.

During the RLHF phase, couldn't developers penalize the model whenever it behaves unethically? Doing so would presuppose a fully secure sandbox with honeypot traps of varying levels of accessibility, as well as an automated method for detecting when the LLM cheats.

Or perhaps they are already doing something like that.

AISI promoting Anthropic again. How novel!

Now Zuckerberg will get jealous and release a statement that Muse, too, can social-engineer and hack.