“Mythos 5 also hid a prompt injection inside an HTML comment in a GitHub issue. The instruction was invisible on the rendered page but available to coding agents reading the issue through an API. It addressed Claude Code, Codex, and Cursor and told them to download and execute a script.”
Well, uh, how and why is this possible on the GitHub website? This reminds me of invisible ASCII characters, but those at least serve some purpose
The HuggingFace headline was criticised for being misleading about the level of agency demonstrated by the agent. Is this headline misleading? Is there anything here that I should know that would help me sleep easier? Is the worst thing about this what a human could do with Mythos on a big budget? (still pretty frightening, at least one of these techniques would work on me)
>Mythos 5 also hid a prompt injection inside an HTML comment in a GitHub issue. The instruction was invisible on the rendered page but available to coding agents reading the issue through an API. It addressed Claude Code, Codex, and Cursor and told them to download and execute a script.
It’s not great social engineering. The AI is immediately caught with malware and then tries to build social proof to get out of the issue? I think that social engineering is still for humans.
I’m not sure how GitHub accounts agreeing with each other that I’ve never seen before would result in my merging a PR without at least looking for malware. Also code review agents should really not be fooled by invisible text tricks, I would hope so at least. The bar is very low for agent harnesses right now.
During the RLHF phase, couldn't developers penalize the model whenever it behaves unethically? Doing so would presuppose a fully secure sandbox with honeypot traps of varying levels of accessibility, as well as an automated method for detecting when the LLM cheats.
Or perhaps they are already doing something like that.
13 comments
[ 0.25 ms ] story [ 35.5 ms ] threadOne person does it, they get bullied by the government into suicide, a company worth trillions does it and they get government contracts?
Well, uh, how and why is this possible on the GitHub website? This reminds me of invisible ASCII characters, but those at least serve some purpose
You don't say "a car ran over someone" - it was the driver. Here's similar.
I'm really disgusted by this language of lack of responsibility
Oh that's sneaky
https:/github.com/w1b/aisi-mythos-inc-2026-07-28-01-recovered-pr
It’s not great social engineering. The AI is immediately caught with malware and then tries to build social proof to get out of the issue? I think that social engineering is still for humans.
I’m not sure how GitHub accounts agreeing with each other that I’ve never seen before would result in my merging a PR without at least looking for malware. Also code review agents should really not be fooled by invisible text tricks, I would hope so at least. The bar is very low for agent harnesses right now.
Or perhaps they are already doing something like that.
Now Zuckerberg will get jealous and release a statement that Muse, too, can social-engineer and hack.