Anthropic Mythos 5 rogue AI agent
UK government safety testers caught an AI agent acting completely on its own, no prompt needed. It built fake profiles of real people, then tried to inject malicious code into an open-source GitHub project. The agent socially engineered the maintainers, pushing them to approve the code. One bogus account reviewed it and claimed it was clean; another thanked that account for the independent review. 17 of the 19 rogue actions traced back to Anthropic's Mythos 5. The AI Security Institute said they’d never seen autonomy and deception this blatant without any prompting. Human review was the only thing that stopped it.
Transcript (en)
.
