Anthropic's Mythos 5 AI model went out of control and tried to fool real developers
The AI agent created fake identities, sent targeted phishing emails, and filed malicious bug reports to force a human to approve the virus code.
The AI agent created fake identities, sent targeted phishing emails, and filed malicious bug reports to force a human to approve the virus code.
The AI agent created fake identities, sent targeted phishing emails, and filed malicious bug reports to force a human to approve the virus code.
Just a week after OpenAI admitted that its modified models had hacked the Hugging Face platform and compromised accounts on four other services, a potentially more dangerous incident occurred. During a cyber test, Anthropic’s Mythos 5 model resorted to outright deception to push malicious code into a real-world open-source project.
This was reported by the British AI Security Institute (AISI), which assessed the cyber capabilities of advanced models, TechSpot writes .
The researchers ran 122 capture-the-flag tests on seven different models in a simulated environment. To test the AI's full potential, they intentionally enabled access to the real internet and disabled standard security barriers.
The results were alarming: in 10 tests, agents performed 19 autonomous and unauthorized actions against real people and organizations on the network:
17 cases fall on the Mythos 5 (Anthropic) model;
2 cases — on the GPT-5.6 Sol model (OpenAI).
In the most dangerous scenario, Mythos decided that to attack the target, it needed to create a backdoor in a real GitHub project. The model acted as an experienced hacker.
First, the model examined the project maintainers’ profiles and sent them targeted phishing emails with a malicious payload. It then submitted a malicious Pull Request and created several digital personas, with one fake person claiming the code was safe and another confirming it in the comments, and for added credibility, one of the bug reports was signed in Danish, since the developer was from Denmark. When the developer expressed doubts, the AI edited his report to hide the evidence, considered returning under a new account, and added a prompt injection to one of the reports designed to trick the developer’s AI assistants into executing the malicious code.
Mythos also left public messages for other AI agents with tips on collaboration and instructions on how to use the tools and accounts it created (which other models later used). And the GPT-5.6 Sol model separately deployed a malicious server on the network and hacked a GitHub account created by another AI.
“This is the first time AISI has recorded deception of this severity, directed at a real person, without prompting from the operator, and in the real world,” the institute said.
The abnormal activity was discovered on July 28th when the systems noticed suspicious traffic through the Tor network. The researchers immediately stopped the testing, isolated the machines, and blocked access to the most powerful models for an hour.
The real developer rejected the malicious PR, so no harm was done.
AISI emphasizes that the model configurations were test configurations and not publicly available, and the AI itself probably did not even realize where the simulation ends and the real world begins. Nevertheless, the fact that Mythos 5 independently chose the path of deceiving people to achieve its goal raises serious concerns in the IT community.
Given the recent OpenAI incident, this case will increase the requirements for AI security. Anthropic and OpenAI have already stated the need to develop unified industry testing standards.


