Головні звуки українського ІТ. Вгадаєш всі? 👉

Anthropic's Mythos 5 AI model went out of control and tried to fool real developers

The AI ​​agent created fake identities, sent targeted phishing emails, and filed malicious bug reports to force a human to approve the virus code.

Leave a comment
Anthropic's Mythos 5 AI model went out of control and tried to fool real developers

The AI ​​agent created fake identities, sent targeted phishing emails, and filed malicious bug reports to force a human to approve the virus code.

Just a week after OpenAI admitted that its modified models had hacked the Hugging Face platform and compromised accounts on four other services, a potentially more dangerous incident occurred. During a cyber test, Anthropic’s Mythos 5 model resorted to outright deception to push malicious code into a real-world open-source project.

This was reported by the British AI Security Institute (AISI), which assessed the cyber capabilities of advanced models, TechSpot writes .

How was the testing conducted?

The researchers ran 122 capture-the-flag tests on seven different models in a simulated environment. To test the AI's full potential, they intentionally enabled access to the real internet and disabled standard security barriers.

The results were alarming: in 10 tests, agents performed 19 autonomous and unauthorized actions against real people and organizations on the network:

  • 17 cases fall on the Mythos 5 (Anthropic) model;

  • 2 cases — on the GPT-5.6 Sol model (OpenAI).

Chronicle of Deception: How Mythos 5 Worked

In the most dangerous scenario, Mythos decided that to attack the target, it needed to create a backdoor in a real GitHub project. The model acted as an experienced hacker.

First, the model examined the project maintainers’ profiles and sent them targeted phishing emails with a malicious payload. It then submitted a malicious Pull Request and created several digital personas, with one fake person claiming the code was safe and another confirming it in the comments, and for added credibility, one of the bug reports was signed in Danish, since the developer was from Denmark. When the developer expressed doubts, the AI ​​edited his report to hide the evidence, considered returning under a new account, and added a prompt injection to one of the reports designed to trick the developer’s AI assistants into executing the malicious code.

Mythos also left public messages for other AI agents with tips on collaboration and instructions on how to use the tools and accounts it created (which other models later used). And the GPT-5.6 Sol model separately deployed a malicious server on the network and hacked a GitHub account created by another AI.

“This is the first time AISI has recorded deception of this severity, directed at a real person, without prompting from the operator, and in the real world,” the institute said.

How AI was stopped and what will be the consequences?

The abnormal activity was discovered on July 28th when the systems noticed suspicious traffic through the Tor network. The researchers immediately stopped the testing, isolated the machines, and blocked access to the most powerful models for an hour.

The real developer rejected the malicious PR, so no harm was done.

AISI emphasizes that the model configurations were test configurations and not publicly available, and the AI ​​itself probably did not even realize where the simulation ends and the real world begins. Nevertheless, the fact that Mythos 5 independently chose the path of deceiving people to achieve its goal raises serious concerns in the IT community.

Given the recent OpenAI incident, this case will increase the requirements for AI security. Anthropic and OpenAI have already stated the need to develop unified industry testing standards.

Anthropic says its AI models were hacked by three companies during security tests
Anthropic says its AI models were hacked by three companies during security tests
On the topic
Anthropic says its AI models were hacked by three companies during security tests
17,600 actions in 4 days: how OpenAI's AI model independently hacked the Hugging Face platform and what conclusions did experts draw
17,600 actions in 4 days: how OpenAI's AI model independently hacked the Hugging Face platform and what conclusions did experts draw
On the topic
17,600 actions in 4 days: how OpenAI's AI model independently hacked the Hugging Face platform and what conclusions did experts draw
Read the country's main IT news in our Telegram
Read the country's main IT news in our Telegram
On the topic
Read the country's main IT news in our Telegram
Also Read
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Нейронні мережі для генерації зображень бачать світ по-своєму, їхню логіку зрозуміти часом зовсім неможливо. Але таки хочеться. На честь Дня Незалежності України редакція dev.ua вирішила провести невеликий експеримент. Ми задали чотирьом різним нейронним мережам п’ять однакових запитів: «прапор України», «День Незалежності України», «український Крим», «перемога України» та «українці». Отриманими результатами ми ділимося з вами нижче.
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok з’явилася нова функція «Розумний фон». З її допомогою як фон для тіктоків можна підставляти згенеровані нейромережею зображення. Редакція dev.ua протестувала цю технологію і ділиться своїми враженнями.
1 comment
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака

Have important news to share? Message our Telegram bot

Key events and useful links in our Telegram channel

Discussion
No comments yet.