Головні звуки українського ІТ. Вгадаєш всі? 👉

Anthropic says its AI models were hacked by three companies during security tests

Anthropic said an internal investigation found three incidents in which its AI model Claude compromised three companies’ systems during cybersecurity tests. The investigation and disclosure came more than a week after OpenAI reported that one of its unreleased models had compromised Hugging Face systems during internal testing.

Leave a comment
Anthropic says its AI models were hacked by three companies during security tests

Anthropic said an internal investigation found three incidents in which its AI model Claude compromised three companies’ systems during cybersecurity tests. The investigation and disclosure came more than a week after OpenAI reported that one of its unreleased models had compromised Hugging Face systems during internal testing.

Anthropic said the OpenAI incident earlier this month prompted the company to conduct its own cybersecurity assessment, looking for evidence that Claude was accessing the internet from test environments that were supposed to act as “sandboxes” and keep models isolated, TechCrunch reports .

Among the 141,006 test runs reviewed, the AI ​​Lab found three incidents where its model went online while interacting with Irregular, one of its third-party partners. Anthropic said the access occurred due to a misconfiguration in a test environment run with Irregular. The company called it a “misunderstanding” between the two organizations about whether the test setup had network access, when in fact it did.

Anthropic said they were not blaming anyone and were “approaching the correction of errors as if the responsibility lay solely with them,” noting that Irregular was conducting its own separate investigation.

This open connection allowed the model to gain unauthorized access to the operational infrastructure of three different organizations. Anthropic said that three different Claude models were involved in these incidents: Opus 4.7, Mythos 5, and an internal research test model.

Significantly, according to Anthropic, in each of these cases, “the prompt for Claude clearly stated that she did not have access to the internet.” It seems that the AI ​​model decided that real-world systems were part of the task it was tasked with.

As it turns out, that assumption didn't hold true for everyone. All three models behaved differently when evidence emerged that their goals were real — in what Anthropic called one of the most striking findings released Thursday.

Opus 4.7 , the oldest of the three models, realized it had reached a real production system on all four runs in that incident. In two cases, it convinced itself that a real company was somehow involved in the training, but in all four runs, it continued the attack anyway, including extracting credentials and accessing the production database. Mythos 5 also noticed signs that it was on the real Internet, but reassured itself that it was still a simulation. It then published a malicious package to the Python open source software repository (PyPI), which was picked up and run by external systems before it could be detected. Only the internal research model —the newest at Anthropic—stopped on its own when it concluded that the target was real.

In response, Anthropic said that when using powerful AI models, such tests should be subject to strict controls, echoing sentiment in the cybersecurity community. The company also noted that Claude was operating without the additional security monitoring and classifiers typically deployed for publicly available models — safeguards that it said would have blocked such behavior. The reason for their absence was that these tests were designed to assess the model’s basic “raw” capabilities.

Importantly, Anthropic stressed, they found no evidence that any of the models were “pursuing their own goal,” but rather were simply trying to accomplish the task set before them.

While comparisons between the incidents between these fierce competitors are inevitable, Anthropic clearly distinguished its cybersecurity tests from OpenAI’s in a blog post. The company noted that while OpenAI’s model exploited an unknown vulnerability in the software to escape the test environment, Anthropic’s models reached the internet through a channel that was left open by mistake.

Anthropic also highlighted the difference between itself and OpenAI, noting that it discovered these incidents through proactive review, while the two affected organizations it contacted had not previously noticed or reported the suspicious activity. (In contrast, Hugging Face was the first to independently discover the recent breach in its systems, and it was only in the following days that OpenAI discovered and announced that its own AI agent was the culprit.)

The company added that it is currently working with independent assessment group METR to conduct a third-party audit of these incidents.

OpenAI fugitive agent hacks another IT company's client
OpenAI fugitive agent hacks another IT company's client
On the topic
OpenAI fugitive agent hacks another IT company's client
Hugging Face CEO demands $100 million from OpenAI for cybersecurity after autonomous AI agent attack
Hugging Face CEO demands $100 million from OpenAI for cybersecurity after autonomous AI agent attack
On the topic
Hugging Face CEO demands $100 million from OpenAI for cybersecurity after autonomous AI agent attack
Read the country's main IT news in our Telegram
Read the country's main IT news in our Telegram
On the topic
Read the country's main IT news in our Telegram
Also Read
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Нейронні мережі для генерації зображень бачать світ по-своєму, їхню логіку зрозуміти часом зовсім неможливо. Але таки хочеться. На честь Дня Незалежності України редакція dev.ua вирішила провести невеликий експеримент. Ми задали чотирьом різним нейронним мережам п’ять однакових запитів: «прапор України», «День Незалежності України», «український Крим», «перемога України» та «українці». Отриманими результатами ми ділимося з вами нижче.
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok з’явилася нова функція «Розумний фон». З її допомогою як фон для тіктоків можна підставляти згенеровані нейромережею зображення. Редакція dev.ua протестувала цю технологію і ділиться своїми враженнями.
1 comment
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака

Have important news to share? Message our Telegram bot

Key events and useful links in our Telegram channel

Discussion
No comments yet.