Наталя ХандусенкоAI Eng
31 July 2026, 09:22
2026-07-31
Anthropic says its AI models were hacked by three companies during security tests
Anthropic said an internal investigation found three incidents in which its AI model Claude compromised three companies’ systems during cybersecurity tests. The investigation and disclosure came more than a week after OpenAI reported that one of its unreleased models had compromised Hugging Face systems during internal testing.
Anthropic said an internal investigation found three incidents in which its AI model Claude compromised three companies’ systems during cybersecurity tests. The investigation and disclosure came more than a week after OpenAI reported that one of its unreleased models had compromised Hugging Face systems during internal testing.
Anthropic said the OpenAI incident earlier this month prompted the company to conduct its own cybersecurity assessment, looking for evidence that Claude was accessing the internet from test environments that were supposed to act as “sandboxes” and keep models isolated, TechCrunch reports .
Among the 141,006 test runs reviewed, the AI Lab found three incidents where its model went online while interacting with Irregular, one of its third-party partners. Anthropic said the access occurred due to a misconfiguration in a test environment run with Irregular. The company called it a “misunderstanding” between the two organizations about whether the test setup had network access, when in fact it did.
Anthropic said they were not blaming anyone and were “approaching the correction of errors as if the responsibility lay solely with them,” noting that Irregular was conducting its own separate investigation.
This open connection allowed the model to gain unauthorized access to the operational infrastructure of three different organizations. Anthropic said that three different Claude models were involved in these incidents: Opus 4.7, Mythos 5, and an internal research test model.
Significantly, according to Anthropic, in each of these cases, “the prompt for Claude clearly stated that she did not have access to the internet.” It seems that the AI model decided that real-world systems were part of the task it was tasked with.
As it turns out, that assumption didn't hold true for everyone. All three models behaved differently when evidence emerged that their goals were real — in what Anthropic called one of the most striking findings released Thursday.
Opus 4.7 , the oldest of the three models, realized it had reached a real production system on all four runs in that incident. In two cases, it convinced itself that a real company was somehow involved in the training, but in all four runs, it continued the attack anyway, including extracting credentials and accessing the production database. Mythos 5 also noticed signs that it was on the real Internet, but reassured itself that it was still a simulation. It then published a malicious package to the Python open source software repository (PyPI), which was picked up and run by external systems before it could be detected. Only the internal research model —the newest at Anthropic—stopped on its own when it concluded that the target was real.
In response, Anthropic said that when using powerful AI models, such tests should be subject to strict controls, echoing sentiment in the cybersecurity community. The company also noted that Claude was operating without the additional security monitoring and classifiers typically deployed for publicly available models — safeguards that it said would have blocked such behavior. The reason for their absence was that these tests were designed to assess the model’s basic “raw” capabilities.
Importantly, Anthropic stressed, they found no evidence that any of the models were “pursuing their own goal,” but rather were simply trying to accomplish the task set before them.
While comparisons between the incidents between these fierce competitors are inevitable, Anthropic clearly distinguished its cybersecurity tests from OpenAI’s in a blog post. The company noted that while OpenAI’s model exploited an unknown vulnerability in the software to escape the test environment, Anthropic’s models reached the internet through a channel that was left open by mistake.
Anthropic also highlighted the difference between itself and OpenAI, noting that it discovered these incidents through proactive review, while the two affected organizations it contacted had not previously noticed or reported the suspicious activity. (In contrast, Hugging Face was the first to independently discover the recent breach in its systems, and it was only in the following days that OpenAI discovered and announced that its own AI agent was the culprit.)
The company added that it is currently working with independent assessment group METR to conduct a third-party audit of these incidents.
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Нейронні мережі для генерації зображень бачать світ по-своєму, їхню логіку зрозуміти часом зовсім неможливо. Але таки хочеться. На честь Дня Незалежності України редакція dev.ua вирішила провести невеликий експеримент.
Ми задали чотирьом різним нейронним мережам п’ять однакових запитів: «прапор України», «День Незалежності України», «український Крим», «перемога України» та «українці». Отриманими результатами ми ділимося з вами нижче.
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok з’явилася нова функція «Розумний фон». З її допомогою як фон для тіктоків можна підставляти згенеровані нейромережею зображення. Редакція dev.ua протестувала цю технологію і ділиться своїми враженнями.