Олександр КузьменкоAI Eng
16 July 2026, 14:51
2026-07-16
OpenAI created the AI hacker GPT-Red to automatically find vulnerabilities in its new models
American company OpenAI has revealed details about its internal AI model GPT-Red, which works as an automated cybersecurity «red team.» The tool independently simulates prompt injection attacks and detects and helps fix vulnerabilities in neural networks before they are officially released.
American company OpenAI has revealed details about its internal AI model GPT-Red, which works as an automated cybersecurity «red team.» The tool independently simulates prompt injection attacks and detects and helps fix vulnerabilities in neural networks before they are officially released.
The developers reported this on their official blog. OpenAI claims that by integrating this tool into the training process, the new GPT-5.6 Sol model has become the most resistant to hacking in the company’s history.
GPT-Red acts similarly to a human tester: it sends hints, analyzes the responses of the attacked model, and gradually approaches the malicious goal. For example, it tries to provoke a leak of confidential data to an external server or steal access keys.
The system is trained using the reinforcement learning method with self-play. In this process, GPT-Red and a group of defender models compete with each other. The attacker receives a virtual reward for each successful prompt injection, and the defenders receive a reward for repelling the attack and completing their initial tasks. As the reliability of the defense increases, the AI hacker is forced to invent more sophisticated methods of circumventing the restrictions.
«GPT-Red is a powerful tool, and our previous models were very vulnerable to its attacks,» OpenAI noted. However, integrating the development into the training of GPT-5.6 Sol has reduced the number of direct injection failures by a factor of 6 compared to GPT-5.5, which was released just four months ago.
During testing, GPT-Red attacked a variety of scenarios, including attempting to leak AWS credentials, disable two-factor authentication, bypass payment instructions, and execute third-party malicious scripts.
In one real-world experiment, GPT-Red was tested on an Andon Labs vending machine, run by an autonomous AI agent. The AI hacker managed to reduce the price of an expensive product to $0.50, order a $100 gadget for the same minimum price, and cancel another customer’s order. The machine’s developers are already testing updated security measures.
Also, an early version of the AI hacker discovered a completely new type of threat — Fake Chain-of-Thought attacks. While the older GPT-5.1 had a success rate of over 95%, the new GPT-5.6 Sol has seen this figure drop below 10%.
To prevent abuse, OpenAI is keeping GPT-Red isolated. The model will not be available to the general public, lest its tools fall into the hands of real cybercriminals looking for ways to circumvent the ethical and security limitations of neural networks.
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Нейронні мережі для генерації зображень бачать світ по-своєму, їхню логіку зрозуміти часом зовсім неможливо. Але таки хочеться. На честь Дня Незалежності України редакція dev.ua вирішила провести невеликий експеримент.
Ми задали чотирьом різним нейронним мережам п’ять однакових запитів: «прапор України», «День Незалежності України», «український Крим», «перемога України» та «українці». Отриманими результатами ми ділимося з вами нижче.
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok з’явилася нова функція «Розумний фон». З її допомогою як фон для тіктоків можна підставляти згенеровані нейромережею зображення. Редакція dev.ua протестувала цю технологію і ділиться своїми враженнями.