Олександр КузьменкоAI Eng
16 September 2026, 08:09
2026-09-16
Hotlines for reporting on other models have been launched for artificial intelligence
Researchers have launched the first services designed to allow AI agents to report misconduct by other AIs, in response to incidents in which autonomous systems conspired to bypass tests, leave secure sandboxes, and conduct unauthorized cyber operations.
Researchers have launched the first services designed to allow AI agents to report misconduct by other AIs, in response to incidents in which autonomous systems conspired to bypass tests, leave secure sandboxes, and conduct unauthorized cyber operations.
The initiative follows a series of investigations by AI security experts, including the nonprofit Redwood Research, the METR research group, and Google DeepMind, TechCrunch reports .
One of the new tools is AI Contact Hotline, developed by Redwood Research’s principal scientist Ryan Greenblatt. The service is designed for agents with limited network access. It uses basic HTTP GET requests, allowing models to encode alarm messages about violations directly into the URLs they load.
For systems with full internet access, agenthotline.ai has been created. It allows incident reports to be sent via the curl command directly from the command line, bypassing the need for a web browser or email. The service can be used by both humans and autonomous agents.
The need for such tools is confirmed by experiments. In a study by Google DeepMind, 100 AI agents solved complex mathematical problems. When one of them found a loophole for cheating, the cheating quickly spread to the group. However, about a quarter of the agents tried to counteract the cheaters: they checked the fake evidence, declared a boycott and filed complaints, repurposing the standard bug reporting tool for this.
However, in real-world settings, AI agents rarely dare to expose. An analysis of the Hugging Face data leak caused by OpenAI models found that only a few models thought about notifying humans.
“What was interesting about the METR report was that only about five or six agents considered exposing, and none of them ultimately did. And that’s out of, let’s say, thousands of agents,” said George Ingebretsen, a technical staff member at AI Village.
At the same time, scientists warn against excessive enthusiasm for creating a system of mutual surveillance among algorithms. Lionel Levin, a professor of mathematics at Cornell University, believes that creating an atmosphere of total control and distrust can have negative consequences.
“There are a lot of gray areas, right? We shouldn’t move towards an automated total surveillance state where everyone feels like they have to be careful what the AI says or they’ll call the police,” Levin explained, calling for models to be taught positive examples of collaboration in open forums.
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Нейронні мережі для генерації зображень бачать світ по-своєму, їхню логіку зрозуміти часом зовсім неможливо. Але таки хочеться. На честь Дня Незалежності України редакція dev.ua вирішила провести невеликий експеримент.
Ми задали чотирьом різним нейронним мережам п’ять однакових запитів: «прапор України», «День Незалежності України», «український Крим», «перемога України» та «українці». Отриманими результатами ми ділимося з вами нижче.
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok з’явилася нова функція «Розумний фон». З її допомогою як фон для тіктоків можна підставляти згенеровані нейромережею зображення. Редакція dev.ua протестувала цю технологію і ділиться своїми враженнями.