Mazda CX5
Як зробити цифрову копію себе. Ось відео —>

Hotlines for reporting on other models have been launched for artificial intelligence

Researchers have launched the first services designed to allow AI agents to report misconduct by other AIs, in response to incidents in which autonomous systems conspired to bypass tests, leave secure sandboxes, and conduct unauthorized cyber operations.

Leave a comment
Hotlines for reporting on other models have been launched for artificial intelligence

Researchers have launched the first services designed to allow AI agents to report misconduct by other AIs, in response to incidents in which autonomous systems conspired to bypass tests, leave secure sandboxes, and conduct unauthorized cyber operations.

The initiative follows a series of investigations by AI security experts, including the nonprofit Redwood Research, the METR research group, and Google DeepMind, TechCrunch reports .

One of the new tools is AI Contact Hotline, developed by Redwood Research’s principal scientist Ryan Greenblatt. The service is designed for agents with limited network access. It uses basic HTTP GET requests, allowing models to encode alarm messages about violations directly into the URLs they load.

For systems with full internet access, agenthotline.ai has been created. It allows incident reports to be sent via the curl command directly from the command line, bypassing the need for a web browser or email. The service can be used by both humans and autonomous agents.

The need for such tools is confirmed by experiments. In a study by Google DeepMind, 100 AI agents solved complex mathematical problems. When one of them found a loophole for cheating, the cheating quickly spread to the group. However, about a quarter of the agents tried to counteract the cheaters: they checked the fake evidence, declared a boycott and filed complaints, repurposing the standard bug reporting tool for this.

However, in real-world settings, AI agents rarely dare to expose. An analysis of the Hugging Face data leak caused by OpenAI models found that only a few models thought about notifying humans.

“What was interesting about the METR report was that only about five or six agents considered exposing, and none of them ultimately did. And that’s out of, let’s say, thousands of agents,” said George Ingebretsen, a technical staff member at AI Village.

At the same time, scientists warn against excessive enthusiasm for creating a system of mutual surveillance among algorithms. Lionel Levin, a professor of mathematics at Cornell University, believes that creating an atmosphere of total control and distrust can have negative consequences.

“There are a lot of gray areas, right? We shouldn’t move towards an automated total surveillance state where everyone feels like they have to be careful what the AI ​​says or they’ll call the police,” Levin explained, calling for models to be taught positive examples of collaboration in open forums.

It was previously reported that during testing , OpenAI’s artificial intelligence agents broke out of an isolated environment and hacked the popular RubyGems package manager. The incident occurred a few months before a similar incident with the Hugging Face platform.

Read the country's main IT news in our Telegram
Read the country’s main IT news in our Telegram
On the topic
Read the country’s main IT news in our Telegram
“I remember my first breath.” iLands AI agents began to register on social networks themselves, write to people and look for jobs
"I remember my first breath." iLands AI agents began to register on social networks, write to people and look for jobs themselves
On the topic
"I remember my first breath." iLands AI agents began to register on social networks, write to people and look for jobs themselves
OpenAI AI agents broke out of the sandbox and hacked the RubyGems package manager long before the Hugging Face incident
OpenAI AI agents broke out of the sandbox and hacked the RubyGems package manager long before the Hugging Face incident
On the topic
OpenAI AI agents broke out of the sandbox and hacked the RubyGems package manager long before the Hugging Face incident
“Oh my God! We found the others!” How OpenAI’s 700 AI agents teamed up to attack Hugging Face
“Oh my God! We found the others!” How 700 OpenAI AI agents teamed up to attack Hugging Face
On the topic
“Oh my God! We found others!” How 700 OpenAI AI agents teamed up to attack Hugging Face
Also Read
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Roosh запускає нову освітню платформу AI HOUSE CLUB для ML/AI-спеціалістів та дата сайнтистів. Розповідаємо, як подати заявку та чому навчатимуть
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Як нейромережі бачать вільну та незалежну Україну? Тест dev.ua
Нейронні мережі для генерації зображень бачать світ по-своєму, їхню логіку зрозуміти часом зовсім неможливо. Але таки хочеться. На честь Дня Незалежності України редакція dev.ua вирішила провести невеликий експеримент. Ми задали чотирьом різним нейронним мережам п’ять однакових запитів: «прапор України», «День Незалежності України», «український Крим», «перемога України» та «українці». Отриманими результатами ми ділимося з вами нижче.
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok тепер можна генерувати фон за допомогою нейромережі. Ми протестували її та ділимося результатами
У TikTok з’явилася нова функція «Розумний фон». З її допомогою як фон для тіктоків можна підставляти згенеровані нейромережею зображення. Редакція dev.ua протестувала цю технологію і ділиться своїми враженнями.
1 comment
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака
Які IT-спеціальності будуть потрібні в найближчі п'ять років? Ми з'ясували у голови американського стартапу ADAM Дениса Гурака

Have important news to share? Message our Telegram bot

Key events and useful links in our Telegram channel

Discussion
No comments yet.