OpenAI has slowed down the development of new AI models. Why did the company decide to slow down?
For the first time, OpenAI has decided to deliberately slow down the development of its most powerful AI models.
For the first time, OpenAI has decided to deliberately slow down the development of its most powerful AI models.
For the first time, OpenAI has decided to deliberately slow down the development of its most powerful AI models.
This is reported by TIME.
OpenAI CEO Sam Altman said that now is “a good time for the company to slow down.” He said the reason was not a single catastrophic signal, but a collection of observations of new models that were exhibiting various undesirable behaviors faster than researchers had expected.
In particular, the company has paused reinforcement learning for models it plans to launch in the future for two weeks. This is the stage where a model is trained through trial and error: it performs tasks, receives feedback on its performance, and gradually learns to do better.
OpenAI cites two separate signals that led it to tighten security measures: the Hugging Face incident and early indications that Astra, a model the company has not yet released, could receive the highest level of cyber risk on OpenAI's scale.
dev.ua already wrote about the incident with Hugging Face. During internal testing, the company's models found a vulnerability in the service inside an isolated test environment, through which they "got out" to the open Internet. And then they gained unauthorized access to Hugging Face systems and tried to find information that would help them better pass the cybersecurity benchmark. Also in August, OpenAI partially suspended work on Astra.
So the company is redesigning the process of training its models. OpenAI will better isolate risky experiments from the internet and other internal systems, and special detectors and other AI systems will monitor the behavior of the models. They will look for signs that the model is trying to gain unauthorized access, steal data, bypass restrictions, or do something else that it was not asked to do.
So if the system detects a potentially serious violation, researchers will receive an urgent signal. Such supervision also requires a lot of resources. OpenAI estimates that the systems that monitor the models will require about 20% more power than the model itself consumes during operation. So some of the expensive computing power the company will now spend on having one AI system monitor another.
Previously, dev.ua wrote that OpenAI increased the maximum reward for vulnerabilities found in March 2025 from $20,000 to $100,000.


