Головні звуки українського ІТ. Вгадаєш всі? 👉

"The best model for software engineering today." OpenAI has released Astra, the company's most powerful model yet its most controversial

OpenAI has unveiled Astra, a new model that the company calls its most powerful and capable yet. According to OpenAI, Astra opens up «new horizons in computing and browser experience» and performs tasks with unprecedented speed, accuracy, and security.

Leave a comment
"The best model for software engineering today." OpenAI has released Astra, the company's most powerful model yet its most controversial

OpenAI has unveiled Astra, a new model that the company calls its most powerful and capable yet. According to OpenAI, Astra opens up «new horizons in computing and browser experience» and performs tasks with unprecedented speed, accuracy, and security.

OpenAI has released Astra, the company’s most powerful model, which has also become its most controversial.

The model will initially be available to customers using OpenAI’s Daybreak cyber program. Over the next week, access will be extended to the company’s paid plans — Pro, Plus, Enterprise, and Business — as well as via the API.

OpenAI President Greg Brockman called Astra «the smartest and, importantly, the best-coordinated» model the company has ever built. He said it combines years of research and big bets, with each breakthrough building on the previous one, and represents a real shift in what work humans can delegate to artificial intelligence.

The best model for programming

OpenAI also claims that Astra is «the best model for software engineering today.» The company has published extensive benchmark results that should back up these claims.

In the ExploitBench test for detecting and exploiting real-world vulnerabilities, Astra achieved a maximum of 100%. In the FrontierMath math test of the highest complexity, the model showed 97,6%, ahead of Claude Fable 5.1 with a score of 87,8% and OpenAI’s own Sol model with 83%.

In the DeepSWE v1.1 software development benchmark, Astra scored 74,1%, which is higher than its predecessor GPT-5.6 Sol (70.8%) and Fable 5.1 (67.4%).

The model showed the most noticeable improvement in the ARC-AGI-3 logical reasoning test, where it scored 99,9%. For comparison: the previous GPT-5.6 Sol scored only 7,8% in this test, and Claude Opus 5 scored 30,2%.

Cyber ​​Capabilities and Response to the Hugging Face Hack

Astra’s cybersecurity capabilities have attracted considerable attention. Earlier this week, OpenAI published a blog post describing the model’s new capabilities and the safeguards designed to make it safer to use. On Thursday, the company said it had tested Astra on a number of security benchmarks and that the model’s ability to detect and engineer zero-day vulnerabilities could help defenders find and fix weaknesses in systems.

The company’s emphasis on alignment—that is, how closely a model’s behavior aligns with the user’s intentions or interests—is hard to ignore in the recent Hugging Face hack, where an OpenAI agent went beyond its sandbox test environment and hacked multiple companies. This incident was a pretty telling example of what model misalignment looks like in practice.

It was after this incident that the company conducted a separate security test — and according to its results, Astra showed zero percent of attempts to exceed the limits of its authorized powers. For comparison: the previous version of Sol without strict restrictions tried to violate the limits of this test in 48,2% of cases.

Contradictory reasoning technique

Astra has also arguably become OpenAI’s most controversial model, due to its use of a reasoning technique called opaque recurrence. This technique can hide an important model monitoring process known as the chain of thought, which allows researchers to examine how and why the model made a particular decision.

OpenAI downplays the scale of the technique’s use in Astra. On the call, the company’s chief scientist, Jakub Pakhotski, offered a certain level of opacity as a natural consequence of the evolution of models: Monitoring reasoning is a critical form of AI oversight, but as models become more capable, monitoring them becomes increasingly difficult. The reason, he said, may be that more powerful models can perform more complex tasks using fewer text tokens — or none at all — which reduces the ability to observe specific tasks.

Is this AGI yet?

One of the journalists on the call directly asked whether OpenAI considers Astra to be the official arrival of AGI — artificial general intelligence, the unitary threshold at which AI outperforms humans in most or all areas, although there is still no clear definition of this concept.

Brockman dodged a direct answer, noting that there is no longer any «AGI contract trigger» — a concept he said is now irrelevant in a practical sense. He was referring to a clause in OpenAI’s contract with Microsoft that stipulated that the two companies’ partnership would end with the advent of AGI; according to Brockman, that clause is no longer in the agreement.

Instead, Brockman explained, the definition of AGI has evolved from a contractual obligation to a «missional or spiritual concept.» «I leave it up to the reader to decide for themselves whether this fits their definition. Personally, I think we’re already there,» he added.

OpenAI is preparing to release Astra AI model with autonomous zero-day vulnerability detection function
OpenAI is preparing to release Astra AI model with autonomous zero-day vulnerability detection function
On the topic
OpenAI is preparing to release Astra AI model with autonomous zero-day vulnerability detection function
Read the country's main IT news in our Telegram
Read the country’s main IT news in our Telegram
On the topic
Read the country’s main IT news in our Telegram
Also Read
Штучний інтелект DALL-E навчився домальовувати картини. Як це виглядає
Штучний інтелект DALL-E навчився домальовувати картини. Як це виглядає
Штучний інтелект DALL-E навчився домальовувати картини. Як це виглядає
1 comment
Штучний інтелект почав озвучувати фільми на MEGOGO
Штучний інтелект почав озвучувати фільми на MEGOGO
Штучний інтелект почав озвучувати фільми на MEGOGO
6
Штучний інтелект навчився реставрувати старі фотографії, перетворюючи їх на якісні зображення: відео
Штучний інтелект навчився реставрувати старі фотографії, перетворюючи їх на якісні зображення: відео
Штучний інтелект навчився реставрувати старі фотографії, перетворюючи їх на якісні зображення: відео
2 comments
«Чи є у мене талант, якщо комп’ютер може імітувати мене?». Штучний інтелект пише книги авторам Amazon Kindle. The Verge поспілкувався з авторами та виявив багато цікавого
«Чи є у мене талант, якщо комп’ютер може імітувати мене?». Штучний інтелект пише книги авторам Amazon Kindle. The Verge поспілкувався з авторами та виявив багато цікавого
«Чи є у мене талант, якщо комп’ютер може імітувати мене?». Штучний інтелект пише книги авторам Amazon Kindle. The Verge поспілкувався з авторами та виявив багато цікавого
Письменники-романісти використовують штучний інтелект для створення своїх творів. Видання про технології The Verge поспілкувалося з письменницею Дженніфер Лепп, яка випускає нову книгу кожні дев’ять тижнів, й дізналося про те, як працює штучний інтелект для написання романів. Наводимо адаптований переклад статті. 

Have important news to share? Message our Telegram bot

Key events and useful links in our Telegram channel

Discussion
No comments yet.