Оксана СтепураAI Eng
9 October 2026, 12:09
2026-10-09
OpenAI has published hundreds of solutions to complex mathematical problems. But scientists have found that the AI could change the meaning of the proof during verification
OpenAI has published hundreds of solutions to complex math problems generated by its AI model. But researchers found that in one proof, the textual explanation didn't fully match the code that was supposed to prove it correct. The researchers found at least two such discrepancies.
OpenAI has published hundreds of solutions to complex math problems generated by its AI model. But researchers found that in one proof, the textual explanation didn't fully match the code that was supposed to prove it correct. The researchers found at least two such discrepancies.
This week, OpenAI published 719 papers with AI-based solutions to complex mathematical problems. But scientists have raised questions about how the company verified these results.
To check a mathematical proof, an AI can first describe the solution in ordinary words, and then translate it into a special Lean language. It allows the computer to check whether the proof is constructed correctly. However, during such a “translation”, the AI may not accurately convey the meaning of the proof in the Lean language. As a result, the code may not prove exactly what is written in the textual explanation. And even if the computer confirms the correctness of the code, this does not mean that the original proof will be correct.
It was this problem that researchers at the University of Cambridge and King's College London drew attention to, studying OpenAI's work on a problem related to the Navier-Stokes equations.
Navier–Stokes equations
The company also consulted with the AGMAI group of mathematicians before publishing, which developed recommendations for AI labs. However, it did not implement all of them.
In particular, scientists have called for the use of closed AI models to study complex mathematical problems. And OpenAI confirmed that the vast majority of published results were obtained using its own internal model, which it has not yet released into the public domain.
In addition, OpenAI published AI reasoning records for only 10 out of 719 papers, and presented approximately 42% of the proofs in formal form.
Mathematician Terence Tao has pointed out another problem: humans who get solutions from AI don't always understand them well enough to explain them to other scientists. Harvard professor Melanie Wood also noted that when AI gives a ready-made solution, it doesn't mean that humans understand it.
Earlier, dev.ua wrote that 25 Fields Medal winners, including Ukrainian mathematician Maryna Vyazovska, criticized the race of AI companies for mathematical breakthroughs. Scientists warned that companies are rushing to announce the solution of complex problems before properly verifying the results.