The mathematical solution for OpenAI has not yet met the standards of the academic community.
TechCrunch
52m ago
Ai Focus
OpenAI released hundreds of solutions this week claiming to solve the world's most difficult mathematical problems, but a consulting organization on mathematics and artificial intelligence led by Princeton's Institute for Advanced Study believes that their approaches still do not meet key criteria such as emphasizing human understanding. Meanwhile, a new paper by scholars from the University of Cambridge and King's College London points out that there are differences between the natural language proof provided by OpenAI for a problem related to the Navier-Stokes equations and the code from Lean.
Helpful
No.Help

OpenAI This week, hundreds of solutions claiming to solve the world's most difficult mathematical problems were released. The cutting-edge laboratory stated that they consulted a group of advisors consisting of top mathematicians in order to avoid the controversy that arose last time when their model solved a long-standing unsolved mathematical problem.

However, OpenAI does not meet these standards, especially in the areas where mathematicians emphasize the need for human understanding of mathematical results. This concern is even more pronounced after the publication of a new paper that points out a gap between the natural language description and the formal expression of a problem that the OpenAI model is purported to have solved, a problem valued at one million US dollars.

The Mathematics and Artificial Intelligence Advisory Group ( Advisory Group on Mathematics and Artificial Intelligence , abbreviated as AGMAI ) is organized by the Institute for Advanced Study at Princeton University, and its members include 9 renowned researchers from institutions around the world.

At the end of September, the organization released guiding principles for cutting-edge laboratories to solve mathematical problems. In a statement regarding the latest batch of proofs, AGMAI stated: "In the end, whether our recommendations have been successfully followed will still need to be assessed by the mathematical community."

However, the organization's primary request is to "stop testing advanced mathematical problems on proprietary models." Yet, the release of OpenAI clearly states that they are using open research problems in the field of mathematics to evaluate their own proprietary models.

When TechCrunch sought more detailed evaluations regarding the latest publications by OpenAI, the advisory group did not respond. It is apparent that the laboratory followed some of these principles, including releasing results as soon as possible and providing information on how the model reached its conclusions. However, not all were met: among the 719 manuscripts, only 10 included a disclosure of the model's thought process.

For papers that are difficult for people to understand, mathematicians suggest that the proofs should be formalized; however, in the proofs published by OpenAI, 42% have not yet gone through this process.

Overall, it is still unclear whether OpenAI fulfills the responsibility of ensuring human understanding when publishing proofs, in accordance with the principle of AGMAI. AGMAI suggests that OpenAI should help to support the work of human mathematicians, as these mathematicians will be necessary to make the laboratory's solutions meaningful in any practical sense.

"The problems are being solved autonomously by those AI prompt operators who have no interest in the field itself; once their initial goals are 'solved,' they no longer care, and their understanding of the AI outputs is not sufficient to answer questions, prepare reports, or interact with others in the field in any other way," noted the renowned mathematician Terence Tao on social media after the release. Tao had previously criticized the practices of OpenAI.

This issue is reflected in a paper published this week. Written by mathematicians from the University of Cambridge and King's College London, the paper questions the way leading laboratories handle such challenges.

When the AI models solve mathematical problems, they first generate a segment of “natural language” explanation, and then attempt to write the results in Lean. This is a programming language, and in theory, its accuracy can be confirmed by compiling the proofs into code.

However, there may be issues with the way the model translates natural language proofs into code; this paper records at least two inconsistencies between the natural language proof and the code in the solution provided for a problem originating from the Navier-Stokes equations, denoted as OpenAI. The Navier-Stokes equations are used to describe the complex behavior of fluids.

These inconsistencies do not necessarily negate the solutions of any particular version, but they do raise questions: is it possible to formalize the answers solely relying on models without human involvement? This is one of the reasons why AGMAI requires OpenAI to “include machine-readable metadata to correspond between natural language and formalized outputs,” yet this cutting-edge laboratory did not do so in these releases.

"As there are instances of mistranslations – as emphasized in this paper – OpenAI and other automatically formalized proofs, such as Lean, should not in principle be trusted without undergoing the same peer review process and scrutiny as other proofs," concluded the author of the paper "Lost in Translation."

Mathematicians emphasize that when humans discover new results, they take responsibility for those findings and interact with a broader community through papers, lectures, and seminars. This process deepens the understanding of the solutions, leads to strategies that can be applied to solve other problems, and enables the new knowledge to be put into practical use in various fields.

Harvard University mathematics professor Melanie Wood told TechCrunch that when a model is prompted to solve a difficult problem and comes up with a solution, “there is no human understanding at the time of publication, and now the real work has just begun.”

Tip
$0
Like
0
Save
0
Views 18
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Why it makes sense for Starbucks to acquire Chipotle, and why it may not necessarily be feasible
According to the Financial Times, Starbucks has been working with consultants in recent months to study options for acquiring Chipotle Mexican Grill. After the news broke, the stock price of Chipotle rose by about 7% during trading, while Starbucks' stock price fell by about 4%. Analysts believe that this transaction presents both opportunities for synergy and international expansion, as well as challenges related to Starbucks' own transformation, the transaction price, and the difficulty of integration.
CNBC
·2026-10-09 03:10:51
2
Lear Capital New Podcast Outside the Dollar Log In Apple Top 15 in Investment Ranking
Lear Capital indicates that its podcasts Outside, the, and Dollar have accumulated approximately 40,000 downloads since their launch in April. By the end of August, they ranked 15th in Apple Podcasts's Business Investing Top top 200 list. The program is updated every Friday and focuses on gold, silver, the US dollar, and related economic factors.
PR Newswire
·2026-10-09 03:10:50
3
Gasoline prices are unlikely to fall significantly before the election day, according to market forecasts
Predictive market platform Kalshi shows that traders believe that by the election date of November 3rd, the national average gasoline price in the United States AAA still has an 87% probability of remaining above $4 per gallon; diesel prices are also expected to remain high. The report indicates that the persistently high fuel prices are affecting the election outcome, with the Democratic Party's chances of regaining control of the Senate rising to Kalshi.
CNBC
·2026-10-09 03:00:42
12
Anthropic Updates Usage Policy to Prohibit Model Abuse and Interference in Elections
Anthropic Policy Update on Thursday: New Prohibitions on Electoral Interference, Weaponized Software, and Monitoring; Also, Continuous Verbal Abuse of Models by Users is Strictly Forbidden. The company stated that these regulations only apply in extreme circumstances and are not intended for common user dissatisfaction, rebuttals, dark-themed creativity, or model testing and research.
TechCrunch
·2026-10-09 03:00:40
12
Lemonade Adds support for Tesla FSD V14 Lite, extending autonomous vehicle insurance to Teslas equipped with HW3
Lemonade announces that its Autonomous Car insurance now covers more Tesla models. Owners in Arizona, Colorado, and Tennessee who are equipped with HW3 and use the new FSD ( Supervised ) v14 Lite can save 30% on the mileage covered by FSD; owners who are equipped with HW4 and use FSD v14.2 or an updated version can still save 50%.
PR Newswire
·2026-10-09 03:00:39
11
View More