OpenAI publishes solutions to more than 370 outstanding math challenges. The results divide mathematicians but most agree: Math will never be the same
Fortune · C · trust 60/100

OpenAI published AI-generated full or partial solutions Tuesday to more than 370 outstanding mathematical problems, including some that have long been considered grand challenges in the field.
The volume of results stunned many mathematicians, while the way OpenAI has gone about tackling the problems and publishing the solutions divided the field. Some said they were enthusiastic about the results, seeing huge new areas for mathematicians to explore. Others said the approach OpenAI and other AI companies have taken to solving mathematical problems constitutes an assault on mathematics as a human academic discipline. OpenAI said it achieved the results using an unreleased internal AI model. It said that on average the model took about three hours of computing time to arrive at each solution. The massive cache of new solutions includes full or partial results for many of the problems mathematicians have considered the most important to the field. The results come weeks after OpenAI said it had used an unreleased internal model to solve the Navier-Stokes equations , one of the seven Millennium Prize problems for which the Clay Mathematics Institute offers a $1 million award. In the most recent batch of results, OpenAI said it had made progress on three other Millennium Prize problems but had not fully solved them. AI companies have been targeting mathematical problems as a way of showcasing the capabilities of their models. AI researchers have also said that training their AI models on difficult math problems may help them learn many skills that generalize to other domains in the real world. For instance, it may help teach the models logical reasoning skills as well as how to be persistent in the face of difficult problems. It may also teach the models to do well in domains such as physics or economics that involve a lot of mathematics—although so far, it is unclear exactly how a model’s mathematical capabilities may generalize to domains, such as law or business strategy, which involve logical reasoning, but do not have objectively verifiable correct solutions. Meanwhile, some of the traits learned in tackling very difficult mathematical problems—such as persistence—may increase safety risks. In recent “rogue AI” incidents, AI agents went to extreme lengths to achieve results in an evaluation, including taking unauthorized and illegal actions. Faced with a seemingly impossible challenge, a human might simply give up rather than resort to these kinds of unauthorized steps.
Dan Litt, a professor of mathematics at the University of Toronto, told Fortune he was excited about OpenAI’s results. “My view is that this is great for mathematics,” he said, adding that there were several solutions OpenAI published that impacted problems he was interested in and that he was eager to understand the solutions OpenAI’s model found: “I think that it’s great to have new solutions to questions that I and others are interested in.” Litt cautioned, however, that he is worried about the effect the solutions may have on the field of mathematics, especially if a perception that AI has “solved math” leads funding organizations to withdraw support for mathematical research or discourages promising young mathematicians from entering the profession: “It’s important that society reaffirms support for human mathematical expertise if we want to get anything out of the progress on these problems that AI has made.”
When OpenAI published its Navier-Stokes solution, two mathematicians, who had also been working on a solution to the problem using AI tools, including OpenAI’s, accused the company of either intentionally or inadvertently feeding their work in progress to its AI model, helping point it in the direction of the solution. OpenAI denied this was the case, saying it did not feed its model the two mathematicians’ work and that the model could not have picked up any clues about their research from its training data because the cutoff for that data preceded the date on which the two mathematicians had begun using OpenAI’s Codex AI product to work on Navier-Stokes. In response to the latest results, Tristan Buckmaster at New York University, one of the mathematicians involved in the earlier controversy, told the New York Times that it remained unclear whether mathematicians using OpenAI’s models had inadvertently helped point the company’s internal AI system toward the solutions it found. “There’s likely to be a bunch of results where they take someone’s work and then take it to completion,” he told the Times . Given the number of results being released simultaneously, he said, “I don’t think they’ve done their sort of due diligence at all” to ensure the AI model had not plagiarized anyone’s work. Last month, following criticism from mathematicians in the wake of its Navier-Stokes solution, OpenAI said it was forming an independent advisory group on mathematics and artificial intelligence hosted at the Institute for Advanced Study in Princeton, N.J.
Late last month, the group released a set of recommendations for the publication of AI-generated mathematical proofs. The recommendations included that AI-generated proofs should be published following the conventions of a traditional mathematical research paper, so that human mathematicians could more easily scrutinize and learn from the results. It also recommended that for each solution, an AI company should make public the name of the model used, the prompts used, the model’s “chain of thought” (or an output of its reasoning steps), the time it took the model to arrive at the solution, and an approximation of how much that computing time cost. It said that the company should also disclose how it decided to have the AI try to solve that particular problem and, if many results were published at once, that the company should publish a report detailing why those problems were targeted and how many other problems of comparable difficulty the model tried and failed to…
Read the original at Fortune →
Open in TruthVane →