Mathematicians marvel, and grumble, at OpenAI’s trove of new results

It’s been two days since OpenAI’s “slop drop” of AI-generated results on 372 open math and computer science problems. Scientists are just beginning to parse through the heap and pick out the most exciting breakthroughs in their own fields.
It will be months, perhaps years, before we understand just how much novel mathematics was added to the literature this week, but everyone agrees it’s been a week without precedent. “Yesterday will be the day when a new era of math and mathematical physics began,” Harvard University mathematician Michael Douglas told Scientific American on Tuesday. It was an even bigger moment, he thinks, than last month’s solution of the Navier-Stokes problem, one of math’s million-dollar challenges, by the same new internal model at OpenAI.
READ MORE: The most exciting claims from OpenAI’s heap of new proofs
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
Already certain aspects of this new era are becoming clear. Many of the results, experts say, are far better written than the Navier-Stokes proof, suggesting that OpenAI’s math team is working to overcome its model’s inability to explain its math clearly to humans. Hector Pasten, a mathematician at the Pontifical Catholic University of Chile, says the papers he’s looked at improve on the company’s notorious lack of citations. “I think they are doing a better job of giving credit,” he says.
But the quality remains extremely uneven—mathematicians are finding some papers readable and others utterly nonsensical. “Others are much, much harder to make sense of—they’re sort of this AI slop,” says Roland Bauerschmidt, a mathematician at New York University. “If I had received these by e-mail from a nobody, I probably would have deleted it, it’s so badly written.”
Even much of that seeming nonsense, though, is ostensibly correct. At this point, 42 percent of the results are coded in the programming language Lean, which means their logic is almost certainly sound, even if their English-language write-up is too hard to follow.
On the other hand, three of the papers have already been withdrawn after experts uncovered a significant error, causing some to question the more than half of the proofs that aren’t yet Lean-verified. “While I expect most of these will ultimately be found to be correct, or at least correctable, we should expect some of the proofs to contain mistakes, possibly serious ones,” says Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology.
The community’s reaction to the way OpenAI chose to drop the results has been mixed, to say the least. Sutherland argues that OpenAI should also release the problems its model failed to solve, which the company has said number in the thousands. “I expect they will not want to do that,” he adds, “but that is how open science works. You need to report your negative results along with your positive results.”
The lack of transparency goes further. A number of mathematicians have told Scientific American that they doubt anyone on OpenAI’s team would have thought to ask some of the questions the model answered, either because they’re extremely niche or because they contradict a belief most experts previously held. Several independently suggested that it seems the model itself is coming up with problems to solve, perhaps in an extended conversation between a swarm of agents. This idea seems corroborated by the fact that a number of papers contain strikingly similar ideas and methods—as if their authors were communicating. OpenAI didn’t reply to a request for comment on this question.
Others in the community have been more outspoken in their criticism. A statement from the Association for Human Mathematics decries the company’s actions. “Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power,” it reads. “We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centers human understanding.”
Subscribe to Support Independent Journalism
Great science journalism requires human expertise, time, effort and creativity. And it costs money. That’s why I and the journalists here at Scientific American hope you’ll join our community.
When you subscribe, you are supporting staff and freelance journalists who are passionate about telling science stories that are true, important and compelling. Our editors and reporters are often experts in their fields, which means they understand the nuances of big discoveries and can untangle the breakthroughs from the hype. With a subscription, you are also supporting rigorous fact-checking to ensure the words we publish are precise and accurate. And you’re supporting original illustrations, graphics and photos that bring you closer to an advanced laboratory, an ice sheet in Antarctica or a space mission in orbit. You’re helping us craft other types of high-quality journalism as well: Our newsletters are carefully written, edited and curated by staffers you have or will come to know and love. Our Science Quickly podcast is based on original reporting, collaboration with editors and scientists and exacting production.
Subscriptions keep this engine running so we can continue to deliver thoughtful, rigorous and independent science journalism to you. In an era of viral misinformation, this work is crucial. If you value what we do, I hope you’ll consider joining us as a subscriber



