October 8, 2026

Sam Altman’s company just dropped 722 manuscripts onto GitHub. They group into 372 result families. The output comes from one unreleased internal model. And mathematicians don’t know whether to applaud or sound the alarm.

Some results advance long-stalled questions in number theory. Others target topology, algebraic geometry, analysis. A few touch three of the remaining Millennium Prize Problems. None claim full victory on those million-dollar puzzles. But the sheer volume staggers the field. OpenAI says the average solution required roughly three hours of compute at ChatGPT Pro scale. The company tested the model on about 4,000 open problems.

The Scale of the Release

This October dump follows last month’s claim that the same model resolved the Navier-Stokes existence and smoothness problem. That announcement alone triggered accusations of intellectual trespass. Researchers who had fed their own partial work into OpenAI systems wondered aloud whether the model simply completed their thoughts. Tristan Buckmaster, a New York University mathematician who worked on Navier-Stokes, told The New York Times he suspects similar dynamics here. “There’s likely to be a bunch of results where they take someone’s work and then take it to completion,” he said. “I don’t think they’ve done their sort of due diligence at all.”

The new cache includes claimed progress toward the Riemann hypothesis, a solution to the four-dimensional Kakeya conjecture, and counterexamples to several longstanding beliefs. OpenAI provided summaries of the model’s reasoning for a handful of cases. It withheld the exact model name, the precise prompts, and per-problem compute figures. Those omissions matter. They frustrate verification.

But here’s the thing. Many of the results come with Lean formalizations. The proof assistant language offers machine-checked certainty for at least 162 of the manuscripts, according to the repository. Others remain unverified. OpenAI itself cautions in the README that some unformalized results “could have issues.” The company promises updates and corrections. Still, the responsibility for deep scrutiny now lands on human mathematicians. Terence Tao once likened such dumps to “dumping carcasses of raw meat onto our communal village table.” The metaphor lingers.

The Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study, tried to shape this moment. Nine prominent researchers formed the independent panel after OpenAI approached them. They issued recommendations in late September. Release results promptly. Use established academic channels. Disclose the model name, exact prompts, full chains of thought, compute costs, and selection rationale. Above all, stop treating advanced math problems as proprietary benchmarks for model capability.

OpenAI says it drew on that advice. The GitHub repository includes some reasoning summaries and average compute estimates. It avoids marketing hype in the announcement itself. Yet the company rejected the core plea. It will keep testing frontier models on hard math. Dan Roberts, OpenAI’s research lead, told reporters the evaluations help build better tools for the field. The proofs emerge as byproduct. Melanie Wood, a Harvard mathematician on the advisory group, described the release as “the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge.”

Critics see a pattern. Earlier this year OpenAI announced solutions to ten longstanding problems. Some statements required correction after mathematicians pointed out prior progress. The Navier-Stokes claim produced sharp exchanges over credit and unpublished ideas. Andreas Thom and others raised questions about whether chatbot conversations seeded the training or prompting in subtle ways. The latest release amplifies those tensions. One prompt, according to an OpenAI spokesperson quoted in Scientific American, generated almost every result when handed to a single AI agent. MIT’s Andrew Sutherland responded bluntly. Treat such one-shot claims as unverified until the model is released and others can replicate the work.

The numbers themselves sparked minor confusion. Some outlets reported 377 results. OpenAI’s repository uses 372 families. The discrepancy remains unexplained in public statements. It adds to the sense that speed outpaced precision.

Yet excitement runs alongside the frustration. Alex Kontorovich of Rutgers posted on X that a human achieving one of the headline results would earn an instant Fields Medal. If the proofs hold, they represent genuine advances. Several dozen appear to be counterexamples rather than proofs. They overturn assumptions in group theory, geometry, and more. The field gains new conjectures, new directions, new data.

But who will check all this? Peer review already strains under growing submission volumes. Young mathematicians worry their career paths narrow when AI claims the flashy breakthroughs. Senior figures debate whether the purpose of mathematics shifts from discovery to interpretation. Kevin Buzzard at Imperial College London notes that full formalization takes time. It requires formalizing references too. OpenAI says it will add more Lean proofs over time.

The company also pledged funding for workshops, conferences, and programs to help mathematicians absorb the output. It plans eventual release of the model. Those steps respond, at least partly, to the advisory group’s call for responsible communication.

Verification and Responsibility

Even so, the power imbalance persists. The model stays closed. The data it trained on stays opaque. Mathematicians cannot easily probe why it succeeded where humans stalled for decades. Did it recombine existing ideas in novel ways? Did it discover truly original techniques? Or did it, as some fear, simply accelerate the final steps of paths humans had already scouted?

Bryna Kra at Northwestern attended early meetings with OpenAI. She described mixed feelings of excitement and concern. The advisory group itself insists it does not endorse the practice of using secret models on frontier problems. Its statement reads like a careful line drawn in shifting sand.

More than 4,000 mathematicians signed an open letter titled “A Severe Misalignment of AI With Mathematics.” They argue the rush to solve problems for marketing or benchmarking undermines the discipline’s core values. Understanding matters as much as the result. OpenAI’s approach, they say, risks eroding that.

The company disagrees. Continued evaluation on hard science, it maintains, accelerates tool-building that ultimately benefits researchers. The tension will not resolve soon. New results will keep arriving. The GitHub repository already invites citation via BibTeX entries for each manuscript. Some papers have been lightly edited for readability.

Mathematicians now face a flood. They must verify, absorb, extend, or refute. Some will celebrate the acceleration. Others will mourn the loss of solitary heroic effort. Most, perhaps, feel both at once. The era when a single human could claim an entire major conjecture may be fading. In its place comes collaboration with systems that generate ideas faster than any person can check them.

OpenAI’s release improves on prior efforts by including more context and some formalization. It still falls short of the transparency many demanded. The next chapters will be written not in corporate announcements but in the patient work of human experts poring over these 722 documents. That work has only begun.

OpenAI Floods Mathematics With Hundreds of AI-Generated Solutions first appeared on Web and IT News.

Leave a Reply

Your email address will not be published. Required fields are marked *