- On September 8, 2026, OpenAI said an internal model running about 10,000 agents proved that 3D Navier–Stokes flows with a smooth applied force can blow up in finite time. A Lean formalization is public.
- The Clay Institute says the problem has 'apparently been settled' but that its review is 'deliberately unhurried'. The proof is not peer reviewed, and the model is not available.
- For builders, the lesson is architectural: very long, parallel reasoning becomes trustworthy only when a machine verifier checks it, and the remaining risk sits in the specification.
Contents
One of the most consequential AI research results of September 2026 is a math result, and it's the clearest demonstration yet of what reasoning models can do with enough compute and a strict checker. On September 8, 2026, OpenAI said an unreleased internal model, running about 10,000 concurrent agents, had shown that the three-dimensional Navier–Stokes equations, which describe fluid flow, can develop a singularity in finite time when a smooth external force is applied. OpenAI also published a formal proof in the Lean proof assistant (OpenAI). On October 6, it followed up with a catalogue that lists 719 more AI-generated manuscripts as of October 9 (OpenAI; GitHub).
You don't need fluid dynamics to take something useful from this. Below: what OpenAI claims, what has and hasn't been verified, and what it means if you build with agents.
What OpenAI claims
The paper's main theorem says that for every positive viscosity, there is a solution of the 3D incompressible Navier–Stokes equations that "starts from rest and develops unbounded velocity in finite time while maintaining uniformly bounded kinetic energy". The force driving it is smooth and compactly supported in space and time (OpenAI paper, PDF). The accompanying Lean repository states two formal versions, one on all of 3D space and one on a periodic torus. It says these correspond to alternatives (C) and (D) in the Clay Millennium Prize problem description (GitHub). In plain terms, this is a blowup answer: the claim is that smooth solutions do not always stay smooth.
OpenAI also claims a separate result for the unforced Euler equations (fluid with no viscosity and no external force), which took about 100 agents roughly 50 hours (OpenAI).
The scale, per OpenAI:
| Item | OpenAI's figure |
|---|---|
| Model | Unnamed internal model, "significantly more capable than GPT‑6 Astra"; training ongoing |
| Agents (Navier–Stokes) | "On the order of 10,000 concurrent agents" |
| Time to result | About 88 hours after launch (reached September 5, 2026) |
| Formalization | 17 more hours, using GPT‑6 Astra, to produce the Lean proof |
| Messages / output tokens (Navier–Stokes) | 2.7 million / about 130 billion |
| Messages / output tokens (all problems) | 4.9 million / about 300 billion |
The agents had tools including a cached copy of the internet and code execution. OpenAI used Codex to consolidate insights across agent groups. OpenAI says plainly: "We do not intend to claim the Millennium Prize for this result" (OpenAI).
What is confirmed, and what isn't
Confirmed: OpenAI has published a human-readable paper and Lean source code. According to the repository's README, the Lean project builds with a specified toolchain (Lean 4.34.0-rc2 and Mathlib), and the README points to instructions for checking it with the Comparator tool (GitHub).
Not yet confirmed: that the result is accepted mathematics. On September 11, 2026, the Clay Mathematics Institute said it was contemplating "the announcement that the Navier-Stokes problem has apparently been settled". It added: "The process is deliberately unhurried, but we will provide updates" (Clay Mathematics Institute). The statement names no researcher or team, and Clay's prize rules require publication and community vetting.
The key subtlety: a Lean check guarantees that the proof proves the formal statement in the file. It does not guarantee that the formal statement means what the famous problem means. As Quanta Magazine noted, humans still have to confirm that the Lean statement matches the intended claim (Quanta).
Contested: credit and timeline. About 12 hours before OpenAI's announcement, NYU's Tristan Buckmaster and Levent Alpöge posted related results. OpenAI describes Alpöge as an Anthropic employee and says their work used an internal Anthropic model. OpenAI "recognize[s] the priority of their work on forced Euler", but claims the Navier–Stokes result as its own. In a September 10 update, it said an internal investigation found no influence from Buckmaster's Codex prompts (OpenAI). Quanta calls the details "murky" and reports Buckmaster's suggestion that OpenAI researchers may have benefited from work he and Alpöge did with OpenAI's models. Charles Fefferman, who wrote Clay's official problem description, told Quanta he was "thrilled that the problem was solved". He called mathematicians Diego Córdoba and Luis Martínez-Zoroa, whose earlier blowup techniques this line of work builds on, "the heroes of the story" (Quanta).
Physically, little changes. According to The Tufts Daily (September 24, 2026), Brown University's George Karniadakis told Nature that for air the singularity appears when the vortex is "around 70 nanometres wide". That is roughly the distance air molecules travel between collisions, where treating air as a continuous fluid, as the equations do, stops being a good approximation. Nature's September 18 news analysis covers what the result does and doesn't mean for physics.
The October 6 follow-up: volume, with caveats attached
On October 6, OpenAI released "a broad range of new mathematical results" from the same internal model (OpenAI). As of October 9, 2026, the repository says: "The current catalogue contains 719 manuscripts organized into 372 families." OpenAI posed the model about 4,000 problems and applied a significance threshold. About 42% of top-line results have Lean formalizations, and each result used on average about three hours of ChatGPT Pro thinking compute. The README is unusually candid: "Some of the unformalized results could have issues," and results are at "different stages of verification" (GitHub). Media counts vary: some reports cite 722, and Nature reported "more than 700" preprints on October 7. The repository's figures may change. Nature also reported a backlash: number theorist Alvaro Lozano-Robledo called the volume "staggering", and some mathematicians complained of being scooped (Nature).
What it does not show
- It does not show that you can buy this capability. The model is internal, and OpenAI says only that it is "working to responsibly release" it. OpenAI's API flagship as of October 9, 2026, GPT‑6 Astra, did the formalization, not the discovery.
- It does not show cheap reasoning. About 130 billion output tokens went into the Navier–Stokes run. For scale only: at GPT‑6 Astra's public list price of $50 per million output tokens, that many tokens would cost about $6.5 million. The internal model has no public price. Quanta reports OpenAI's Sébastien Bubeck estimated the cost at "several million dollars".
- It does not show reliable unsupervised research. Even Buckmaster, whose own work used AI, told Quanta that the first LLM-generated proof he received was "the most horrendous I have ever read". The system worked because a strict checker sat at the end, not because every output was good.
- It is not peer reviewed. Clay's review, and the wider community's, is ongoing. Clay's announcement page showed no update as of October 9, 2026.
What builders should take from it
1. The verifier carries the trust. The real headline isn't "AI solved math". It's that a huge, noisy, parallel search produced something trustworthy because a machine verifier was the final gate. The software equivalents are tests, type checkers, schema validators, linters, property-based tests and reproducible builds. If your agent pipeline has no hard verifier at the end, more reasoning effort just produces more confident output, not more correct output. As we explain in our reasoning-models explainer, a model's own account of its reasoning is not a substitute.
2. Your remaining risk lives in the specification. Lean moved the uncertainty from "is the proof right?" to "is the statement right?". In software, that means "does the test check what we actually care about?". Agents are very good at satisfying the letter of a check. Spend human review time on specs and tests, not on reading every line an agent writes. This is also the core argument in the case against unattended agents.
3. Parallel width is now a design choice. OpenAI's run combined thousands of agents, a coordinator that consolidated findings, and a separate model for formalization. That's a routing pattern: expensive exploration, a cheaper consolidation step, and a specialized model for verification-friendly output. Our guide to routing between cheap, mid and frontier models covers the everyday version.
4. Budget for tokens, not for prompts. At this scale, output tokens (mostly reasoning) are the whole cost. If your own agent workloads scale toward this pattern, even at a tiny fraction of the size, track reasoning tokens per completed task, not per call.
5. Treat vendor research announcements like system cards. Read the primary document, find out what was independently checked, and note what was left out. Our guide on how to read a system card applies here too: these are company claims until outside experts finish checking them.
About this storyBased on the sources linked below. Editorial standards




