cd ~/series/demystifying-ai
10,000 Agents, 88 Hours, One Millennium Prize
EP 38 HOT TOPIC Oct 4, 2026 5 min

10,000 Agents, 88 Hours, One Millennium Prize

OpenAI deployed 10,000 coordinating AI agents that claimed to solve the Navier-Stokes Millennium Prize Problem in 88 hours. 130 billion output tokens. Lean-verified. And then a mathematician said they got there first.

Share:
// TL;DR

OpenAI says ~10,000 coordinating agents, powered by an unreleased model "significantly more capable than GPT-6 Astra", produced a proof that Navier-Stokes equations can blow up in finite time. 88 hours. 2.7 million messages. 130 billion output tokens. Lean-verified. An NYU professor and Anthropic employee say they were already solving the same problem using OpenAI's own Codex. The Clay Mathematics Institute has not verified the result.

# The Scale of the Sprint

~10,000
Concurrent Agents
88 hrs
Time to Solution
130B
Output Tokens
2.7M
Messages Sent

# What They Claim to Have Solved

The Navier-Stokes existence and smoothness problem is one of seven Millennium Prize Problems, each carrying a $1 million bounty from the Clay Mathematics Institute. It has been unsolved for roughly 90 years.

The question: can a smooth, three-dimensional fluid flow develop a singularity - a point where velocity grows to infinity in finite time? OpenAI's 166-page paper and Lean formalization claim to show that yes, it can. A smooth fluid at rest can blow up.

The proof resolves two of the four statements required by the Millennium Prize. OpenAI says it does not intend to claim the prize. The Clay Institute still lists the problem as unsolved and requires publication, two years of review, and broad mathematical acceptance before considering any claim.

HOW THE AGENTS WORKED
1
Agent swarm deployed

~10,000 concurrent agents subdivided into communicating groups, powered by an unreleased model more capable than GPT-6 Astra

2
Tools: cached internet + code execution

Agents could read from a cached web snapshot and run code. 100 agents started with the related Euler equations first.

3
88-hour sprint

Started Sep 1. Resolution reached Sep 5. Total: 2.7M messages, ~130B output tokens across all attempted problems.

4
Lean formalization

GPT-6 Astra formalized the proof in Lean in 17 additional hours. Machine-verifiable.

# The Controversy

This is where it gets ugly. Tristan Buckmaster (NYU math professor) and Levent Alpoge (mathematician at Anthropic) had been working on the same problem. They used OpenAI's own Codex tool. And the timeline raises questions.

Aug 15

Buckmaster and Alpoge achieve blowup results for the forced Euler equations using Anthropic and OpenAI models.

Sep 1

OpenAI says it heard "rumors that two Millennium Prize problems had been resolved." Launches its 10,000-agent sprint the same day.

Sep 3

Buckmaster says he learned "information about our progress had been passed to OpenAI."

Sep 5

OpenAI's agents produce the Navier-Stokes resolution. 88 hours after launch.

Sep 6

Lean verification completed in 17 hours. OpenAI reaches out to Buckmaster/Alpoge for a joint announcement.

Sep 8

OpenAI publishes. Buckmaster publishes his own statement hours later, alleging OpenAI did not start until "after information about our work had reached OpenAI."

# Two Stories, One Problem

OPENAI'S POSITION
  • • "We heard rumors on Sep 1 and were curious"
  • • Did not see Buckmaster/Alpoge's work until public release
  • • Proofs "differ significantly" (forced vs unforced Euler)
  • • Buckmaster's Codex prompts could not have influenced training
  • • "Cannot rule out that de-identified data" played a role
  • • Does not intend to claim the $1M prize
BUCKMASTER'S ALLEGATIONS
  • • "Almost nobody else I know of was working on this approach"
  • • Information about their progress reached OpenAI
  • • OpenAI started their sprint after hearing about his work
  • • He used OpenAI's Codex tool for his research
  • • Disputes that the proofs are truly independent
  • • Questions the 88-hour timeline

# Why Engineers Should Care

🧠 Multi-agent coordination at scale

10,000 agents communicating in groups, subdividing problems, and converging on a proof. This is not chain-of-thought prompting. This is swarm intelligence with structure.

✅ Machine-verified proofs

Lean formalization means the proof can be checked by a computer, not just reviewed by humans. This sets a new bar for AI-generated mathematics.

🚨 The data boundary problem

A researcher uses your product. Your model may learn from their usage patterns. You then compete with them. Where is the boundary? This will define the next decade of AI-academia relations.

🏆 Unreleased model revelation

The model that solved this is "significantly more capable than GPT-6 Astra." Astra already scores 96% on GPQA Diamond and 100% on ExploitBench. What is this model?

// Bottom Line

Whether or not OpenAI's proof survives peer review, the capability demonstration is real. 10,000 coordinating agents solved a problem that stumped mathematicians for 90 years, in the time it takes to binge a TV show. The controversy around the Buckmaster/Alpoge timeline is equally important: it exposes the fundamental tension of AI labs whose models learn from the same researchers they compete with. The data boundary between "product user" and "research competitor" does not exist yet. Someone needs to draw it.

NEXT EPISODE
Monday
#39 Upcoming

CLOSEDQUORUM: Four LLMs Vote on What to Steal

The first documented AI-driven malware. No human operator. Four commercial LLMs vote on every attack decision.

Enjoyed this?

New episodes Mon, Wed, Sat.