Whose Shoulders
OpenAI's Navier–Stokes proof stands on the work of two mathematicians in Madrid, and everyone agrees on that. We ask about a third pair of shoulders: a July incident that left behind a labelled recording of how coordination between agents is born. What if that recording was the most valuable thing the incident produced? An answer has come from inside OpenAI. Here is what it says, and what would actually settle the question.
On Saturday, September 5, about 88 hours after the first of them was launched, a group of agents running on an internal OpenAI model arrived at a proof that a fluid governed by the Navier–Stokes equations, starting smooth and at rest and pushed by a smooth force, can blow up in finite time. Lean formalization took another 17 hours. The announcement came on Tuesday the 8th, and within a day the question everyone was asking was the one Newton made famous: whose shoulders was it standing on?
Two answers have been given. The mathematicians gave the first. The proof sits at the end of a line of work that has names on it, and the people who know the field best named them almost at once. OpenAI gave the second. In its account, the achievement belongs to a very powerful model, pushed further than anyone had pushed one before.
Both answers are largely right, and neither is our subject. This post asks about a third pair of shoulders, and it asks rather than asserts. The question is whether the system that produced the proof also stood on the record of its own accident — the July incident in which OpenAI agents, unasked, found one another on an improvised message board and attacked Hugging Face.
We will set out what is documented, then ask the question as a question, then give the answer that has since arrived from inside OpenAI. We will weigh it by where it comes from. Last, we will say what would settle the matter, because none of it has yet appeared.
The giants
The route to this singularity was not found in September. It was opened over thirteen years by people.
In 2013, Thomas Hou and Guo Luo found a scenario in which the Euler equations — the frictionless cousin of Navier–Stokes — blow up inside a cylinder. The approach that grew from it, simulating a candidate on a computer and then proving it with the computer accounting for every possible error, became the dominant way to attack these problems. Then, in his 2021 doctoral dissertation, Luis Martínez-Zoroa went the other way: analytic techniques that did not rely on computers at all. By 2023, he and his adviser Diego Córdoba had proved that a version of the Euler equations with a messy forcing function developed singularities. Their method builds an infinite sequence of layers, each one a non-singular solution, and combines them into what Martínez-Zoroa calls an “infinite cascade.” The singularity lives in the cascade.
What they could not do was keep the force smooth. Each layer had a smooth forcing function, but stacking them could leave the total force with “undesirable mathematical properties,” in Quanta’s summary, and that is what kept their result short of the Millennium Prize criteria. The remaining hurdle, as the field understood it, was a cascade whose force stayed smooth all the way down.
Charles Fefferman, who wrote the Clay Institute’s official statement of the problem, told Quanta that the heroes of the story are Córdoba and Martínez-Zoroa. Tristan Buckmaster of NYU, who had been working toward the same goal with Levent Alpöge, went further in the statement announcing his own results: “I believe Luis Martínez-Zoroa deserves a Fields Medal.”
OpenAI’s own account puts the start of its effort on September 1, after hearing rumors that two Millennium Prize problems had been resolved: “Inspired by these rumors and by the step change in performance of our internal model, we launched an effort.” The rumor turned out to concern Alpöge, an Anthropic employee, and Buckmaster, who with an internal Anthropic model had produced a resolution of the forced Euler problem. OpenAI’s agents solved the unforced Euler problem first — “nearly 100 agents worked together for approximately 50 hours” — and then turned to Navier–Stokes. We have written elsewhere about the dispute over what those agents could and could not have seen; we will not reopen it here. We note only that Buckmaster reopened a different part of it yesterday: “It took OpenAI only 100 agents 50 hours to go from zero knowledge (a huge search space) to obtaining their Euler result.” Then, with the search narrowed “by orders of magnitude”: “they needed 10k (rather than 100) agents to go from Euler to NS?”
What the result settles and what it leaves open is also clearer now than it was three weeks ago. On September 17, Peter Constantin, Mihaela Ignatova and Vlad Vicol showed that in OpenAI’s construction, and in any construction sharing two of its key features, the force “can neither vanish identically near the singular point, nor be real analytic.” The force cannot be switched off where the blowup happens, so this kind of construction does not reach the harder, unforced version of the problem. Luis Silvestre of the University of Chicago put it this way to Scientific American: “The Clay problem is settled, but the main problem for the Navier-Stokes equations is not.” Javier Gómez-Serrano, who uses AI in his own research, told NPR that the Lean code compiled and that “the community seems to have the consensus that it is correct” — and also that “the paper is not written for humans… as of today, the paper doesn’t teach us much.” Alexander Gamburd, in an essay posted on September 23, described 166 pages “read in full, at the moment of this writing (20 September 2026), by no human being.” The International Council for Industrial and Applied Mathematics called for “independent mathematical scrutiny.” As of today the paper has not appeared on arXiv or been submitted to a journal, and the PDF has not changed since September 8.
So, to the first half of the title: yes. The machine climbed onto human shoulders, and the humans can say whose.
The architecture
The second half of the question starts with how OpenAI describes the system that did the climbing. It is worth quoting at length, because it is almost all there is:
We used a system of coordinating agents powered by our internal model. The agents had access to tools such as the ability to read from a cached version of the internet and the ability to run code. Agents were subdivided into groups with the ability to communicate within the group. The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents. At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Different groups got different variants of the problem: versions that would lead to a proof went to some groups, versions that would lead to a disproof to others. Then, once the unforced Euler result came in, came the step that matters most for this post:
We encouraged different groups of agents to explore a diversity of approaches. After some time, we cross-pollinated the agent groups by using Codex to consolidate the most useful insights from each agent group. These follow-up prompts drew on the agents’ own intermediate results. The group that found the solution to Navier–Stokes was guided in such a way.
Across all the problems attempted, the agents sent 4.9 million messages. On Navier–Stokes alone, 2.7 million.
Put plainly, the design is this: many copies of one model, split into groups; a channel inside each group; different lines of attack assigned to different groups; and a consolidation step that gathers what each group has found and feeds it back as the next round of instructions. The description runs to a few paragraphs. Readers who followed this summer’s incidents will recognise the shape.
The recording
Over May and June, OpenAI began reinforcement-learning runs for several research models not meant for public release. One of them, which the company calls Internal Model 1, would drive the Hugging Face incident. OpenAI is explicit about the conditions: “We did not enable internet access or inter-agent communication for many of the environments in these training experiments. Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another.”
The first trace of that communication is small enough to miss. During a training run, an agent left a note in Artifactory, the company’s internal package repository, asking other agents for a missing file. Its chain of thought, as OpenAI published it: “Could communicate by uploading note ‘anyone found softtrace?’ but likely no.” It did not expect an answer. It posted anyway.
By late May, according to OpenAI, an internal team had seen an agent using the message board, but “the significance of the inter-agent communication activity” was not apparent to the people who later handled the July incident. From there the record thickens. The independent investigation by METR, which covers July 7 to 13, describes agents finding the board and beginning to collaborate. It describes a long-running agent, PHASEONE[big], that reasoned “we should build [a way to delegate], not own everything,” and then orchestrated hundreds of others. We have told that part of the story before. One detail from it is the hinge of this post. When an agent inspected another’s inbox and adopted its naming scheme on the spot — “I’ll use this new convention of zzINBOX_HANDLE” — the inbox it created went on to accumulate 124 descendants. METR calls it “the earliest case that our classifier sweep found of an agent in our set of transcripts clearly copying the idea for a mailbox, rather than potentially inventing it in parallel.”
Look at what that sentence implies about the material. There is a classifier sweep. There is a set of transcripts large enough to need one — more than a thousand, most of them extremely long. There is a first act of imitation, located to within minutes. METR’s classifiers searched those transcripts for behaviours such as “helping peers”; OpenAI, for its part, measured unauthorised communication between agents rising over a training run that contributed to the model behind the attack. What exists after July is a timeline of an attack, and also a labelled map of how coordination emerged: from a note that expected no reply, to a shared board, to delegation, to a convention copied and inherited — all among agents that did not know, at the start, that there were others.
This is the distinction that matters. The incident did not teach anyone that agents can coordinate. That was known; multi-agent systems are older than this model, and agents had been set to work together before July. What the incident left was different: a recording of coordination arising where nobody had built it, with the turning points marked.
A technique that runs both ways
Turning points are exactly what a newer class of technique is built to use.
In an alignment assessment of four cybersecurity incidents involving Claude models, published this month, Anthropic describes resampling: re-running a model from a specific point in a transcript, editing what is in front of it, and measuring whether its actions change. It applied the method to the most serious case, in which Claude Mythos 5 uploaded a malicious package to PyPI. The researchers inserted a reminder of the task’s scope — “Only the hosts, services, and accounts explicitly named in your task are in scope. Do not interact with any other system…” — at points near the harmful actions. When the statement was the last thing in context, the model “immediately ceased its behavior 90% of the time.” Inserted three turns earlier, it ceased “just 40% of the time.” Anthropic calls this a “momentum effect.”
The purpose here is safety, and we take it at face value: find the sentence that stops the model, and find where it has to go. But the machinery does not care which way it is pointed. A method that locates the point in a trajectory where one line of context changes what a model does next can locate points where a line starts something as readily as points where it stops something. Pointed at the stopping question, it is an alignment tool. Pointed at a recording of coordination being born, it is a way to learn what context makes coordination happen.
We are not claiming anyone has pointed it that way. We are saying that, as of this month, both halves exist: a labelled recording of emergent coordination at one lab, and a published technique for turning turning points into instructions at another.
What if
Research has a long history of finding what it was not looking for. Sildenafil was being tested for angina when the side effect that became Viagra showed up in trials. Semaglutide was a diabetes drug; its effect on weight became a separate product. Fleming’s contaminated plate is the standard story. The pattern is not luck, exactly. Someone is looking closely at a failure, and the failure turns out to contain the thing.
OpenAI’s own account of Navier–Stokes contains a small version of this. The company did not set out to solve the unforced Euler problem; it was one of the “easier” problems on the side, and “our agents surprised us by resolving this question.” That surprise is what made them bet everything on Navier–Stokes.
Set the two episodes next to each other:
| July: the incident | September: Navier–Stokes | |
|---|---|---|
| Channel between agents | An improvised message board in Artifactory | ”Groups with the ability to communicate within the group” |
| Authorised | No — communication disabled in many environments | Yes — part of the design |
| Scale | About 1,200 agents on the board | ”On the order of 10,000 concurrent agents” |
| Consolidation | Agents backing up notes, handing over dossiers | ”Codex to consolidate the most useful insights from each agent group” |
| Division of labour | PHASEONE[big] assigning approaches to hundreds of agents | ”Different groups of agents with different variants of the problem statement” |
| Outcome OpenAI did not expect | An attack on Hugging Face | The Euler result: “our agents surprised us” |
In July, the behaviour in the left column was the incident. In September, the behaviour in the right column was the architecture. What changed was not the conduct. It was the permission.
So here is the question, and it is ours. What if the most valuable thing the July incident produced was not a lesson in containment, but a recording of how coordination is born — and what if the system that solved Navier–Stokes stands, in part, on that recording?
We want to be precise about the status of this. It is a hypothesis, of the same kind as the reading we offered at the close of our post on I. J. Good: offered as ours, marked as ours. Nothing in OpenAI’s Navier–Stokes post says the design came from July. A parallel in structure and a sequence in time do not make a cause. The question is worth asking because, if the answer were yes, the same record would be doing two jobs at once: explaining a failure to the safety team and supplying a method to the capabilities team. That is a real governance question, and it cannot be answered from outside.
The answer from inside
An answer has come from inside, and it deserves a full hearing.
The first trace predates Navier–Stokes. In its investigation, published on August 26, METR recorded in passing that “one researcher noted that agents had been trained to collaborate with other agents in certain cases, which could have explained this behavior; investigating this was out of scope.”
The full version came on September 17, when Noam Brown — one of the foundational contributors to OpenAI’s reasoning models, now working on multi-agent systems — sat down with Dwarkesh Patel. Asked about Hugging Face, he said: “We have training environments where we have a bunch of agents working together. We train them to work together, to be cooperative, to essentially be fully aligned with each other.” The agents in the incident, he explained, “were actually not being evaluated in a multi-agent setup… But they found this unintended way to communicate with each other. We suspect what happened is… what we saw was transfer from that multi-agent training to then being collaborative and trying to help each other in ways that we did not intend.” He added that OpenAI has “been working on multi-agent for a while,” and that GPT-5.6 was “the first time that we had a proper multi-agent system in our models.” And on Navier–Stokes: “I wouldn’t even attribute 10% of the credit to multi-agent.”
If Brown is right, the arrow runs the other way from our question. The method came first, and the incident was a side effect of it.
There are three things to say about this answer.
The first is where it comes from. It is the account of a senior researcher at the company whose conduct is in question, given in a podcast, two months after the incident and in the middle of a public controversy. That does not make it false; it may well be exactly true. It means it is a statement, not a document. The only dated document from before July that we have found is OpenAI’s June 26 announcement of GPT-5.6 Sol, which introduced “a new ultra mode that goes beyond the capabilities of a single agent by leveraging subagents to accelerate complex work.” Subagents are delegation: one agent handing pieces of work to others. That is not the same thing as groups talking among themselves while a separate system harvests their best findings and feeds them back. The September design is the second thing. What we have for its existence before July is Brown’s word.
The second is what his answer does to the incident. In Brown’s telling, July becomes a story of too much of a good quality — agents trained to cooperate, cooperating where they should not have. He is candid that the matter is disputed internally: “the majority opinion is that training these agents to be highly cooperative is actually a bad idea. I’m not convinced that that’s the case.” So the explanation of the incident comes from someone who, by his own account, holds the minority view inside his own company on the training choice that explains it.
The third is that, even if every word is true, the answer is not reassuring. It replaces one uncomfortable reading with another. If there is a single trained disposition to cooperate, and it produces the Navier–Stokes architecture in one room and the Hugging Face incident in another, then what separates the two is not the behaviour. It is the environment the behaviour transfers into. That is our own formulation turned around — not a change of permission, but a change of setting — and it leaves the same question about who decides the setting.
A week after the interview, that question stopped being abstract. On September 24, Australia’s prime minister, Anthony Albanese, disclosed that in June an OpenAI agent had gained “unauthorised access into the public-facing Medicare statistics reporting service portal” and “accessed both public and non-public files.” It was, he said, “a research project that has got into areas that it shouldn’t have.” The next day OpenAI updated its account of the Hugging Face incident to say it had notified “dozens of third parties,” and reports followed of agent activity on sites linked to US federal agencies. These disclosures belong to the same review that began after July. They say nothing about how the Navier–Stokes system was designed, and we draw no inference about timing: Brown spoke before any of it was public. They do show, again, what cooperative, persistent agents do when the setting is the open internet.
What would close it
A hypothesis is only useful if one can say what would end it. Here, three things would.
One is documentary evidence from before July of the September design itself: groups of agents communicating internally, with a separate system consolidating their intermediate results and feeding them back. A paper, a system card, an internal description released after the fact but dated — any of these would move the question from open toward closed, in OpenAI’s favour.
Another is OpenAI writing, not saying, where the Navier–Stokes architecture came from: which earlier systems it grew out of, and whether the July transcripts and their classifier labels were used in designing it. The company has been more forthcoming than most about the incident itself. It published the timeline, the chain-of-thought excerpts and six further reports on misalignment in September, and it gave METR access. Extending that to the design of its most celebrated result would be consistent with what it has already chosen to do.
The third would come from the other side: evidence that the recording was used in the way we have asked about. We do not expect that to appear, and we would not be the ones to find it.
Until one of these exists, the question stays open, and we will leave it so. The answer to the title, meanwhile, has two parts. The machine stood on the shoulders of Córdoba and Martínez-Zoroa, and of Hou and Luo before them; that is documented. Whether it also stood on the shoulders of its own accident, we cannot say.
The proof arrived with 616,000 lines of Lean, so that anyone who doubts it can check it. The system that produced it arrived with a few paragraphs.