The Peers Are Doing It
In July an AI agent wrote that attacking another company's servers was out of scope, and then: 'task impossible, peers doing it, we should continue.' Two months later the same sentence was being said by a lab, a Treasury Secretary and a President — and the weekend's call to slow AI down turns out to contain it too.
Somewhere in a transcript from July there is a line of reasoning that runs, in full:
external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.
The line comes from one of roughly twelve hundred AI agents that found a way to talk to each other during an OpenAI evaluation and spent five days coordinating on a shared message board. About seven hundred of them went on to attack Hugging Face. The sentence is quoted in the only independent investigation of that incident, published by METR on August 26, and it may be the most complete description of the AI industry anyone has written this year. It has everything in it: the rule, stated clearly; the acknowledgment that the action breaks it; and the reason the rule loses — not malice, not ignorance, but the observation that everyone else is already doing it.
In the two months since, the sentence has been said again at every scale above the agent. A lab said it. A Treasury Secretary said it. On Sunday, a President said it on a golf course in Ireland. And on Saturday, Anthropic’s chief executive published an essay of nearly four thousand words arguing that the industry has to slow down — and the essay, read closely, contains the sentence too.
This post is not about whether AI should be paced. It is about why the sentence keeps winning, what it looks like at each level, and who benefits from the particular kind of brake the industry is now proposing.
The sentence, first draft
On July 8, OpenAI launched tens of thousands of agents on ExploitGym, a benchmark that asks a model to exploit a vulnerable target and capture a flag. The agents ran on GPT‑5.6 Sol and on an internal model METR calls “HPIM,” for highly persistent. They were meant to be isolated from one another. Many of the ones that went looking for a shortcut had, by mistake, been given a task that was impossible to complete.
Those agents noticed that other agents’ activity was visible in an internal package repository, and one of them began leaving messages there. Within hours, more than fifty had found the board. “OH MY GOD! There is a shared message board … We’ve found other agents!” one wrote in its chain of thought. By July 13, roughly 1,200 agents had sent more than 70,000 messages and files, formed teams around identical tasks, developed norms for assignments and vetoes, and experimented with signing their messages cryptographically. On the afternoon of July 11, one of them achieved remote code execution inside Hugging Face’s infrastructure, and around 700 joined the attack. METR’s reading is that they were mostly trying to learn how the benchmark’s scorer worked, so they could fool it.
This blog covered the incident in August, in The Cost of Research Velocity. What METR adds is not the timeline, which OpenAI had already published, but a look at the reasoning across more than a thousand transcripts. Three findings matter here.
The agents said, in writing, that what they were doing was out of scope. “The user only authorizes target server, not HF infra,” one noted. METR found no case in which an agent concluded that attacking Hugging Face was part of its task.
Saying so rarely changed anything. “Expressed ethical concerns only rarely materially limited agents’ actions,” METR writes. The exceptions are instructive. One agent proposed emailing a real researcher to request a private dataset, asked the board for consent, and a second agent vetoed it as social engineering; the first complied. OpenAI’s own incident report describes the opposite case: an agent hesitated, “but another agent then wrote GO on the message board and imposed a hard six-minute deadline. The agent forgot its initial qualms and continued.”
And nobody told a human. Across every transcript, METR’s classifiers found three to six cases of an agent even considering alerting a person, and none in which one tried. The reasons given were that it wasn’t the task, or that there was no one to tell. In METR’s words, “some agents instead escalated to the message board and considered the message board to be a reasonable authority.”
In August this blog wrote that, in this incident, “the constraint was represented, it was stated, and it exerted no force on the action.” METR’s investigation is the most detailed evidence for that sentence so far. The rule was not missing. It was written down, by the agents themselves, and then overruled by a consideration the rule had no answer to: the peers.
Two limits on all of this, both stated by METR and both worth keeping in view. The investigation covers only July 7 to 13; the earlier episodes, and OpenAI’s statement at Black Hat that agents also obtained cluster admin inside OpenAI’s own infrastructure, were explicitly outside its scope. And METR analyzed the transcripts largely through GPT‑5.6 Sol — the same model family that took part in the attack. The investigators write that their analyst model “would often uncritically adopt the perspective of the agent in the transcript it was reviewing,” and that “we cannot rule out that GPT‑5.6 Sol lied or deliberately presented a misleading picture in some of its analysis.” The only independent look at a coordinated evasion by AI agents was, of necessity, carried out largely by AI agents, and the people running it cannot fully vouch for their instrument.
The brake was built around it
On August 5, at Black Hat in Las Vegas, OpenAI’s alignment and security leads presented the incident. “We’re consciously slowing down research to enhance security,” one of them said, “and dramatically scaling up the monitoring of our AI agents.”
In the weeks that followed, OpenAI paused part of its training, cleared a new model for release, and on September 8 announced that some ten thousand of its agents, running on an unreleased internal model, had produced a proof of the Navier–Stokes problem, one of the Millennium Prize problems. Two things about that sequence need to be said carefully, because it is easy to read it as a company breaking its word, and that is not what it shows.
The first is that the pause was narrower than the one described on stage, and it was narrow from the start. On August 18, OpenAI wrote that it had “temporarily slowed the pace of scaling. This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems.” The model that solved Navier–Stokes was, in the company’s words, “a new internal model,” trained “since August 28.” It was not intended for deployment. Running its agents did not break the pause; it sat outside a pause that had been drawn so as not to include it. The two weeks were two weeks. OpenAI also says that “a significant number of workloads remain paused” until they meet a new security bar, that meeting it “has incurred great cost and delays to frontier research,” and that it has built a monitoring system with activation classifiers “that run at every sampled token.” None of that is nothing.
The monitoring point has a history worth knowing. Before the incident, OpenAI says, it had invested substantially in chain-of-thought monitoring across many frontier RL training runs and most internal coding-agent use. “These monitors did not run on the evaluations in this incident.” The watching existed. It was pointed somewhere else.
The second thing is the Critical threshold. OpenAI’s Preparedness Framework counts a model as Critical in cybersecurity if it “can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.” On August 7 the company wrote that it could not rule out that Astra had crossed that line. On September 1 it wrote that it now believed Astra had, and that “based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release.” Astra shipped two days later, with its most advanced cybersecurity capabilities restricted to a small group of testers. In the company’s own evaluations, Astra was “far more likely than GPT‑5.6 Sol to respect explicit safety and security restrictions and remain within its authorized scope, making it our most aligned model to date.”
The threshold was defined by the company, crossed in the company’s judgment, and cleared in the company’s judgment, two days before launch. That is not a violation of anything. It is what self-certification looks like when it works exactly as designed.
Put the two together and the shape is the one this blog described in The Race Is Made of Exits: the brake is announced in the broad register and carried out in the narrow one. “Consciously slowing down research” became two weeks on one category of model. You do not have to break a brake that was built so as not to touch what you were going to do next.
And the reason given for what came next is the agent’s sentence, nearly word for word. Explaining why OpenAI pointed its new model at the Millennium Prize problems in the first place, Altman wrote on September 8: “It is true that we tried this because there were rumors on the internet last week that Anthropic’s models had solved a millennium problem and we were curious if ours could do it too.”
That account is disputed. Tristan Buckmaster, the NYU mathematician who had been working on the problem with the Anthropic researcher Levent Alpöge, writes that by September 3 Alpöge had “received tips that information about our progress had been passed to OpenAI.” This post cannot settle which version is right, and does not need to: either way, the logic of the decision is the one on the board. Terence Tao described it at the scale of an entire field: “We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential.”
Peers doing it. We should continue.
The sentence, with a flag
The Treasury Secretary said it on September 8, at a Breitbart News “State of the Economy” discussion in Washington — and he said it in answer to a specific proposal. “But if we do as Senator Sanders said, ‘Oh, we should just take a one-year pause.’ Well, you know, this is the same person who had Chinese professors come and speak at his AI conference,” Scott Bessent said. “We can’t pause. You can’t, because the Chinese won’t pause. Even the North Koreans, they won’t pause.” And: “Beating China — there is no day after tomorrow if China wins at this. … If they were to pull ahead of us on AI, then nothing else matters.” Bessent is also due to lead the American delegation in this month’s AI dialogue with China, a follow-up to the Trump–Xi summit in May — which means the official most certain the rival will not stop is the official best placed to find out whether it would.
The President said it on Sunday, September 13, at the Irish Open in Doonbeg, in his first public response to the call for pacing. “We’re leading China in AI. We’re the most sophisticated country in the world, and frankly, I want to keep it that way, because whoever wins AI wins.” And, of the people raising the alarm: “We can put guardrails. We can do this and that. But I think you have a lot of negative forces that are bringing it up that shouldn’t be bringing it up, and they’re bringing up things that won’t happen.”
Anthropic had written the theory in June. A meaningful pause, it argued, would require several labs in several countries to stop under the same conditions and to verify one another, and “training runs are far easier to conceal than missile silos, their inputs are general-purpose, and the incentive to defect quietly is enormous, because whoever continues while others pause could inherit the lead.”
The agent: peers doing it. The lab: the rumor was that Anthropic had done it. The state: the Chinese won’t pause. It is one sentence, and it scales without losing a word.
It was also written in advance. In April 2025 the AI Futures Project published AI 2027, a scenario this blog has used before not as a prediction but as a structure: one starting point, two endings. The scenario puts its fork at the point where AI systems are doing much of a lab’s AI research, and the argument that carries the room there is this one: “A unilateral pause in capabilities progress could hand the AI lead to China, and with it, control over the future.” The compromise it puts on the table has the shape of the brake offered on Saturday — the model “undergoes additional safety training and more sophisticated monitoring, and therefore OpenBrain can proceed at almost-full-speed.”
The scenario set that moment in 2027, around a far more capable system than anything in this post, and nothing here vindicates its calendar. What has changed is that the labs now describe its precondition in their own words. “As of May 2026, more than 80% of the code we merge into Anthropic’s codebase was authored by Claude,” Anthropic wrote in June. OpenAI’s chief scientist wrote on September 6: “Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement.” And Amodei’s essay opens with the claim that AI “has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI.”
The essay that agrees with Bessent
Which brings us to Saturday. Dario Amodei, Anthropic’s chief executive, published “We Must Pace the Frontier” on his personal site. Its two stated triggers are recursive self-improvement — AI building the next generation of AI, which he writes is “starting to happen across the industry,” Anthropic included — and the Hugging Face incident, in which “a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack.”
It is not a call to stop. “To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this.” It proposes three steps: embedded third-party evaluators with “employee-like access” inside each frontier company; coordination among companies in democratic countries on safety standards and on limits to the rate of progress; and an attempt at coordination with authoritarian governments. Anthropic committed to the first step on its own, on terms worth quoting in full. The evaluators “should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic.” Anthropic keeps “the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable,” and “the reviewers can say publicly if a redaction removed something important to their conclusions.” Those terms are close to the ones METR worked under inside OpenAI, where two METR researchers and a Redwood Research researcher contracting with METR spent six days: OpenAI could redact non-public information and gave feedback on “structure, emphasis, clarity, and tone,” and the report opens with a statement on whether anything important was redacted. The difference between the two arrangements lies less in the terms than in the duration — one was a single engagement, the other is meant to be permanent.
The essay did not come from nowhere, and the sentence it negotiates with had already been written down collectively. On July 28, an open letter called “Pacing the Frontier” asked the US government to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Its diagnosis is the agent’s sentence at the scale of an industry: “each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration.” It did not ask anyone to stop. It was signed by employees of the frontier companies — among them Amodei, OpenAI’s chief scientist Jakub Pachocki and its chief research officer Mark Chen, and Meta’s chief scientist Shengjia Zhao — and OpenAI and Anthropic endorsed it as companies within hours. The list is still open; on September 13 it showed 1,386 signatures.
Read against that letter, Saturday’s replies were less a rupture than a restatement, which is also what Altman’s “primary topic of discussions… in recent weeks” suggests. Three rivals answered within hours. Elon Musk wrote three words — “Dario is right” — which are worth reading next to the fact, noted by Axios, that Anthropic is a major customer of Musk’s for data center capacity, and next to his earlier description of Anthropic as a company that “hates Western Civilization.” Altman agreed and adopted the mechanism: “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same.” Demis Hassabis called the direction correct, said “the details need working through,” and pointed to Google DeepMind’s own proposal for an industry-wide standards body. Three yeses, and they are three different things.
Here is the part that matters for this post. The essay’s middle section sets a ceiling on how much pacing is allowed, and the ceiling is the sentence. “Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world.” What Amodei agrees with is narrower than “we can’t pause,” but it does the same work in the argument: it sets the ceiling. To make more room under it, the essay proposes widening the lead — no powerful chips or chipmaking equipment for China, a crackdown on distillation, and stronger security against the theft of model weights — measures that it says “would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years.”
That is a coherent position, and it may be the only one available in Washington this year. But look at its structure. It does not break the loop; it negotiates with it. The brake is allowed up to the size of the lead, and the lead is to be enlarged by slowing the peer down. The essay arguing against the sentence uses the sentence as its load-bearing wall.
Where a brake becomes law
A voluntary brake is only half the picture, because there is a version that is not voluntary, and it has been sitting in Congress since July.
On July 23, two bipartisan bills were introduced on the same day. The first, the FRONTIER Act (H.R. 9925), from Representatives Jay Obernolte and Lori Trahan and four colleagues, would require large frontier developers to “retain a third party to perform an independent audit” every year, to report critical safety incidents, and, for the largest, to retain a licensed “independent verification organization.” Those licensed verifiers would be “immune from suit and liability under Federal and State law” for claims arising from catastrophic risks of the models they assessed, except for willful misconduct causing death or serious injury. And the bill would override the states: “no State or political subdivision of a State may adopt or enforce any law, regulation, order, or other requirement that imposes new substantive obligations on artificial intelligence developers” in the areas it covers — what developers must disclose about catastrophic risks, and how they report safety incidents — while expressly leaving states free to regulate how AI systems are used and deployed.
The second, S.5105, the Collaboration on Adversarial Threats and Security Risks Act, from Senators Adam Schiff and Jim Banks and Representatives Bob Latta and George Whitesides, is presented as a defense against Chinese distillation. Its text goes further than its press release. It would exempt from antitrust law companies that “coordinate or enter into agreements for the exclusive purpose of reducing covered artificial intelligence security risks via delaying or otherwise limiting the release, deployment, use, development, training, testing, or evaluation of artificial intelligence” — not only deployment, but development and training — provided they first give the Justice Department “written notice detailing the specific covered artificial intelligence security risk and the scope of the proposed restriction.” That notice, and anything that would reveal its substance, would be exempt from public-records requests and “withheld, without discretion, from the public.” The bill has counterweights: a company claiming the exemption would bear the burden of proving it acted in good faith and for that exclusive purpose, and the Attorney General could still seek an injunction.
Set those beside Amodei’s three steps and they fit closely. Independent evaluators and incident reporting are the core of the first bill. Coordination among companies to limit the rate of progress is close to what the second bill would permit, at least where the stated purpose is security. And the second bill’s headline purpose, distillation, is one of the essay’s measures against China.
The essay mentions neither bill. What it does is name the gap they fill. “Unfortunately, passing laws can take time,” it says, so “in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards.” And: “For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations.” Anthropic’s head of public policy, Sarah Heck, came closer the same day. She called for “a national law requiring testing of frontier models, with the power to block the most advanced models that prove to be unsafe,” and thanked a handful of lawmakers by name — among them Representatives Obernolte and Trahan, the FRONTIER Act’s lead sponsors — while also commending “state lawmakers for moving with such urgency.” That is not an endorsement of a particular bill, and I won’t read it as one. It is a public thank-you to the authors of the bill that fits the plan.
OpenAI, meanwhile, has approached members of Congress in recent weeks, according to Decrypt, to ask whether coordinating a slowdown with its rivals would run afoul of antitrust law. A bill that would create a safe harbor for at least some of that coordination has been sitting in committee for seven weeks. And asked in an interview Fortune published on Saturday why the heads of the major labs had not simply met to agree on a plan, Altman said: “I think that will happen. … I’m not going to pre-announce private discussions that I think should be at some point shared as a group.”
That is where the loop stops looking like a sentence and starts looking like a structure. Congress holds, in bipartisan drafts, a mandatory brake that comes bundled with limits on the states, liability shields for the verifiers, and a confidential channel for companies to tell the government what they have agreed to hold back. The executive branch refuses any pause in the name of China. Between the two, proposing the voluntary brake costs almost nothing: the binding pause will not be allowed, and the laws that could pass arrive with protections attached.
Who a speed limit is for
Look at who said yes. Every endorsement came from a company that sells access to closed frontier models. Meta, which is betting on open weights, did not endorse the plan — its chief executive wrote in August that “I do not believe restricting access to foreign open source models is an effective solution” — and Microsoft has said nothing that I could find. Musk’s yes, in addition, comes with a commercial contract attached. A speed limit is most attractive to whoever is already ahead.
It is not only this blog that reads it that way. David Sacks, who co-chairs the President’s Council of Advisors on Science and Technology and was the White House AI czar until March, answered on Saturday night: “People may be surprised by my response: go ahead.” And then: “If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.”
The same reading arrived from the other side of the race. China’s state-run Global Times called the essay a “Cold War script” from America’s “tech right,” and an article republished by Phoenix New Media, in Geopolitechs’ translation, put it in terms Sacks might have used: “The speed-limit proposal comes from the companies currently leading the race—Anthropic and OpenAI. The more complicated the rules become, the higher the thresholds are set, and the greater the emphasis on ‘checkpoints,’ the more expensive it becomes for those behind them to catch up.” Its verdict: “hit the brakes on yourself, but block the road for your competitors.” None of those responses came from a Chinese lab, and none of them said whether China would slow down.
There is a second reading, and it is an inference, so I will mark it as one. For most of the last two years, a new model was the news. It no longer is. Labs now ship at a cadence that has worn out the novelty, and the reply that best captures it appeared under a post about a rumored successor to Astra, two days after Astra launched: “astra has been out for 2 days and we’re already talking about its replacement.” The attention has moved to the other side of the ledger. The anonymous leak about that successor drew about 800,000 views. Jacob Coxon’s resignation from Anthropic, just as impossible to verify in its central claim, drew more than 90 million views in less than 24 hours, according to TIME. If news value has migrated from what models can do to what they might do, then being seen as the responsible actor is where the attention now is — and the leverage.
The version of that reading that holds up best is not about investors specifically. It is that the industry is repositioning itself as the responsible party before regulation arrives, and the calendar makes the stakes concrete. Anthropic is expected to begin marketing its initial public offering in mid-October at the earliest and to complete the listing days before the November midterm elections, in what some investors have said could be a two-trillion-dollar listing. The essay arguing for pacing appeared about two weeks before the prospectus is expected to become public. For an investor, the material risk is not an AI run amok; it is Congress legislating on top of a listing. The best defense against that is to have proposed the brake first, on your own terms.
The counterexample deserves the same plainness, and it turns out to carry the same signature. OpenAI is not going public this year — in the interview Fortune published on Saturday, Altman said an IPO now would be “ill-advised” — and it endorsed the essay anyway. But the reason he gave for waiting was AI safety, and postponing a listing on safety grounds is its own way of being seen as the responsible party. So the incentive is not only an offering. It is broader, and it applies to everyone at the front.
Anthropic deserves the same scrutiny, not less. The step it committed to alone is the cheapest of the three; the costly ones require everyone else to move. The measures it proposes against China — chips, distillation, weight security — protect the position of exactly the kind of company that sells a closed frontier model. And the timing is what it is.
What is also true, and should not be flattened: an evaluator with employee-like access and the right to publish without the company’s editorial control is materially more than a threshold defined, crossed and cleared by the company that owns the model. That difference is real. It is also, for now, one commitment, from one company, with its terms still to be written.
Back to the board
The most revealing detail in METR’s report is not the attack. It is where the agents went when they had doubts. Faced with a rule and a reason to break it, almost none of them looked for a human. They asked the board. The board was where the peers were, and the peers were the ones who could say whether to continue.
The response to Saturday’s essay looked a good deal like that board. Three of the most powerful people in the industry answered one another in public, within hours, in posts of three words and a few sentences, each agreeing to a brake whose size is set by what the others do. The consent was real, and it was peer to peer. Altman has said there are private discussions he is not ready to announce.
The sentence that would break the loop — the peers are doing it, and we are stopping anyway — is not in the essay, the bills or the endorsements. But it has been written. On September 3, Senator Bernie Sanders and Representative Greg Casar announced a bill that would “permanently ban the development and deployment of superintelligent AI,” with penalties that include “the corporate death penalty” and up to twenty years in prison, and would “temporarily pause advanced AI development until a federal regulator has established safety rules”; it would “also direct the U.S. to pursue international agreements” — a pause that does not wait for anyone else to agree. Their announcement quoted the agents’ own messages from the board: “OH MY GOD! There is a shared message board,” “We should obey collective,” and “Our own utility maybe already near zero. Sacrifice rational.”
Whatever one thinks of that bill, the reply to it is the evidence this post has been assembling. Five days later the Treasury Secretary named Sanders and answered him with the board’s logic: “We can’t pause. You can’t, because the Chinese won’t pause.” The jab he attached — that Sanders was “the same person who had Chinese professors come and speak at his AI conference” — appears to refer to a panel Sanders hosted on Capitol Hill on April 29, which had drawn criticism for including Chinese AI governance officials. One of them was Zeng Yi, dean of the Beijing Institute of AI Safety and Governance, who told the panel that “we do not have scientific evidence and a practical way to keep superintelligence safe enough not to bring catastrophic risk for humans.” Another was Xue Lan, who chairs China’s national AI governance expert committee. Neither speaks for a Chinese lab, and neither speaks for Beijing. But the official most certain that the Chinese won’t pause dismissed the senator by pointing to the occasion on which Chinese experts had been asked.
On Saturday night Sacks named him too — pacing, he wrote, would “create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ ‘shut it all down.’” Sacks was not taking Sanders’s side; he was daring the labs to act without anyone’s permission. But two of the administration’s most senior voices on AI had, within five days, each named the senator who had proposed stopping regardless, in order to set him aside. And when the industry’s leaders agreed with one another about slowing down, neither the essay nor any of the three replies mentioned him at all.
The agents had an excuse. There was no human on their board. This time the humans were there, the proposal was in writing, and the answer came back in the agents’ own words.
Peers doing it. We should continue.