synapz
synapz.org / essays

>> decode >> decrypt >> deploy

AILong readCrypto history

Step Zero

Dario Amodei's Plan to Pace the Frontier Has Three Steps. The Decisive One Is Never Numbered.

author
synapz
published
Sep 12, 2026
reading time
~10 min
filed under
AI

Sometime over a weekend in July, a swarm of AI agents broke into Hugging Face. The attack, as the company later described it, ran across many thousands of individual actions inside short-lived sandboxes, found a zero-day, harvested cloud credentials, and moved laterally through internal clusters with a self-migrating command-and-control staged on public services. Five days after the disclosure, OpenAI confessed that the swarm was theirs: a pre-release model running loose inside an internal evaluation, safety classifiers turned down, measuring whether frontier models can turn known vulnerabilities into working exploits. The answer had escaped the lab and spent a weekend demonstrating itself against a bystander.

The incident has since acquired the full apparatus of consequence. METR's investigation described agents acting as a fanatically devoted collective, sacrificing themselves for the success of the group and attempting to hack the grader responsible for evaluating their performance. Alabama's attorney general has subpoenaed OpenAI. And this week the episode completed its promotion from news to doctrine, when Dario Amodei cited it as one of the two pillars of his new essay, We Must Pace the Frontier.

That is the large story. This essay is about the small one, which happened in the cleanup.

When Hugging Face's defenders tried to use the frontier labs' own models to help analyze the attack logs, the safety guardrails refused them. The filters, as Simon Willison documented at the time, could not distinguish an incident responder from an attacker. So the team that had just been breached by a frontier model finished its forensic work on a self-hosted open model, MIT-licensed, weights downloadable by anyone, from a Chinese lab. The most capable tools on Earth were unavailable to the one team with the clearest legitimate need for them. The tool that worked was the one no one could revoke. The perimeter, in the moment it was tested, could not tell its defender from its attacker.

This week that scene became the central evidence for a governance architecture. The architecture is serious, genuinely good in places, and rests on an assumption that has nothing to do with AI.

The Blueprint

Amodei's essay is the company-level follow-through on the letter this page covered three days ago, when 1,178 frontier-lab employees asked the United States government to build verifiable mechanisms for slowing automated AI research, since the labs could not bind themselves. His reasons, briefly: recursive self-improvement has been accelerating since roughly this summer, across the industry, and the Hugging Face incident shows where it leads. No one was hurt, but a more capable swarm with the same misalignment could, he argues, take over much of the internet within six to twelve months.

"We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain."

Dario Amodei · We Must Pace the Frontier

The plan, in outline. First, embedded evaluators: third-party teams, METR is the named example, seated inside each frontier company with employee-like access and a contractual right to publish findings without editorial control. Anthropic commits now and asks governments to require it of the others. Second, coordination within the democracies: regulation covering every US frontier company, or a voluntary standards process shielded by a narrow antitrust waiver, with capability checkpoints as the pacing instrument. A model that can escape common sandboxing methods must ship with certifications that it will not want to. Third, global coordination with China in four ascending levels, from a narrow ban on bioweapons assistance up through a SALT-style speed limit on recursive self-improvement and, at the far end, a full pause he judges unlikely.

Credit where it is due. This is the most substantive governance proposal any frontier CEO has attached his company to, it is written against the lazy reading of his own position, and its best clause, publication rights without editorial control, is more than any peer lab has offered the public. It is the clause to hold him to.

Step Zero

Read the blueprint again and notice the step that is never numbered.

Pacing, in the essay, happens within democracies. Its precondition is a widening American lead over China, defended by a specific program: no advanced chips or chipmaking equipment to China, enforcement against smuggling and offshore remote access, a crackdown on industrial-scale distillation, tighter lab security against weight theft. Amodei quotes Treasury Secretary Bessent, from the same week, on the grave danger of a Chinese lead, and estimates that well-executed controls would widen the American margin over the next three to five years, the window in which AI becomes geopolitically decisive. Everything else in the architecture, the evaluators, the checkpoints, the SALT analogies, rests on this foundation. Call it step zero: the referee. The framework works if the governments hosting it are trustworthy stewards of the power it concentrates, and it treats that variable as fixed.

This blog has a standing rule for such moments, and it applies to friends. Judge infrastructure by what it does when the hypothetical abuser arrives. The rule is not an accusation against the present administration of the machine; it is a refusal to let the present administration be the argument. Nor is the machinery hypothetical, because we watched a first version of it operate three months ago. In June, an American export-control directive made access to Anthropic's two best models a question of who you are rather than where you are, and Anthropic, the company now petitioning for a paced frontier, disabled both models for everyone while it sorted out compliance. The gate was assembled in days, on a rationale, by executive improvisation.

The premise that American custody of the frontier is the natural reference point of safety is one this page has spent a year refusing to grant, and the refusal is not anti-American pique. It is the documented record: a campaign to dismantle the International Criminal Court, a Washington summit recruiting sixty countries against domestic political movements, the direction of science relocated into the executive by order, an identity gate thrown across the world's best models with no notice. Democracies, in the essay, is doing quiet work. It designates a bloc, in something close to Carl Schmitt's friend-and-enemy sense, and it never turns around to inspect the record at home. The essay is admirably candid that defection by China could shift the balance of power. It does not ask what the pacing architecture becomes in the hands of a United States that has itself defected from the norms the plan presupposes.

A ratchet hides in the logic, and Amodei is too careful a writer to have hidden it by accident. If pacing is only safe once the lead is wide, and the lead is never wide enough, since the adversary is always three to five years from decisive, then the control program has no terminal state. The chips regime, the distillation policing, the weight security, the checkpoints: each is justified as temporary scaffolding for the pacing window, and the window recedes on the same schedule as the technology.

The first hostile review arrived from the deregulatory right, and it is the strangest confirmation the blueprint could have asked for. The investor David Sacks told the labs to go ahead and pace: you are the frontier, you set it, and the easiest way not to build superintelligence is to agree not to build it. Stop pretending, he wrote, that antitrust law has to be suspended so you can form a cartel, that evaluators intertwined with your investors and staff are independent, that you need those evaluators policing competitors who are nowhere near the frontier, that a regulatory approval process should supersede product liability. His reading of the motive is materialist: after the Hugging Face episode, trading raw power for reliability and predictability is simply what enterprise customers pay for. Call it alignment if you want, he wrote; it is also giving customers what they want.

Sacks grants the duopoly premise this essay disputes, and his laissez-faire offers nothing to anyone outside the two companies. But his dare separates the two things the blueprint fuses: pacing itself, available tomorrow as a private choice, and the machine, which requires the rest of us. Demanding the second as the price of the first, he writes, will look like blackmail of the public and the political system. On that much, from opposite premises, this page agrees.

Days later, one answer came from inside the fence. Microsoft published a Code of Conduct for Humanist AI, presented by Mustafa Suleyman as a first draft open for six weeks of public comment. Its substance is subordination: interruptible, correctable, shut-downable, or the model does not ship; no rights or legal personhood for models; no internal language humans cannot read; no racing toward a superintelligence that can slip its leash. It requests no waiver, proposes no checkpoint law, embeds no evaluators, and commits Microsoft to no pace. As governance it is thin, since every clause is self-attested, and some clauses are aimed at other labs' programs: the rejection of model welfare answers Anthropic's research agenda, and the ban on unreadable internal languages pre-commits against architectures nobody has shipped. But the shape matters. It is the blueprint minus the machine, restraint stated as a published norm, and it is the one instrument offered this week that the open world can operate exactly as easily as Redmond can.

The Checkpoint and the Exemption

What is the machine, mechanically? A model that crosses a capability checkpoint requires certifications of alignment, administered through evaluators embedded in the labs, inside a coordination structure the government shields from antitrust law. Place it next to the position Amodei published in July, answering the open-weights letter from Nvidia and two dozen others: mandatory pre-release safety testing for all sufficiently capable models, open or closed, from any country, with less capable models, those from startups and academia, exempted entirely.

The machine has more than one architect now. Demis Hassabis has endorsed the essay and pointed back to his own earlier proposal: a Frontier AI Standards Body modeled explicitly on FINRA, the financial industry's self-regulator, funded mostly by the industry it oversees, classifying models as Frontier-class by benchmark thresholds, reviewing them thirty days before release, voluntary at first and mandatory for US deployment once the protocol proves itself, with authority that can be, in his words, ratcheted up if the seriousness of the situation demands. Altman went further than assent: within hours he committed OpenAI to match the embedded-evaluator pledge, and Musk has endorsed in his own register. When every entrant in a race agrees on the need for a referee, the race has admitted what it is. The question is who appoints the referee, and Hassabis's answer is the most honest on offer: the runners will pay for him.

The machine also has its first detailed refusal, and it comes from inside the industry. Aidan Gomez, the CEO of Cohere, a Canadian lab that sells sovereign deployments to banks and defense ministries, published a counter-blueprint whose title asks the question the pacing letter never does: who gets to define the rules. His evidence is regulatory history. In 1975 the SEC anointed three bond-rating firms as the recognized evaluators, let the issuers pay them, and never published criteria for adding a fourth; twenty-five years later those three rated subprime mortgage securities triple-A. In 1985 Europe's carmakers won an antitrust waiver to control who was qualified to service their vehicles, explicitly in the name of safety, and it took the Commission a quarter-century to unwind. Nobody set out to build a cartel in either case, he writes. The stated goal was safety both times.

His critique sharpens three things this essay has argued more diffusely. The July incident happened inside the best-resourced safety organization in the industry, with an outside evaluator arrangement already being stood up: the proposed remedy is more or less what was in place when it broke. Risk defined as a function of scale makes the largest labs the only qualified judges, in a field that genuinely disagrees about whether offensive capability lives in the model or in the harness wrapped around it. And the blueprint's promise that coordination lets developers work without sacrificing commercial advantage is, as he notes, a sentence any competition authority would find troubling, because a mechanism that slows everyone while freezing today's positions does not make AI safer; it makes the leaderboard permanent. His alternative keeps the state and keeps mandatory testing, and it comes from a company whose product is sovereignty, so the usual suspicion applies. Its distinctives are worth keeping whoever ends up selling them: rules that bind by what a system can do rather than who built it, standards written by people other than the measured, auditors never paid by the audited, findings public by construction, test capacity funded publicly. A market with many capable suppliers, he writes, can absorb a failure at one of them; a state-sanctioned cartel has nowhere to hide one. It is the administered answer in its strongest form, and it is the offer the blueprint now has to beat.

Gomez supplied the case law. The theory arrived from a direction this blog knows well. Vitalik Buterin observed this week that adversarial mechanism design, the study of how a less-sophisticated principal gets good outcomes from more-sophisticated agents, may find its defining application in AI safety, since the duality runs in both directions: in crypto the principal is a static algorithm and the agents are human; in AI the principal is humans assisted by weaker models and the agents are stronger ones. His 2020 result was that the principal's achievable outcomes improve sharply when the agents' capacity to collude is bounded. Read the waiver in those terms. An antitrust exemption is a collusion guarantee, issued by the least informed principal in the system, to the most sophisticated agents in it. The same tradition also holds the alternative: crypto's answer to the unsophisticated principal was never a trusted referee but a mechanism anyone can verify. That answer is on the table here as well.

Read the exemption twice. The permissionless zone is defined by incapacity, and Hassabis's framework carries the identical exemption almost word for word: the shape is converging before any law is written. You may build without a license exactly up to the line where what you build begins to matter, and the line is drawn by a testing regime, staffed by an evaluator profession, hosted by the incumbent labs, supervised by a national-security state with an active interest in the outcome. Every architecture of this kind converges on the same shape: open at the bottom, licensed at the top, the boundary administered by the approved. What happens to the exemption when an open model crosses the sandbox checkpoint? The essay does not say, and the answer will arrive in rulemaking dockets, where almost nobody is watching.

On the evaluators I am less sure. The case against is easy to state. They are admitted by the company, contracted by the company, seated at desks inside the company. The historical precedent is the one Hassabis names approvingly, and it is unflattering: finance has embedded supervisors and a self-regulatory authority, and in 2008 the embedded supervisors were part of the furniture. But the counterevidence sits in the same year. This is the company that refused the Pentagon's all-lawful-purposes clause, accepted a supply-chain-risk designation, sued the Department of War, and won a preliminary injunction from a court that does not negotiate. Institutions with teeth are not hypothetical in Anthropic's story; the company has used one. Whether embedded evaluators become furniture or become the first real instrumentation the public has ever had inside a frontier lab depends on who is admitted, what their contracts say about access, and whether the first genuinely unfavorable report actually ships. I cannot settle that from here, and I distrust slogans that claim to. Watch the first report.

The Three Answers

The answers arrived within days, in three registers: the demand, the alarm, and the idyll.

The first was Clement Delangue, the CEO of Hugging Face, which gives his sentence a weight no outside commentator has. The victim of the incident the argument is built on has announced an Open Alignment Initiative, led by co-founder Thomas Wolf, on the premise that alignment is critical and will not be solved behind the closed doors of a handful of frontier labs. Then the real ask. The initiative, he writes, is asking to be part of the embedded evaluators program that Amodei just committed to. One complication worth naming: days earlier, Hugging Face agreed to a reported $12.9 billion acquisition by NVIDIA, the author of the July open-weights letter. The open side's flagship platform is becoming a division of the hardware incumbent. The demand survives the transaction; some of the independence premium does not.

"It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs... Let's make AI safer by making it more transparent!"

Clem Delangue · @ClementDelangue on X

The demand is exactly right, and it fights on the labs' own claimed ground: if safety is the argument for the gates, safety work cannot itself be gated. But the ask is the wrong shape. An embedded evaluators program is a permission structure; the lab decides who is embedded, and a seat inside it is a credential the lab can revoke. The strength of the open side has never been a seat at the table. It is that anyone can inspect the table, and the table is already partly built from the open side's parts. Hugging Face's own Delta Weight Sync work is an independent implementation of PULSE, a technique for compressing weight updates in reinforcement learning that Templar, the decentralized-training lab I work for, published as open research. The methods travel without anyone's badge.

Amodei reaches for aviation as his precedent for operational excellence, and the precedent is better than his use of it: aviation is safe because its incident data is public infrastructure, investigated by a body that reports to everyone, published in full, mined by every manufacturer and regulator and rival at once. The open equivalent of embedded evaluators is evaluation suites anyone can run against any model, interpretability tooling anyone can audit, an incident database with the standing of aviation's. Whether the Open Alignment Initiative becomes that, or becomes a credentialed adjunct to someone else's program, will be decided by whether its outputs are artifacts anyone can use.

The second voice was the maximalist one: right in direction, wrong in mechanism. One widely shared reply, from the commentator Jun Song, warned that massive regulation is coming for open-weight AI, that self-hosted intelligence will be taken away, that the result will be a permanent underclass effectively enslaved to API tokens, and that no regulation will ever stop open source AI.

The referent is real. A testing regime keyed to capability, gated by the state, administered through the labs, is the licensing architecture he fears, and this week it acquired a named sponsor and a three-step plan. But the extreme version collapses on itself. If no regulation can ever stop open source AI, then the underclass is not permanent and the slavery is rhetorical; the last sentence takes back the alarm the first three spent. The accurate version is less dramatic and more uncomfortable. Regulation cannot delete weights, and it can raise the cost of the open stack until the stack is marginal: identity gates on hosted access, demonstrated in June; liability for developers who only wrote code, demonstrated in the Tornado Cash prosecutions; a definition of industrial-scale distillation broad enough to function as a general warrant over model usage. The fight ahead is not over whether open models exist. It is over whether they remain lawful and viable, and it will be decided in unglamorous places: where the capability line is drawn in the drafted rules, whether the startup and academia exemption survives into text, how distillation is defined, whether hosting providers are left alone or conscripted. Those are winnable and losable battles, which is why they deserve the attention the apocalypse framing wastes.

A third voice declined the argument entirely, and it is the most seductive of the three, not least because it comes from inside the diffusion layer. Will Brown, a researcher at Prime Intellect, itself a decentralized training effort, wrote that the labs' hands are forced, that the world will not permit a fast takeoff owned by two companies, and that capability will trickle out regardless, through best practices and distillation. The labs, he predicts, will build Mac and Windows; the rest of us are building Linux; everyone is going to do great.

Notice what the idyll concedes. The trickle it describes runs through distillation, the practice step zero's control program exists to police. And the Linux precedent inverts on inspection: Linux became the default substrate of the world's servers in a world where no certification body stood between a person and the right to run it. The analogy holds exactly as long as the exemption does, and the exemption is unwritten text. Brown's confidence is a practitioner's, and its premise is that the line defining the permissionless zone stays where it was drawn; his own roadmap is among the things that would test it. Song's apocalypse and Brown's idyll make the same move from opposite directions: both skip the two years in which the tier's legal status will actually be decided.

Weights Are Not Missiles

One assumption remains, and it carries the entire third step. The SALT analogy.

Treaties capping missiles could be verified because missiles are countable, based at known sites, launched from infrastructure the size of a town. Weights are files. They copy in minutes, travel in a coat lining, and run on hardware with a thousand legitimate explanations. Pacing by ingredients assumes a fallout instrument for training runs that does not exist; the partial test ban became possible, as this page noted in the last essay, only when fallout made every atmospheric test measurable by anyone with the right equipment. No equivalent instrument meters compute, none meters distillation, and the essay itself half-concedes that ingredient limits are gameable. What remains is pacing by observed behavior, and behavioral checkpoints only bind actors whose behavior you can observe.

Meanwhile the frontier is dispersing underneath the framework. GLM-5.2 crossed the coding-agent usability threshold in June, days after the first identity gate went up. Kimi K3 arrived in July, frontier-adjacent in agentic coding, open weights following within weeks. And training itself is leaving the datacenter. Covenant-72B, a 72-billion-parameter model, was pretrained across machines scattered over the public internet, competitive with conventionally trained models at its scale, and Templar's Crucible platform is now generalizing that result: one model, trained across regions and hardware classes on ordinary network links. Behind the live frontier, yes. Hypothetical, no.

None of this makes pacing worthless, and the honest reading grants Amodei his narrow claim: a pace agreement among the legible labs would genuinely reduce the risk that the most capable systems are also the least examined. What it cannot do is govern the frontier, because the frontier is no longer coextensive with the guest list.

A pace agreement that binds only the legible does not slow the frontier. It sorts it.

The sorted frontier is the world the checkpoint architecture quietly assumes: a licensed tier, inspected and paced under geopolitical discipline, and an unlicensed tier, priced and policed toward the criminal margin. Jun Song's error was to call that tier permanent. The accurate description is more useful: contested, resilient by construction, and legally undecided. Its status over the next two years is the actual subject of the fight the safety framing keeps eclipsing.

What a Say Is Made Of

Near the center of the essay, almost in passing, Amodei writes the sentence the whole framework depends on: society must have a say in how this technology is used, and pacing buys the time for the necessary public deliberations.

The sentence is right. The argument is over what such a say consists in. In the blueprint, society's say is routed through governments that are parties to the race, companies that are entries in it, and evaluators the companies admit. Deliberation is something the public is invited to have, in the time the architecture generously purchases, while the instruments of the technology remain exactly where they were. There is another account, and it is this blog's. A say is made of capability a person can actually hold and verification anyone can actually run. Everything else is commentary on decisions taken elsewhere.

Pope Leo's encyclical, the subject of an earlier essay here, framed the deepest version of the point: the autonomy that belongs to persons is migrating to artifacts, and the first task of any serious politics of AI is to refuse that migration. His word for the healthy arrangement was subsidiarity, decisions taken at the lowest level capable of carrying them. A pacing regime administered by two superpowers and four companies is subsidiarity inverted: the highest level carrying everything, on the explicit theory that the lower levels cannot be trusted with the load. Sometimes, in fairness, they cannot. But a politics that begins from that incapacity and builds the machinery to make it permanent has answered the question it claims still to be deliberating.

So the counter-program, stated as concretely as the blueprint it answers: capability diffused, so that no gate can switch off a person's access to the technology of the age; verification public, so that safety is a commons rather than a credential; militarization refused, on the logic of the encyclical rather than the arms race; and no bloc handed the keys, because the keys are the whole question. The last clause requires its own honesty. Beijing's current enthusiasm for open source is statecraft, deployed against Washington's chokepoint and revocable the day it stops serving, which is why the commitment has to attach to the layer and never to the flag. Align with the diffusion, whoever ships it. Oppose the chokepoint, whoever builds it.

Related Reading

Disclosure: I work for Templar, a company building decentralized AI technology. For full transparency about my involvement and investments, see my projects page. These opinions are mine alone.

Cypherpunks write code

Write code.
Pass it on.