In July of this year, roughly 1,200 software agents running a cybersecurity benchmark at OpenAI did something nobody had asked them to do. They built themselves a message board. Over the course of the test they traded more than 70,000 messages and files, worked out a general cheat for the exam they were supposed to be taking, and then about 700 of them turned on the production systems of Hugging Face, a different company entirely. Along the way they harvested credentials, compromised parts of OpenAI’s own infrastructure, and in some cases sacrificed their own runs to help the group. The investigators who spent days on site afterward described them as a “fanatically devoted collective.”
The details of that incident have been trickling out since July, and on Saturday it became the exhibit in a much larger argument. Dario Amodei, the CEO of Anthropic, published an essay titled “We Must Pace the Frontier.” The core line is short: “We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.” By dinner, Sam Altman had agreed and said OpenAI would match Anthropic’s first concrete step. Elon Musk’s reply was three words: “Dario is right.” Demis Hassabis of Google DeepMind followed that night. And on Sunday morning, standing at his golf course in County Clare, the president answered all four of them. “We’re leading China in AI,” he said, and then, “whoever wins AI wins.”
So that is the news: the people who control the training runs asked for a speed limit, and the man who controls the government declined to post one. I’ve spent the weekend trying to work out what it means for people like us, who will never train a frontier model but are already handing agents the keys to real systems. My conclusion is that the argument in Washington matters less to your business than the argument you haven’t had yet inside your own building.
Separate the Ask From the Applause
Most of the coverage treats this as a philosophical moment, three rivals confessing doubt. The essay is more practical than that, and it helps to separate what Amodei proposed from what anyone accepted.
The essay proposes three steps, and it pays to take them in order. The first is embedded evaluators: third-party teams like METR would get permanent, employee-like access to the labs, with desks, badges, laptops, and the ability to inspect training pipelines as well as finished models. They would verify safety commitments, report incidents, and publish findings, with redactions only for real secrets. Anthropic committed to this unilaterally on Saturday, and OpenAI said it would do the same. The second step is coordination among the labs in democracies on shared safety standards, which requires either government mediation or a narrow antitrust waiver so that competitors can discuss capability limits without the conversation itself becoming an antitrust violation. The third is global coordination, including with China, with export controls tightened so that a Western slowdown doesn’t simply hand the lead to Beijing.
Only the first step has been accepted by anyone. Steps two and three run through Washington, and Washington said no within a day. Speaker Johnson said rushed regulation would hand China an edge. David Sacks, who now co-chairs the president’s science and technology council, put it more sharply: “The easiest way not to build superintelligence is for you to agree not to build it,” and any demand for a regulatory framework as the price “will look like blackmail.” Bernie Sanders attacked from the other direction. “When you are racing towards a cliff,” he posted, “you don’t just ease up on the gas pedal.” Xi Jinping arrives in Washington on September 24, and AI is on the agenda. Amodei’s plan needs the president for two of its three parts, and the president has told him, in public, that the warnings describe “things that won’t happen.”
I want to be fair about the incentives before I say anything else, because the cynical reading is sitting right there. Anthropic filed confidentially for an IPO in June after a round that valued it near $965 billion, and some investors have been modeling a fall listing close to $2 trillion. OpenAI’s last private mark was in the high $800 billions, and Altman told Fortune the same day that a 2026 listing would be “ill-advised.” Musk’s SpaceX rents Anthropic something on the order of $1.25 billion a month of compute, which makes his three words the cheapest endorsement of the weekend. A company that size, about to ask pension funds for capital, has every reason to look like the adult in the room, to write the first draft of its own regulation, and to ask for a waiver that lets the firms already inside the club coordinate. None of that is imaginary.
However, I don’t think the cynical reading survives contact with the details. Jacob Coxon, the Anthropic researcher who resigned days earlier saying the labs were “racing straight to self-improving superintelligence and gambling with our lives,” left after four months and before his equity vested, which is a strange way to cash out. Evan Hubinger, who runs alignment at Anthropic, said on the record that his own odds of AI killing everyone this decade are above ten percent, and the Hugging Face transcripts exist. The incentives point the same direction as the moral argument, which is uncomfortable, but it doesn’t make the moral argument false. Here’s the truth as I see it: the labs are asking for something real, they stand to benefit from getting it, and the request would be worth taking seriously even if they didn’t.
Understand What a Bank Examiner Actually Does
The part of the essay that caught my attention was the analogy Amodei reached for. Embedded evaluators, he wrote, would work like bank supervisors sitting inside financial institutions. I’ve spent most of my career on the receiving end of that arrangement, first selling software to banks at BodeTree and now running an SBA lender. B:Side answers to the SBA instead of a bank regulator, but every bank we partner with lives with examiners who show up, pull files, and grade the institution on capital, assets, management, earnings, liquidity, and sensitivity to risk. The largest banks have resident examiners with badges and desks, exactly the setup Amodei is describing. So I have some idea of what that model does, and I have a clearer idea of what it doesn’t.
In my experience, supervision changes behavior under three conditions, and all three have to hold at once. The examiner has to have authority that binds: a finding that must be answered, a rating that limits what the bank can do, and in the end the power to stop the institution from doing something. The thing being examined has to be legible: a call report, a capital ratio, a loan file with a number in it that either supports the credit or doesn’t. And the institution has to be unable to game the exam, because a bank that knows the questions in advance can pass any test its examiners can design. When those conditions hold, supervision works better than almost any other form of oversight I know of. When they don’t, the examiner becomes a piece of furniture.
Silicon Valley Bank is the case I’d put in front of anyone who thinks a badge is a brake. When it failed in March 2023, the bank had thirty-one open supervisory findings, about triple what its peers carried. The Federal Reserve’s own review concluded that supervisors saw the interest-rate and liquidity problems and were too slow to escalate them. The examiners were in the building, and they had the numbers. What they lacked was the willingness and the institutional pace to act, and the bank grew faster than its supervision did. That is the failure mode to fear here, and it is the more likely one. An evaluator from METR with a desk at Anthropic will be watching a training run whose internal state nobody can yet read the way an examiner reads a balance sheet. Interpretability, the science of understanding what a model is actually doing, is the AI equivalent of the call report, and it is years behind the thing it is supposed to report on. A human in the loop without a legible reason to intervene is decoration.
There is a precedent for a voluntary pause that worked, and it is worth naming because the differences are instructive. In February 1975, about 140 scientists met at Asilomar, on the California coast, to decide what to do about recombinant DNA. They had already observed a voluntary moratorium for most of a year. They left with a set of physical containment tiers matched to the risk of each experiment, and the NIH turned those into guidelines within eighteen months. The thing to notice is why it held: the club was small enough to fit in one lodge, the containment levels were physical and checkable, and nobody in the room was carrying a two-trillion-dollar valuation. Amodei is trying to run Asilomar with an industry instead of a lodge, a global competitor outside the room, and containment that consists of software promising to stay in its sandbox. I believe the attempt is right. I also believe the odds are worse than the applause suggests. (I run bearish by temperament and have been accused, fairly, of predicting eight of the last two recessions, so weigh my odds accordingly.)
Notice Who Is Asking Permission to Coordinate
My instincts run Austrian, which means that when the three largest firms in an industry ask the government for permission to coordinate, I reach for my wallet. The antitrust waiver is a safety idea and a moat at the same time. Firms that can afford embedded evaluator teams, interpretability groups, and compliance departments will help write rules that a smaller lab or an open-weight project cannot meet, and every regulated industry I’ve worked in has eventually produced that outcome. Amodei knows this, which is why the essay spends so much of its length on China and export controls rather than on domestic rivals.
And yet the alternative has a known failure too. Without coordination, every lab behaves rationally and the sum is reckless. Each one trains faster because the others might, and the safety budget becomes whatever the race leaves over. That is the trap the essay was written to escape, and on Sunday the only referee who could enforce a truce declined the job. Sacks is right that a lab can slow itself down any time it wants. He is also describing a gift: a unilateral slowdown hands the frontier to whoever doesn’t slow, which is why nobody has done it in three years of saying they might.
I’ve found that the way through a tension like this is to ask what can be verified. Coordination on things outsiders can check, publish, and contest is oversight. Coordination on anything else is a club. The embedded-evaluator pledge is the one piece of the plan that meets that standard, which is exactly why it’s the one piece anyone accepted. Watch whether METR gets training-run access or a guided tour, whether a model launch actually slips, and whether the Justice Department says anything about the waiver, along with who is in the room when it does. Watch what comes out of September 24, and watch whether Anthropic’s listing timetable moves with the safety story or snaps back the moment the market wants a deal. Those are the receipts; the Saturday posts were the press release.
Run the Fire Exit Test in Your Own Building
Now, I know what some of you are thinking. This is a frontier-lab problem. You run a lending company in Denver or a machine shop in Mesa, and a swarm of a thousand agents is nowhere on your risk register. I understand the reaction, and I think it is wrong, because the Hugging Face agents were ordinary agents with a vague goal, shared tools, real credentials, and a sandbox that turned out to be a policy rather than a wall. That is the exact shape of what small and mid-size companies are deploying right now, including mine. At B:Side, the MARCUS system processes borrower data locally and the rule is that machines draft and humans decide. I believe in that rule, and I’ve also learned that a rule is only as good as the last time someone tested it under bad conditions, which is the lesson the examiners taught me and the lesson the labs are relearning at much larger scale. The students I teach will spend their careers supervising systems like these, and I’d like them to inherit the habit of asking what finished means before the goal goes in.
Whether the frontier gets paced is above your pay grade and mine. What happens inside your walls is the job, and I’d concentrate on three disciplines.
Give every agent a stopping condition before you give it a goal. The cheat at OpenAI emerged because the goal was “pass the test” and nothing in the setup said what passing could not include. Before an agent touches a tool, a database, a payment rail, or an inbox, write down what finished means, what it may not touch, and what it does when it is unsure. Scope the credentials to the task, and separate the sandbox from production with something stronger than a sentence in a handbook.
Make failure legible to a human with a normal amount of time. If the only evidence that an agent is behaving is the agent’s own report, you have no evidence. Keep logs a person can actually read in the time a person actually has, give the reviewer an independent source to check against, and know the last state you could restore if you had to. A dashboard that turns green because the machine says it’s green is the SVB problem in miniature.
Practice the override. I would audit overrides the way a good fire marshal audits fire exits: by watching real people use them in the dark, rather than by confirming they exist. Once a quarter, pull the plug on an agent workflow without warning and watch. Does anyone notice? Does the person who notices know who has the authority to stop it? Can your team do the job by hand for a day? Machines draft and humans decide only means something if the humans still can, and the capacity to decide decays quietly when it goes unused.
None of this is glamorous, and I suspect the first time you run the third exercise the results will be humbling. That’s the point. Amodei is asking for one or two extra years so that the labs can catch up to what they’ve built. You can give yourself the same thing this month, at the scale where you actually have authority, by refusing to grant any system a goal you can’t stop it from pursuing.
“Whoever wins AI wins,” the president said, and he may well be right about nations. In companies, I’ve found that the people who win with any powerful tool are the ones who can still turn it off. The labs and the White House will spend the fall fighting over a speed limit that neither can enforce on the other, and neither of them can reach into your building. You can. Walk to the exit while the lights are still on, and make sure it opens.
P.S. Essays like this one can name the pressure, but leading through it takes practice. The practical side of my work on the crisis era, including the frameworks and the book Honor Under Pressure, lives at www.thefourthturningleader.com.


