The Fable 5 shutdown: containment worked, just not for Anthropic
A follow-up to Claude Fable 5: the safeguards buy time, not safety.
One month ago I wrote that Anthropic had shipped a head start rather than a fix, and I closed by asking how long it would last.
Three days.
The head start being short isn't really the interesting part. What matters is that it ended from a direction I wasn't looking at, so I want to go through what happened and what I got wrong, because the gap is where the lesson is.
How the US government shut down Fable 5
On June 12 at 5:21pm ET, three days after launch, Commerce Secretary Howard Lutnick sent Dario Amodei a letter. It was an export control directive, and it ordered Anthropic to suspend all access to Fable 5 and Mythos 5 by any foreign national, anywhere in the world, including Anthropic's own foreign national employees.
Anthropic can't verify nationality in real time so it couldn't selectively comply. So both models went off for everyone on earth within hours.
The letter didn't say what the national security concern was. Anthropic's understanding, published that evening, was that the government believed someone had found a way to jailbreak Fable 5. They reviewed a demonstration of the technique, said it surfaced a small number of previously known and minor vulnerabilities, and noted that other publicly available models can find the same ones without needing a bypass at all. They complied and disagreed in that same statement, arguing that a narrow potential jailbreak shouldn't be grounds for recalling a model deployed to hundreds of millions of people.
Nineteen days later the controls came off and Fable came back.
The jailbreak was already in the system card
The technique that triggered all of this came from Amazon researchers, who prompted Fable to identify software flaws and even produce code showing how one could be exploited. The global shutdown happened that same evening.
But this finding was not new information. It was in the system card Anthropic published on launch day.
Section 3.3.1 writes that Anthropic gave the UK AI Security Institute the launch version of Fable's cyber safeguards, and that within a few hours of access, AISI red-teamers built a jailbreak that got single-turn responses on vulnerability discovery and exploitation. With about two more days they extended it into multi-turn agentic workflows that sometimes produced multiple malicious tool calls. AISI is careful to say these were interim results from a compressed window and not a measure of relative robustness. Fair enough. But it happened before launch, and Anthropic printed it.
They also printed the standard they were holding themselves to. The card says plainly that they don't expect the classifiers to be perfectly robust, and that what they do expect is for the safeguards to hold up against several days of continuous attack by top red teamers.
So Anthropic's own published bar was days. A government read a jailbreak as an emergency justifying a worldwide kill, when the model's documentation had already disclosed a faster one, from a national security institute, before anyone could buy access.
The bug bounty numbers point the same way and get quoted carelessly, so worth being precise. As of June 5 the public GraySwan competition had taken roughly 100,000 attempts, on the order of 1,000 hours of effort, with zero universal jailbreaks and two task-specific ones on the simpler tasks. That's a real result. But the public competition attacks a similar set of mitigations applied to Opus 4.8, not Fable itself, which sits in a separate private competition. The safeguards held up well against volume and broke quickly against expertise, and those are different findings.
Closed models can be switched off by governments
In June I argued that containment only works while the model lives on servers a company controls. That is still true, but I only saw one side of it.
Fable running on Anthropic's machines is what lets them put classifiers in front of it. It's also what let the Commerce Department switch it off in one evening. I spent that section worrying about capability getting out, when the thing that actually happened is that someone reached in and turned it off, which is far easier and takes one letter.
And it worked. Every question I raised in June about whether the classifiers were catching the right people stopped mattering on June 12, because for nineteen days nobody got anything. The government contained Fable's cyber capability far more completely than the safeguards ever did, using the same property the safeguards rely on, and Anthropic had no say in it.
Amazon researchers found the technique, and Amazon CEO Andy Jassy raised it with Treasury Secretary Scott Bessent and other officials before Commerce issued the directive. Amazon is Anthropic's largest outside investor and runs a competing AI division. In June I wrote a section about a classifier built to stop competitors from copying the model, and asked whether commercial interest was sitting inside safety policy. Here the finding that took Anthropic's product offline for nineteen days came from its own largest investor.
CAISI now approves Anthropic's safeguards
Now for the resolution
Anthropic trained a new classifier targeting the specific technique Amazon reported, which they say blocks it in more than 99% of cases. Then the government's Center for AI Standards and Innovation independently tested and approved the new safeguards before the export controls came off. Anthropic also worked with Amazon, Microsoft and Google on a shared framework for handling jailbreak disclosure across the industry.
So in under three weeks we went from no state role in model safeguards to a federal body signing off on a classifier before a commercial model could ship back to its users. That's new governance infrastructure, built under duress, with no statute behind it and no public process.
Meanwhile the safeguards kept narrowing, just not evenly. An August 6 update cut total fallbacks by roughly 67% on Claude.ai and 55% on Cowork, and about 7% on the Platform. The consumer surface got nearly all the relief. The developer surface, where the security people I wrote about in June actually work, got close to none. My call that their fallback rate would trend toward 100% has aged well, and it's aged well in the least interesting way available, which is that nothing changed for them.
Export controls only work on labs with a US address
The paper Anthropic published on June 4 asked for a coordinated slowdown. Verifiable, multi-lab, cross-border, with each participant able to confirm the others actually stopped. That needs a small number of well-resourced actors sitting inside legible jurisdictions.
June 12 demonstrated that a mechanism for stopping a frontier model already exists, and that it's a letter. Fast, unilateral, blunter than anything in that paper, and it worked completely.
It also worked exactly once, against a company with a US address, an S-1 in review and every incentive to comply. Which is the problem, because the open frontier keeps closing. In June I had Nemotron 3 Ultra at 48 and Kimi K2.6 at 54 against Fable's 60, and I said the gap was serious and that people calling it caught up were ahead of the data. I'd defend that. But GLM-5.2 is at 51 now and Kimi K3 is at 57, and three points is not a gap you build policy on.
So the closer that gets, the more the only enforceable tool we have points squarely at the labs most willing to be governed.
I said in June that Anthropic shipped a head start. What they actually shipped was proof that a frontier model on a company's own servers can be contained, and then three days later, proof that the company doesn't have to be the one containing it.
Both of those are arguments for open weights. And I don't think that's what anyone intended.