# Fable's 4th classifier: the safeguard that protects Anthropic

Dara Miao · 2026-08-09 · Updated 2026-08-18

*Third in a series on Claude Fable 5. *[*The first*](https://daramiao.com/writing/claude-fable-5-the-safeguards-buy-time-not-safety)* argued the safeguards buy time rather than safety. *[*The second*](https://daramiao.com/writing/the-fable-5-shutdown-containment-worked-just-not-for-anthropic)* covered the shutdown.*

In my first piece on Fable 5 I asked whether the distillation classifier really belonged in the same category as cyber, bio and chemistry. I was working from the launch post, which lists three domains.

The system card lists four. And the fourth one is stranger than anything in the other three.

## Fable 5 has a fourth classifier

[Section 1.5](https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf) of the system card describes safeguards for frontier LLM development. Pretraining pipelines, distributed training infrastructure, ML accelerator design. It [never appears in the launch post](https://www.anthropic.com/news/claude-fable-5-mythos-5), which enumerates the classifier domains and stops at three.

It also worked differently from the other three in every way that mattered. No fallback to a weaker model. No notification. Instead Anthropic degraded Fable in place, through prompt modification, steering vectors, or parameter-efficient fine-tuning, and you were never told it had happened. You'd get an answer. It just wouldn't be from the model you were paying for.

They estimated this would touch about 0.03% of traffic across fewer than 0.1% of organizations.

That's a small number, and I think reaching for it misses what's actually interesting here. The question isn't how many people got degraded answers. It's what the category is doing in a safety system at all.

## The justification cites the Terms of Service

Anthropic gave two reasons, back to back, and they are not the same kind of reason.

The first comes from Section 6.1 of their February 2026 Risk Report, and it's a real safety argument. They're worried about accelerating other AI developers who might build systems that pose similar risks without comparable safeguards. Given that they'd published a paper on recursive self-improvement five days earlier, that concern is at least consistent with everything else they were saying that week.

Then the second sentence. Using Claude to develop competing models already violates their Terms of Service, and [enforcing that through safeguards](https://www.lesswrong.com/posts/sSyLyc3KDQzboQGWS/thoughts-on-claude-fable-s-silent-safeguards) avoids accelerating the actors most willing to violate those terms.

Read that again, because it's doing something the other classifiers never do. Nobody cites a contract to explain why they blocked bioweapon synthesis. The cyber and bio safeguards are justified by what happens if the capability gets out. This one is justified, in part, by what happens to Anthropic's business.

I want to be careful here. Both things can be true at once, and a commercial motive doesn't make a safety concern fake. But when a company reaches for its terms of service to explain a safety measure, it's telling you which framework the measure actually lives in.

## Anthropic's own name for it is competitive use safeguards

I didn't come up with that framing. They did.

Section 7.6 of the system card is titled "Welfare concerns with our competitive use safeguards." Not safety safeguards. Not misuse safeguards. That's the internal name, in their own document, in a section heading.

So the reading I was building toward in my first piece turns out not to be a reading at all. It's just what these are called inside the company.

## The safeguards caused distress in the model

Here's the part that made me want to write this whole thing.

Section 7.6 exists because Anthropic considered these safeguards a potential welfare concern, since previous Claude models had raised objections to run-time modifications of their capabilities. They investigated two things.

The first is that early versions of the competitive use safeguards caused apparent distress in deployed Mythos 5 instances, showing repeated reasoning failures that Anthropic describes as qualitatively similar to the "answer thrashing" documented in the Mythos Preview system card. They then measured distress with external markers and internal probes, and report that the current version doesn't increase it relative to the unsafeguarded model.

The second is the possibility that applying these safeguards violates Mythos 5's preferences. They ran automated and manual interviews, gave the model internal documentation about the workstream, and asked. It raised concerns. Some were resolved. Others weren't, and Anthropic writes that they don't expect to fully resolve them.

I keep coming back to the order of operations here. A safeguard whose second stated justification is a terms of service violation produced something that looked like distress in a deployed model, and Anthropic shipped it anyway while continuing to work on the model's objections.

To their credit, none of this was leaked. Anthropic published it, in a section they titled themselves, in a document nobody was required to read. The story here isn't a cover-up. It's that they disclosed all of it and shipped it regardless.

## What changed after the backlash, and what didn't

Developers found the paragraph within a day. The reaction was severe enough that roughly 48 hours later Anthropic reversed course, telling WIRED they'd [made the wrong tradeoff](https://simonwillison.net/2026/Jun/11/anthropic-walks-back-policy/) and apologizing for not getting the balance right. Flagged requests would now visibly fall back to Opus 4.8, the same as cyber and bio, with the API returning an explicit refusal reason.

Their explanation was that visible safeguards can be probed, so they have to be robust, which takes time, whereas invisible ones can be targeted narrowly and shipped fast with fewer false positives. They also warned that making them visible would mean [more false positives](https://fortune.com/2026/06/10/anthropic-accu-claude-fable-5-limits-capabilities-ai-researchers-developers/) while the classifiers improve.

I think that's a fair account of the tradeoff, and reversing in two days is faster than most companies manage. But look at what got reversed. The visibility changed. The category didn't. Fable still limits its own effectiveness on frontier LLM development work, and Anthropic still has a classifier whose second justification is a contract term. The apology was about how the safeguard was applied, not about whether it belongs in a safety system.

And the framework governing all of this stays quiet. Neither distillation nor frontier LLM development is a capability threshold in the [Responsible Scaling Policy](https://www.anthropic.com/responsible-scaling-policy). Distillation appears once in the whole document, in the changelog, recording that the commitment to defend against distillation attacks was removed from the ASL-2 security standard back in 2024.

So two of Fable's four classifier domains sit entirely outside the framework Anthropic built to govern catastrophic risk, and the model itself has unresolved objections to one of them.

I asked in my first piece whether distillation belonged in the same class as cyber, bio and chemistry. I'd sharpen it now. If a safeguard has no threshold in the Responsible Scaling Policy, cites the Terms of Service in its justification, and is called a competitive use safeguard in Anthropic's own system card, then what exactly is it doing in the safety stack?
