Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on the evening of 1 September. The two are technically the same base model; the only difference is the safeguard layer applied on top. Fable 5.1 is generally available and can be called through the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. Mythos 5.1 goes only to vetted US organisations inside Project Glasswing.
The announcement rests on two legs: a jump in benchmark scores, and price. Anthropic cut the cache read rate from $1.00 to $0.25 per million tokens, a 75 percent reduction. The company says this amounts to roughly 25 percent lower cost on typical workloads and up to 45 percent on long-running agentic tasks. Base input and output pricing is unchanged at $10 and $50 per million tokens. For comparison, Opus 5 runs at half that, $5 and $25.
The objection came from the testing partner itself
Artificial Analysis, which took part in Anthropic's pre-release testing, disputes the “45 percent cheaper” framing. By its measurement the cache discount saves about $1.40 per task on agentic workloads. But at the maximum effort setting Fable 5.1 produces roughly 1.7 times as many output tokens as Fable 5. The net result is that per-task cost does not fall — it rises by 20 percent.
The figures: at max effort, Fable 5.1 costs $3.76 per Intelligence Index task. Opus 5 costs $2.34 on the same test and scores only three points lower. At extra-high effort Fable 5.1 reaches a score of 65 for $2.72 per task, closing the gap but still costing more than Opus 5. High pricing had been cited as the main reason Fable 5 saw weak adoption among enterprise customers.
The benchmark table
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 (coding) | 55.8% | 42.0% | 52.3% | 37.3% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
As MarkTechPost notes, Anthropic reports a standard error of 3.5 to 4.5 points per model, so the margin deserves more attention than the ranking. The five-point gap on Terminal-Bench 4.0 between Fable 5.1 at 55.8 percent and Mythos 5.1 at 60.9 percent is the direct cost of the safeguard layer: the same model, a different level of safety intervention.
Three breaking changes waiting for developers
- Forced tool use is gone. Setting
tool_choicetoanyortoolreturns a 400 error. - Thinking blocks are model-bound. Fable 5.1 can read earlier models' thinking blocks, but not the other way round; router and fallback setups lose reasoning when they switch down.
- Editing earlier turns invalidates thinking blocks. The check is enforced for accounts created on or after 31 August 2026.
Anthropic also documents regressions. Parallel tool calling has become more variable, so agent loops may issue one call per turn where Fable 5 batched several. The model narrates less, answers from memory more often at low effort, and prefers whole-file rewrites over targeted edits.
Safeguards and content provenance
Cyber safeguards now permit vulnerability discovery but not exploit development. Anthropic says this cuts unnecessary interventions in Claude Code sessions by roughly 60 percent per session. On the biology side, false refusals on benign requests fell by 85 percent. Penetration testing, exploit generation and binary-based vulnerability scanning still redirect to Opus models. The two releases are also the first Claude models to carry a statistical text watermark in their output, with C2PA credentials attached to generated files.