Anthropic Made Its Priciest Model Cheap Where Agents Repeat Themselves
By Toolbox Ninja · · 6 min read
Fable 5.1 keeps Anthropic's premium base price but cuts cached input by 75%, changing the economics of agents that repeatedly revisit the same context.
Anthropic Made Its Priciest Model Cheap Where Agents Repeat Themselves
Anthropic's newest model still has a premium sticker price. Claude Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, the same base rates as Fable 5. But one small line in the price sheet changes how the bill works for long-running agents: cached input now costs $0.25 per million tokens, down from $1.[1][4]
That discount is more interesting than another crowded benchmark chart. Agents repeatedly reread the same material: system instructions, tool definitions, repository files, policy documents and their own growing history. Fable 5.1 makes those repeated reads 75% cheaper. Anthropic estimates that typical workloads will cost about 25% less, while highly agentic work may save up to roughly 45%.[1]
The pitch is clear. Anthropic wants its most expensive general-use model to make economic sense when the job lasts hours rather than one prompt.
The cheap part is repetition
Prompt caching is easy to overlook because the normal input and output prices have not moved. A fresh million-token input still costs $10. A cached million-token read costs one fortieth of that amount. VentureBeat notes an odd result: Fable 5.1's cached input is half the price of an Opus 5 cache read, even though its ordinary input and output are twice as expensive.[4]
This does not make Fable 5.1 a cheap model across the board. A workflow that constantly feeds it new context, produces long answers or misses the cache will still rack up a large bill. The savings favor a particular shape of work: a stable base of instructions and documents revisited many times while an agent takes steps, calls tools and checks its progress.
Think of a coding agent working through a large repository. It may need the same architecture notes, API contracts and tool descriptions for dozens of turns. Paying the full input rate every time makes persistent context expensive. Caching shifts more of the cost toward new material and output. That is a practical change, not a magical efficiency gain.
The write side matters too. VentureBeat reports that five-minute cache writes cost $12.50 per million tokens and one-hour writes cost $20; the $0.25 rate applies to later reads.[4] Teams still need enough reuse to earn back that setup cost. A short chat may see little benefit. An agent that keeps returning to the same code and rules is where the arithmetic gets friendlier.
Fable and Mythos are two doors to one model
Anthropic released Fable 5.1 and Mythos 5.1 on September 1. They use the same underlying model but have different safeguards. Fable is generally available. Mythos, whose controls permit more work in cybersecurity and life sciences, is limited to vetted participants in Anthropic's trusted-access programs.[1]
The public Fable version can now search source code for vulnerabilities, though Anthropic says it will not help develop exploits. The company also says its latest cyber safeguards produce 60% fewer false positives than the previous version.[1] TechCrunch describes the broader tradeoff more plainly: Mythos is slightly more willing than Opus 5 to cooperate with misuse and accept unverifiable authorization, while improving on some failures found in earlier models.[3]
Those details matter because the lower cache price encourages longer sessions with more tools and more accumulated context. A capable model left running against a repository or browser presents a different risk than a chatbot answering one question. Permissions, logs, checkpoints and human review become part of the product, whether vendors put them on a benchmark chart or not.
Better scores, with the usual caveat
Anthropic reports sizable gains on agent tasks. Fable 5.1 scored 55.8% on Terminal-Bench 4.0, compared with 42.0% for Fable 5, and 31.4% on AutomationBench versus 17.1% for its predecessor.[1] An independent ARC Prize results page reports 97.5% on the ARC-AGI-1 semi-private set and 90.0% on ARC-AGI-2 at maximum effort, with estimated per-task costs of $1.40 and $4.49 respectively.[5]
These numbers need careful reading. Anthropic says safeguards intervened in some evaluations, sometimes producing a zero or routing a task to another model. It also warns that its August OSWorld task release cannot be directly compared with previously published results.[1] Early customer stories are useful clues, but they are still testimonials supplied for a launch, not independent reproductions.[4]
For buyers, the most useful metric may be cost per completed task rather than cost per token or a single success rate. A seemingly cheaper model can lose its advantage if it retries often, burns more tokens or needs frequent rescue. A premium model can be economical if it finishes reliably and reuses most of its context. Fable 5.1's new pricing is built around that bet.
Privacy moves into the customer's cloud
Long-running agents create another problem: they leave a rich trail of prompts, files and actions. Anthropic's default policy for Fable requires 30-day retention for safety monitoring. Its new Enterprise Frontier Safeguards system, due to roll out in phases beginning later this fall, is designed to keep monitoring data in cloud infrastructure controlled by the customer instead.[2]
Under that design, customers retain data under their own encryption keys, access policies and audit logs. Automated systems inspect a rolling window for signs such as serious misuse or stolen credentials, then send flags to the customer. Anthropic says its employees do not need to perform the human review.[2]
This is not the same as turning monitoring off. It changes who holds the logs and who investigates the alerts. Anthropic says it developed the setup with more than 100 customers and cloud partners including AWS, Google Cloud and Microsoft Azure.[2] Eligible customers can use zero data retention until the system is available.[1]
That architecture may be the quieter half of this release. Agents are useful because they can see enough context to act. The same context can contain source code, contracts, internal messages and credentials. Cheap repetition helps the agent keep working; customer-controlled storage gives enterprises a less awkward place to keep the evidence of what it did.
What teams should test
A fair trial should resemble the real workload. Measure how much of the prompt actually hits the cache, how often the model retries, how long it runs and how much output it produces. Then inspect the work, not just the invoice. The model's list price tells very little about a task that spans hundreds of tool calls.
Teams should also decide where the agent can act before chasing benchmark gains. Read-only repository access is a different proposition from production credentials. A lower cache bill is welcome, but it does not make broad permissions sensible.
Fable 5.1 is still expensive when every token is new. Anthropic has simply made repeated context unusually cheap. For the next generation of coding and research agents, repetition may be where much of the money goes. That makes a 75% cache discount one of the few launch numbers worth watching.
Sources
[1] https://www.anthropic.com/claude-fable-and-mythos-5-1 — Introducing Claude Fable 5.1 and Claude Mythos 5.1 [2] https://www.anthropic.com/news/enterprise-frontier-safeguards — Developing Enterprise Frontier Safeguards with our customers [3] https://techcrunch.com/2026/09/01/anthropics-new-fable-release-is-cheaper-less-restrictive — Anthropic's new Fable release is cheaper, less restrictive [4] https://venturebeat.com/ai/anthropics-claude-fable-5-1-and-mythos-5-1-arrive-with-a-75-cost-reduction-for-fable-cache-reads — Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads [5] https://arcprize.org/results/anthropic-claude-fable-5-1 — Claude Fable 5.1 - ARC-AGI Results