Claude Fable 5.1 is priced above the Opus tier. If you're using it through a subscription, that's someone else's problem. If you're building on the API — or your coding tool bills by usage — it's yours, and it's worth understanding.

The numbers, per million tokens

  • Claude Fable 5.1 — $10 in, $50 out
  • Claude Opus 5 — $5 in, $25 out
  • Claude Sonnet 5 — $2 in, $10 out

Cache reads on Fable 5.1 are much cheaper than fresh input — about $0.25 per million — which matters a great deal for anything with a large stable system prompt or a long conversation.

Prices are the Anthropic first-party rates as published; check the current pricing page before budgeting, since they change.

Why it isn't simply "twice Opus"

Cost per token is not cost per task. Fable 5.1 at low effort often completes tasks that Opus needs high effort and several retries for. Fewer round trips, fewer wasted turns, fewer "that didn't work, try again" cycles. Measure what a finished task costs, not what a request costs.

Conversely, at high effort on a routine task it can deliberate well beyond what was needed, and you pay for that thinking. Effort is your cost control — see effort levels.

When Fable 5.1 earns its price

Hard debugging where cheaper models have already failed. Long autonomous runs — hours of agent work where one wrong turn early costs more than the model does. First-shot implementations of well-specified systems. Anything where being right the first time saves you a day.

When it doesn't

A CSS tweak. Renaming things. A quick question. Routine CRUD. Chat. For all of these, Fable 5.1 at low effort is very good — but Sonnet 5 is also very good and a fifth of the price.

A sensible pattern for a project: Sonnet or Opus for the day-to-day, Fable for the hard problems. Or Fable throughout at low effort for routine work and high for the hard bits, using per-message effort so one conversation mixes levels without resetting the cache.

Before reaching for a cheaper model

Try Fable 5.1 at lower effort first. Anthropic's guidance is explicit: lower effort on Fable often exceeds top effort on previous-generation models, and one model means one cache — switching models between requests throws the cache away.

Controls that matter on the API

fable-cost-controls.txt
If you're building a product on Fable 5.1:

- Cache the stable prefix (system prompt, tool definitions). Verify
  with cache_read_input_tokens — if it's zero, something is
  invalidating it.
- Set max_tokens generously at high effort and above; it's a hard cap
  on thinking plus output combined.
- Use per-message effort to run routine turns at low and hard turns at
  high inside one conversation.
- A hard daily spend cap that stops serving. Fable's per-token price
  makes an unmetered public endpoint expensive fast.
- Log usage per request. You want to know which feature spends the
  money.

Judge by cost per completed task. A model that costs twice as much per token and finishes in a third of the turns is the cheaper model.