Anthropic ships several Claude models at once, and the names don't tell you much. Here's what each is for and what it costs, so you can stop defaulting to whichever one your tool picked.

Prices below are the first-party API rates, checked in September 2026, in US dollars per million tokens. Check the pricing page before you budget anything, because these do change.

The families

Haiku is the small fast one. Claude Haiku 4.5 costs $1 in, $5 out, with a 200,000-token context window. It's for high volume and simple judgement: classification, routing, tagging, extraction, cheap sub-agents. It is not the model for architecture decisions.

Sonnet is the workhorse. Claude Sonnet 5 costs $2 in, $10 out, with a 1M context window. For most production workloads and most vibe coding, this is the sensible default. It is now the cheapest current-generation model per token after Haiku, because the introductory price became permanent rather than rising as originally scheduled.

Opus is the capable one. Claude Opus 5 costs $5 in, $25 out, 1M context. Reach for it on hard debugging, long agent runs, and anything where a wrong answer costs more than the model does. Opus 4.8, 4.7, 4.6 and 4.5 are all still served at the same price.

Fable is the most capable widely available model. Claude Fable 5.1 costs $10 in, $50 out, 1M context. Double Opus per token. It earns that on the work nothing else finishes. See what Fable 5.1 changes.

Mythos is the same capability as Fable, available only through a limited-access programme. If you haven't been invited, it isn't a choice you need to make.

The whole price list

Input and output, per million tokens:

  • Claude Fable 5.1 and Fable 5 — $10 and $50
  • Claude Opus 5, 4.8, 4.7, 4.6 and 4.5 — $5 and $25
  • Claude Sonnet 5 — $2 and $10
  • Claude Sonnet 4.6 and 4.5 — $3 and $15
  • Claude Haiku 4.5 — $1 and $5

Note that Sonnet 5 is cheaper than the Sonnet 4.6 it replaced. Newer is not automatically pricier.

Output costs five times input

On every model in that list, output is 5x the input rate. This is the fact most people miss when estimating, and it changes what you optimise. A long system prompt you send repeatedly is cheap, especially cached. A model that rambles is expensive.

Asking for shorter answers is a cost control, not just a style preference.

What this means in practice

A rough sense of scale: a million tokens is roughly 750,000 words. A typical back-and-forth coding session might move a few hundred thousand tokens including all the file contents you paste. On Sonnet 5 that's cents. On Fable 5.1 it's still not much, until you run fifty of them a day or leave an agent running for six hours.

The cost of a personal project is almost never the problem. The cost of a public endpoint anyone can call is where people get hurt.

Picking one

Start at Sonnet 5 and move only when you have a reason. Move down to Haiku for bulk work with simple judgement. Move up to Opus when Sonnet keeps getting it wrong. Move up to Fable when Opus keeps getting it wrong, or when you want a long autonomous run to actually finish.

More on that in which model for which job.

The expensive mistake isn't picking a pricey model. It's picking a cheap one, needing four attempts, and paying more in tokens and your own time than the good model would have cost once.