Newer Claude models take an effort setting alongside the model choice. It controls how much the model thinks before answering, and because thinking is billed as output tokens, it moves your bill a great deal.

Most people set a model and never touch effort. That's leaving both money and quality on the table.

The five levels

From least to most: low, medium, high, xhigh, max. High is the default when you don't set anything.

There is no separate price per level. A level doesn't have a rate; it changes how many tokens get generated. Higher effort means more thinking tokens, and thinking bills at the output rate, which is five times the input rate on every model.

That's the whole cost model, and it has one useful consequence: effort is a multiplier on the expensive half of your bill.

Roughly what changes

Going up a level means more deliberation before answering, more context gathered, more verification of its own work. Going down means fewer and more consolidated tool calls, less preamble, shorter confirmations.

The differences are not small. The same task at low and at max can differ by several times in total tokens generated. You should measure it on your own work rather than trusting anyone's multiplier, including mine.

effort-measurement.txt
I want to measure what effort costs me on real work.

I'm going to run these three tasks at low, medium and high:
<paste three real tasks from your project>

For each run I'll give you the output, the token counts from my usage
dashboard, and how long it took.

Then tell me: where did the extra effort change the result, and where
did it only add tokens? Recommend a default for my routine work and a
level for hard problems.

Which to use

low for chat, small edits, classification, anything high volume or latency sensitive. On current models low effort is genuinely capable, often matching older models running flat out.

medium as the cost-saving step down from the default, where your own testing shows quality holds.

high is the default and the right starting point for real work.

xhigh for capability-sensitive work: hard debugging, architecture, long agent runs.

max when correctness matters more than cost. Reserve it. On routine work it deliberates past the point of usefulness and you pay for every token of it.

The cheaper move before a cheaper model

If a workload is costing too much, try the same model at lower effort before switching to a smaller model. Lower effort on a strong model often beats high effort on a weaker one, and it keeps you on a single model, which matters because caches are per-model. Switching models mid-project throws your cache away and you pay full input price again.

Watch for one trap

At the top two levels, asking for a long deliverable in one go can make the model draft the whole thing in its thinking and then write it out again as the answer. Twice the output tokens, no better result. The Fable effort guide has the prompt that prevents it, and running those requests at high avoids it entirely.

Effort is per request. You can run a whole project at low and raise it for the three hard turns. Most people set it once and forget, which is how routine work ends up priced like hard work.