On Claude Fable 5.1, the effort setting is the main thing you control. It decides how long the model thinks before answering, how thoroughly it gathers context, how many tool calls it makes, and how much you pay. Get comfortable changing it.

The five levels

low — quick answers, fewer tool calls, terse confirmations. Surprisingly capable: at low effort Fable 5.1 often matches or beats older models running at their maximum. Right for chat, small edits, questions, anything where you're iterating fast.

medium — the cost-saving step down from high. Roughly matches the previous Fable model at less expense. Good default for routine building once you trust it.

high — the default, and the recommended starting point. Balances depth against time and tokens for most real work.

xhigh — for capability-sensitive work: hard debugging, architecture, anything where being right matters more than being fast. Turns get noticeably longer.

max — correctness over everything. Reserve it for the genuinely hard problems; on routine tasks it over-deliberates and you wait for thinking the task didn't need.

What changes as you go up

More thinking before writing. More context gathered before acting. More verification of its own work. Fewer, more consolidated tool calls at low; more thorough exploration at high and above.

The flip side of high effort on a routine task: it may gather context and deliberate well beyond what the task needs. If something finishes correctly but took longer than it should have, turn effort down — that's the signal.

A trap at xhigh and max

When you ask for a long deliverable — a full file rewrite, a big document, a large table — at the top effort levels, the model may draft the whole thing in its thinking and then write it out again as the answer. Twice the output tokens, twice the wait, no better result.

Two fixes. Run those requests at high, which is the recommended start anyway. Or, if you need the top level, add this to the end of your request:

long-deliverable-note.txt
Everything you produce in one reply, including reasoning before the
reply, counts toward a single output limit. Composing the entire
deliverable in full as reasoning and then again as the reply would
double the length without improving the result — so don't. Use the
reasoning space to settle the structure and the difficult decisions;
use the output space to write the thing once.

Where to set it

In Claude Code and similar tools, effort is a setting — check your tool's options. On the API it's output_config.effort. In the chat interface you may not get direct control, in which case the model picks adaptively.

How to find your level

Don't guess. Take three or four real tasks from your project and run them at low, medium and high. Look at whether the output was actually different, and how long each took. Most people find their routine work is fine at medium and their hard work needs high — and that they were paying for xhigh out of habit.

Level names don't mean the same thing across models. If you tuned effort on an earlier model, re-tune it. High on Fable 5.1 is not high on Opus.