Guide · Building with Fable 5.1
Effort levels, and which one to use
The one dial that controls speed, cost and depth on Fable 5.1.
On Claude Fable 5.1, the effort setting is the main thing you control. It decides how long the model thinks before answering, how thoroughly it gathers context, how many tool calls it makes, and how much you pay. Get comfortable changing it.
The five levels
low — quick answers, fewer tool calls, terse confirmations. Surprisingly capable: at low effort Fable 5.1 often matches or beats older models running at their maximum. Right for chat, small edits, questions, anything where you're iterating fast.
medium — the cost-saving step down from high. Roughly matches the previous Fable model at less expense. Good default for routine building once you trust it.
high — the default, and the recommended starting point. Balances depth against time and tokens for most real work.
xhigh — for capability-sensitive work: hard debugging, architecture, anything where being right matters more than being fast. Turns get noticeably longer.
max — correctness over everything. Reserve it for the genuinely hard problems; on routine tasks it over-deliberates and you wait for thinking the task didn't need.
What changes as you go up
More thinking before writing. More context gathered before acting. More verification of its own work. Fewer, more consolidated tool calls at low; more thorough exploration at high and above.
The flip side of high effort on a routine task: it may gather context and deliberate well beyond what the task needs. If something finishes correctly but took longer than it should have, turn effort down — that's the signal.
A trap at xhigh and max
When you ask for a long deliverable — a full file rewrite, a big document, a large table — at the top effort levels, the model may draft the whole thing in its thinking and then write it out again as the answer. Twice the output tokens, twice the wait, no better result.
Two fixes. Run those requests at high, which is the recommended start anyway. Or, if you need the top level, add this to the end of your request:
Everything you produce in one reply, including reasoning before the reply, counts toward a single output limit. Composing the entire deliverable in full as reasoning and then again as the reply would double the length without improving the result — so don't. Use the reasoning space to settle the structure and the difficult decisions; use the output space to write the thing once.
Where to set it
In Claude Code and similar tools, effort is a setting — check your tool's options. On the API it's output_config.effort. In the chat interface you may not get direct control, in which case the model picks adaptively.
How to find your level
Don't guess. Take three or four real tasks from your project and run them at low, medium and high. Look at whether the output was actually different, and how long each took. Most people find their routine work is fine at medium and their hard work needs high — and that they were paying for xhigh out of habit.
Level names don't mean the same thing across models. If you tuned effort on an earlier model, re-tune it. High on Fable 5.1 is not high on Opus.