Benchmarks won't tell you which model to use for your project. The useful question is what happens when the model is wrong, and how expensive being wrong is.

Start with the cost of a mistake

A mistake you'll notice immediately — a broken layout, a failing test, a page that won't load. Use a cheap model. You'll catch the error in seconds and the retry costs nothing.

A mistake you won't notice — a subtly wrong calculation, a permission check that passes when it shouldn't, a migration that drops a column. Use a capable model and read the diff carefully. This is where the price difference is irrelevant next to the cost of being wrong.

A mistake that compounds — anything an agent will build on for the next two hours. Wrong assumptions early become two hours of wasted work. Pay for the better model at the start of a long run.

By task

Classification, tagging, routing, extraction — Haiku. The output space is small and checkable, and you're likely doing it thousands of times.

Writing and editing code, day to day — Sonnet 5. This is most vibe coding, and Sonnet is genuinely good at it.

Debugging something you're stuck on — Opus 5. If you've already failed three times with a cheaper model, the problem isn't your prompt.

Long autonomous agent runs — Opus 5 or Fable 5.1. One wrong turn in hour one costs more than the model does all day.

Architecture and data modelling — Opus or Fable. These are the decisions that are expensive to reverse.

Bulk sub-agents inside a bigger task — Haiku or Sonnet, even when the lead agent is Opus. Reading and summarising doesn't need the expensive model.

By what you're paying with

If you're on a subscription through a coding tool, model choice affects your usage allowance rather than a bill. If you're on the API, it's money. The tradeoffs are the same shape but the feedback is slower on a subscription, so people over-use the expensive model without noticing. See subscription or API.

A working default

model-choice.txt
I'm building <what> with <stack>. The part I'm working on is <describe>.

If this goes wrong, what happens: <I'd notice immediately / I might not
notice / an agent would build on it for hours>.

Recommend one model and one effort level. Justify it against the option
one step cheaper. If the cheaper one would do, say so plainly — I'd
rather not overpay out of caution.

Change your mind on evidence

The signal to move up: you've failed several times at the same task with a clear prompt and full context. The signal to move down: the answers are correct and the model is clearly doing more work than the task needs.

Neither signal is "it feels like this deserves the good model".

Don't pick a model per project. Pick one per task. The same afternoon can reasonably use three.