When the cheap model is the right model
Haiku and Sonnet do more than people give them credit for.
There's a habit of reaching for the most capable model available for everything, on the reasoning that better is better. It's understandable and it's usually wrong. Plenty of what you do all day has a small, checkable answer, and paying ten times the rate for it buys nothing.
The shape of a task a cheap model handles well
Small output space. Classifying into one of six categories. Deciding whether something matches. Picking a label. There isn't much room to be creatively wrong.
Checkable results. You can verify it mechanically, so a mistake surfaces immediately rather than shipping.
Well-specified. Little judgement needed. The rules are in the prompt.
High volume. Where the rate actually matters, because you're doing it thousands of times.
Extraction, tagging, routing, formatting, summarising to a fixed structure, first-pass filtering. All of this is Haiku work.
The shape that needs a capable model
Ambiguity. Long chains of reasoning where an early mistake compounds. Anything where being confidently wrong is expensive and hard to detect. Architecture. Debugging something that has already defeated a cheaper attempt.
The middle, which is most things
Most vibe coding is Sonnet work. It writes good code, it's a fifth of Fable's rate, and for ordinary building the difference rarely shows. Starting at Sonnet and escalating on evidence is a better policy than starting at the top and never reconsidering.
Try lower effort before a lower model
If cost is the concern, the first move is the same model at lower effort, not a different model. Lower effort on a strong model frequently beats high effort on a weak one, and staying on one model keeps your cache intact. Switching models throws it away.
Testing the downgrade honestly
I'm using <expensive model> for <this task>. I want to know whether <cheaper model> would do. Here are ten real examples with the answers I consider correct: <paste> Run both models over all ten. Report: how many each got right, where the cheaper one failed, and whether the failures share a pattern. Then tell me honestly whether to switch, switch with a fallback for the hard cases, or stay put. If staying put is right, say so.
Ten real examples with known answers is the whole method. It takes an hour and it replaces an argument you'd otherwise have with yourself every month.
The escalation pattern
The best of both: run the cheap model by default, detect low confidence or a failed validation, and retry those on the capable one. You get most of the savings and most of the quality, and the share escalating tells you when something has drifted.
Using the best model for everything isn't caution, it's an unexamined default. Caution is measuring which tasks actually need it.