Using more than one model in a project
Cascades and sub-agents work well. There's one cost you'll miss.
Once you're comfortable choosing models, the next idea is obvious: use a cheap one for the easy parts and an expensive one for the hard parts. It's a good idea, and there's a specific cost people don't anticipate.
The hidden cost: caches are per model
Prompt caches are scoped to a model. Send the same system prompt to two different models and you have two separate caches, each paid for at write rates.
For a workload with a large stable prefix, this can wipe out the savings from the cheaper model. If you're routing between models on a big shared system prompt, work out the cache cost before assuming you're ahead.
This is the main reason to try lower effort on one model before reaching for a second model.
Where multiple models genuinely pay off
Sub-agents inside a bigger task. A lead agent on a capable model, with reading and summarising delegated to Haiku or Sonnet. The sub-agents have their own short contexts, so there's little cache to lose, and bulk reading is exactly what a cheap model is for.
Escalation on failure. Cheap model first, escalate when confidence is low or validation fails. Most requests never touch the expensive model.
Genuinely different jobs. Classification on Haiku, code generation on Sonnet, architecture on Opus. Different prompts, different caches, no shared prefix to lose.
That last case is the cleanest. The problems appear when you route the same prompt between models.
Deciding the split
My app does <describe>. Right now everything runs on <model>. I'm considering splitting work between a cheap and an expensive model. Tell me: - Which steps genuinely need the capable model, and which don't? - Do the cheap steps share a large prompt prefix with the expensive ones? If so, what does the extra cache cost me? - Would lower effort on a single model get most of the saving with none of the complexity? - What triggers escalation, and what happens if that trigger is wrong? If one model at mixed effort is the better answer, say so.
The operational cost
Two models means two sets of behaviour to understand, two prompts to maintain, two things that can change under you. That's real work, and on a small project it can exceed what you save.
For a side project, one model at a well-chosen effort level is usually right. Cascades earn their complexity at volume.
Before building a cascade, measure the capable model at low effort on the same tasks. It's frequently close enough, and it's one model to maintain instead of two.