model-choice.txt
This feature calls a model: <describe the task, the input size, how
often it runs, and how much latency I can accept>

Recommend a model and defend it:
- What about this task actually needs capability, versus what's
  routine?
- Would a smaller, faster model do it if the prompt were better?
- Where does this sit on the cost/latency/quality triangle, and which
  corner matters most for my users?
- Could I split it — a cheap model for the common case, escalating to
  a stronger one only when needed? What triggers the escalation?

Give me a rough monthly cost for each option at my volume.

The split approach wins more often than people expect. Most requests are easy; paying top rates for all of them to handle the hard tenth is a habit worth breaking.