Snippet · Building with models
Which model should this use?
The biggest model is rarely the right default.
This feature calls a model: <describe the task, the input size, how often it runs, and how much latency I can accept> Recommend a model and defend it: - What about this task actually needs capability, versus what's routine? - Would a smaller, faster model do it if the prompt were better? - Where does this sit on the cost/latency/quality triangle, and which corner matters most for my users? - Could I split it — a cheap model for the common case, escalating to a stronger one only when needed? What triggers the escalation? Give me a rough monthly cost for each option at my volume.
The split approach wins more often than people expect. Most requests are easy; paying top rates for all of them to handle the hard tenth is a habit worth breaking.