Snippet · Building with models
Cut the cost of a model call
Same output, meaningfully less money.
This call costs more than I'd like: <paste the prompt and roughly how often it runs> Work through, in order of impact: - What's in the prompt that isn't earning its place? Show me the cut. - Is the stable part of this prompt cacheable by my provider? If so, how do I structure it to hit the cache? - Would a smaller model handle this? Say which part of the task actually needs the larger one. - Can I avoid the call entirely for some inputs — a cache, a rule, an early exit? - Is the output longer than it needs to be? Constrain it. Estimate the saving for each, and tell me which one to do first.
Check whether you need the call at all before optimising it. A cache on identical inputs frequently beats every prompt-shortening trick combined.