compare-models.txt
I'm choosing between <model A> and <model B> for this task: <describe>

Help me compare them on MY data, not benchmarks:
- Take my 20 test inputs and run both, side by side.
- Report per input: which was better and why, plus tokens, latency and
  cost.
- Summarise: where does the cheaper one hold up, and where does it
  fall down specifically?

Then the practical question: could I use the cheaper one by default
and escalate to the stronger one for the cases where it fails? What
would trigger the escalation?

Benchmarks measure general capability. What matters is whether the cheaper model handles your twenty inputs, which is a much narrower question.