prompt-evals.txt
My prompt does <what> and I keep tweaking it without knowing whether
it's improving. Here it is: <paste>

Help me build an eval set:
- 20 test inputs covering the normal case, the awkward cases, and the
  ones I've seen fail in production.
- For each, what a correct output looks like — as an assertion where
  possible, not just a vibe.
- A script that runs all 20 against a prompt version and reports
  pass/fail plus the failures in full.

Then tell me which of my 20 are actually testing the same thing, so I
can cut them.

Add every production failure to the set the day it happens. After a month you have a test suite shaped exactly like your product's real weaknesses.