Prompts in production are code. They break, they regress, and they change behaviour when the model updates underneath them. Yet most are changed by editing, running one example, and shipping.

Why "it seems better" fails

Model output varies between runs. Improving one case usually degrades another. And you're comparing new output against a memory of old output, which is not a comparison.

Without a test set, every prompt change is an untested deployment.

Build the set

Twenty inputs is enough to start. Include: the ordinary case, several awkward ones, the empty or minimal input, an unusually long one, and every input you've seen fail in production.

That last category is the valuable one. Add each production failure the day it happens, and within a few weeks you have a suite shaped exactly like your product's real weaknesses.

Decide what "correct" means

For structured output — a classification, an extraction, valid JSON — you can assert automatically. Does it parse? Is the category one of the allowed values? Are the required fields present?

For open text, you can still assert quite a lot: length limits, forbidden phrases, required elements, correct language. Beyond that, a side-by-side human comparison is fine, and still far better than reading one output at a time.

build-evals.txt
My prompt does <what>: <paste>

Build me an eval harness:
- 20 test inputs covering normal, awkward, empty, very long, and
  adversarial cases.
- For each: what a correct output looks like, as an automatic
  assertion where possible.
- A script that runs all 20 against a prompt version and reports
  pass/fail plus every failure in full.
- Support for comparing two prompt versions side by side.

Then tell me which of my 20 are testing the same thing, so I can cut
them.

Record the conditions

Store the model version, the temperature and the date with every result. When quality drops, "the model changed" and "my prompt changed" are different problems, and you can only tell them apart if you wrote it down.

Run it before every change

The whole point. Change the prompt, run the set, compare. It takes a minute and it catches the regression that "it seems better" hides.

Keep every prompt version, including the bad ones. Half the value is being able to find exactly which edit caused the drop.