Testing prompts like code
If a prompt is in your product, it needs tests.
Prompts in production are code. They break, they regress, and they change behaviour when the model updates underneath them. Yet most are changed by editing, running one example, and shipping.
Why "it seems better" fails
Model output varies between runs. Improving one case usually degrades another. And you're comparing new output against a memory of old output, which is not a comparison.
Without a test set, every prompt change is an untested deployment.
Build the set
Twenty inputs is enough to start. Include: the ordinary case, several awkward ones, the empty or minimal input, an unusually long one, and every input you've seen fail in production.
That last category is the valuable one. Add each production failure the day it happens, and within a few weeks you have a suite shaped exactly like your product's real weaknesses.
Decide what "correct" means
For structured output — a classification, an extraction, valid JSON — you can assert automatically. Does it parse? Is the category one of the allowed values? Are the required fields present?
For open text, you can still assert quite a lot: length limits, forbidden phrases, required elements, correct language. Beyond that, a side-by-side human comparison is fine, and still far better than reading one output at a time.
My prompt does <what>: <paste> Build me an eval harness: - 20 test inputs covering normal, awkward, empty, very long, and adversarial cases. - For each: what a correct output looks like, as an automatic assertion where possible. - A script that runs all 20 against a prompt version and reports pass/fail plus every failure in full. - Support for comparing two prompt versions side by side. Then tell me which of my 20 are testing the same thing, so I can cut them.
Record the conditions
Store the model version, the temperature and the date with every result. When quality drops, "the model changed" and "my prompt changed" are different problems, and you can only tell them apart if you wrote it down.
Run it before every change
The whole point. Change the prompt, run the set, compare. It takes a minute and it catches the regression that "it seems better" hides.
Keep every prompt version, including the bad ones. Half the value is being able to find exactly which edit caused the drop.