Snippet · Building with models
Build a test set for your prompt
Stop tuning prompts by reading one output at a time.
My prompt does <what> and I keep tweaking it without knowing whether it's improving. Here it is: <paste> Help me build an eval set: - 20 test inputs covering the normal case, the awkward cases, and the ones I've seen fail in production. - For each, what a correct output looks like — as an assertion where possible, not just a vibe. - A script that runs all 20 against a prompt version and reports pass/fail plus the failures in full. Then tell me which of my 20 are actually testing the same thing, so I can cut them.
Add every production failure to the set the day it happens. After a month you have a test suite shaped exactly like your product's real weaknesses.