If your product has a prompt in it, that prompt is code — and right now you're probably changing it, eyeballing one output, and shipping. A playground turns that into something you can actually reason about.

1. The core idea

Two things: a set of test inputs, and versions of a prompt. Run every version against every input, put the results side by side, and read them.

That's it. Everything else is convenience.

playground-build.txt
Build a prompt playground. Stack: <stack>.

- A prompt editor with named, saved versions.
- A set of test inputs, saved and reusable.
- "Run all" — every version against every input, in parallel, with
  progress shown.
- Results in a grid: versions as columns, inputs as rows.
- Each cell shows the output, token counts, latency and cost.
- Diff view between two versions' outputs for the same input.

API key server-side. A hard spend cap per run, with a confirmation
showing the estimated cost before it runs.

2. Collect real inputs

The single most valuable part. Every time your production prompt does something wrong, add that input to the test set. Over a few weeks you accumulate a collection shaped exactly like your product's actual weaknesses — worth more than any generic benchmark.

3. Watch for the regression

The reason this exists: improving a prompt for one case usually breaks another. Without a grid you don't see it. With one, it's immediately visible, and you find out before your users do.

4. Score what you can

For anything with a checkable answer — a classification, an extraction, valid JSON — assert it automatically and show pass/fail. For subjective output, a side-by-side pick is fine and still far better than reading one output at a time.

Track temperature and model version alongside results. "It got worse" is often a model change, not a prompt change, and you can only tell if you recorded it.

Keep every prompt version, including the bad ones. Half the value is being able to go back and see exactly which edit caused the drop.