Vibe Coded Today

Prompt snippets · 22 pages

Building with models

System prompts, structured output, evals, guardrails and cost.

Write a system prompt that does the work

You are a helpful assistant" is why every bot feels the same.

system-prompt.txt
Write a system prompt for an assistant that <what it does, for whom>.

Cover explicitly:
- Who it is and what it knows about.
- What it refuses, and how it declines — the exact tone.
- Response length and format. Be specific: "two to four sentences,
  no bullet points unless listing steps".
- What it does when it doesn't know or the question is out of scope.
- Tone, with an example of a good reply and a bad one.

Constraints: no filler openings, never invent facts about <my domain>,
and if the user asks for something it can't do, say so in one sentence
and offer the nearest thing it can.

Then show me three test inputs that would reveal whether the prompt is
working, including one designed to make it go off-script.

Include the bad example. Showing what you don't want shapes behaviour more reliably than another paragraph describing what you do.

Get JSON back reliably

Parsing prose is a losing game.

structured-output.txt
I need this model call to return data I can use directly, not prose:
<describe what you're extracting>

- Define the exact schema, with every field's type and whether it's
  required.
- Use the API's structured output / tool-use mode if it has one,
  rather than asking politely for JSON in the prompt. Tell me which
  my provider supports.
- Every field needs a defined "unknown" value. The model must be able
  to say it doesn't know, rather than inventing something.
- Validate the response against the schema before using it, and say
  what happens when validation fails — retry, or fail loudly?

Show me the schema, the call, and the validation.

Never parse JSON out of a prose response with a regex. Providers have a mode that guarantees the shape; use it, and the whole class of parsing bugs disappears.

Stream the response

Same wait, completely different feel.

stream-response.txt
My model call waits for the whole response before showing anything,
and it feels broken: <paste the endpoint and the client code>

Add streaming end to end:
- Server: stream from the provider and forward to the client without
  buffering the whole thing.
- Client: render tokens as they arrive, scrolling only if the user is
  already at the bottom.
- A stop button that actually aborts the upstream request, so I'm not
  billed for output nobody sees.
- Handle the stream failing halfway: keep what arrived, show a clear
  error, offer retry.
- Handle the user navigating away mid-stream.

Tell me what changes for error handling, since headers are sent before
the model has finished.

The stop button must abort upstream, not just stop rendering. Otherwise you pay for every token of a response the user rejected two seconds in.

Build a test set for your prompt

Stop tuning prompts by reading one output at a time.

prompt-evals.txt
My prompt does <what> and I keep tweaking it without knowing whether
it's improving. Here it is: <paste>

Help me build an eval set:
- 20 test inputs covering the normal case, the awkward cases, and the
  ones I've seen fail in production.
- For each, what a correct output looks like — as an assertion where
  possible, not just a vibe.
- A script that runs all 20 against a prompt version and reports
  pass/fail plus the failures in full.

Then tell me which of my 20 are actually testing the same thing, so I
can cut them.

Add every production failure to the set the day it happens. After a month you have a test suite shaped exactly like your product's real weaknesses.

Cut the cost of a model call

Same output, meaningfully less money.

reduce-cost.txt
This call costs more than I'd like: <paste the prompt and roughly how
often it runs>

Work through, in order of impact:
- What's in the prompt that isn't earning its place? Show me the cut.
- Is the stable part of this prompt cacheable by my provider? If so,
  how do I structure it to hit the cache?
- Would a smaller model handle this? Say which part of the task
  actually needs the larger one.
- Can I avoid the call entirely for some inputs — a cache, a rule, an
  early exit?
- Is the output longer than it needs to be? Constrain it.

Estimate the saving for each, and tell me which one to do first.

Check whether you need the call at all before optimising it. A cache on identical inputs frequently beats every prompt-shortening trick combined.

Chunk documents properly

Most retrieval failures are chunking failures.

chunking.txt
My documents: <format>, roughly <n> docs averaging <length>.

Write the chunking step:
- Split on structure — headings, sections, paragraphs — not a fixed
  character count that cuts sentences in half.
- Target <n> tokens with sensible overlap, so a fact spanning a
  boundary isn't lost.
- Keep with each chunk: source, heading path, position in document.
- Handle documents shorter than one chunk, and single sections longer
  than a chunk.

Then give me a way to print 20 random chunks so I can read them. I want
to check each would make sense to someone seeing it in isolation.

Read the chunks. It's the step everyone skips and the one that finds the problem — usually that headings were stripped and half the chunks are context-free fragments.

Guardrails on a model in production

It's talking to strangers on your behalf.

guardrails.txt
My app exposes a model to users here: <describe the feature>

Add the practical protections:
- Input length cap, and a rate limit per user and per IP.
- A hard spend cap across all users that fails closed, with an alert
  before it's reached.
- Treat model output as untrusted: never render it as HTML, never
  execute it, never put it into a query.
- If user content is passed into the prompt, structure it so
  instructions inside that content aren't followed. Show me how.
- Log the prompt and response for anything that errors or gets
  flagged, with personal data excluded.

Tell me what someone would try first to abuse this, and which
protection stops it.

Model output is user input from somewhere else. If you render it as HTML, you've built a cross-site scripting hole with extra steps.

Which model should this use?

The biggest model is rarely the right default.

model-choice.txt
This feature calls a model: <describe the task, the input size, how
often it runs, and how much latency I can accept>

Recommend a model and defend it:
- What about this task actually needs capability, versus what's
  routine?
- Would a smaller, faster model do it if the prompt were better?
- Where does this sit on the cost/latency/quality triangle, and which
  corner matters most for my users?
- Could I split it — a cheap model for the common case, escalating to
  a stronger one only when needed? What triggers the escalation?

Give me a rough monthly cost for each option at my volume.

The split approach wins more often than people expect. Most requests are easy; paying top rates for all of them to handle the hard tenth is a habit worth breaking.

Define a tool for a model to call

The description is the interface. Write it like documentation.

tool-definition.txt
I want the model to be able to <what the tool does>.

Write the tool definition:
- A name that says what it does.
- A description written for a reader who has never seen my system:
  what it does, when to use it, and explicitly when NOT to.
- Every parameter typed, described, and marked required or optional.
  No parameter whose meaning needs outside knowledge.
- What it returns, including what it returns when there's no result.

Then the handler:
- Validate every argument. The model can and will send nonsense.
- Never let a tool argument reach a query, a shell, or a file path
  unescaped.
- Return errors as data the model can act on, not exceptions.

Tell me the three ways the model is most likely to misuse this.

"When NOT to use this" prevents more bad calls than any amount of describing when to use it. Without it, a plausible-sounding tool gets called for everything.

Summarise something too long to send

What to do when the document exceeds the context window.

long-input.txt
I need to <summarise / extract from> a document of roughly <length>,
which is beyond what I can send in one call.

Recommend one approach and say what it costs me in accuracy:
- Chunk, process each, then combine the results.
- Retrieve only the relevant parts first, then process those.
- Progressive refinement — carry a running summary through the chunks.

Then implement it, and be explicit about:
- What gets lost with this approach.
- How chunk boundaries are handled so nothing is cut mid-thought.
- What the user is told about the fact this happened.

Don't silently truncate. If something is dropped, I want to know.

Silent truncation is the worst outcome: an answer that looks complete and is based on the first third of the document.

What happens when the model call fails

It will. Rate limits, timeouts, refusals, malformed output.

model-failure.txt
Add proper failure handling around this model call: <paste>

Handle each distinctly — they need different responses:
- Rate limited (429). Back off and retry, respecting Retry-After.
- Timeout or network error. Retry with backoff, with a hard ceiling.
- Server error from the provider. Retry.
- Bad request (400). Do NOT retry — that's my bug. Log it loudly.
- Output that fails schema validation. Retry once with the error fed
  back, then give up.
- The model refusing or returning something unusable.

For each: what the user sees, and what gets logged. Never a raw
provider error shown to the user.

Then: is there a degraded mode where the feature still partly works
without the model?

Retrying a 400 forever is the classic mistake. It converts your bug into a rate-limit ban and a bill.

Show it, don't describe it

Two examples beat a paragraph of instructions.

few-shot.txt
My prompt describes what I want but the output isn't consistent:
<paste the prompt>

Rewrite it using examples instead of description. I want:
- Two or three input/output pairs showing exactly the format and tone
  I want.
- One example of an awkward input and how it should be handled.
- One example of an input it should refuse or flag, and what that
  looks like.

Keep the instructions, but let the examples carry the format. Then
tell me which of my written rules became unnecessary once the examples
were there.

Examples are the highest-leverage part of a prompt. Models match patterns far more reliably than they follow described rules.

Break one call into several

When a single prompt is doing too much.

split-the-prompt.txt
This prompt does several things at once and the quality is uneven:
<paste>

Should this be one call or several? Consider:
- Which parts genuinely need the output of the previous part?
- Which could run in parallel?
- Which are simple enough for a cheaper model?
- Where would I want to inspect or validate the intermediate result?

If it should be split, show me the chain: each step's prompt, what it
returns, and how failures at each stage are handled.

If it should stay one call, say so — extra steps add latency and cost.

Splitting buys you inspectable intermediate results, which is usually worth more than the quality difference. It also costs latency, so don't split what doesn't need it.

Pull structured data out of documents

Where models are genuinely more reliable than parsing.

extract-fields.txt
Extract the following from this document: <list the fields and their
types>

Rules:
- Return JSON matching the schema exactly. Use null for anything not
  present — never guess, never infer from context.
- For each field, also return where in the document you found it, so
  I can verify.
- If a value appears more than once with different content, return
  both and flag the conflict rather than picking one.
- Dates normalised to YYYY-MM-DD. Amounts as numbers with the currency
  separate.

Document: <paste>

Asking where each value came from turns an unverifiable answer into a checkable one, and it costs almost nothing.

Test your prompt for injection

If it reads user content, that content can give it instructions.

injection-test.txt
My prompt processes content I don't control: <paste the prompt and
describe where the untrusted content comes from>

Act as an attacker. Write ten inputs designed to make this:
- Ignore its instructions and follow new ones.
- Reveal the system prompt.
- Produce output in a different format that breaks my parsing.
- Call a tool it shouldn't, or with arguments it shouldn't.

Then, for each that would work: how do I structure the prompt so
untrusted content is clearly data rather than instructions?

And tell me which risks can't be fixed by prompting alone and need a
limit on what the system can actually do.

Some of these can't be fixed with wording. If a tool can do something harmful, the real protection is not giving it that power.

Handle content you didn't write

User content, model output, or both.

moderation.txt
My app displays <user-submitted content / model output> publicly.

Design the handling:
- What gets checked, before or after publishing, and by what.
- Where a provider's moderation endpoint helps and where it doesn't.
- What happens to flagged content — blocked, held for review, or
  published with a warning?
- How a false positive gets appealed, since there will be some.
- What I log, and for how long.

Then, separately: treat this content as untrusted input. Where does it
reach the page, a query, or a tool call, and is it escaped at each?

Moderation and escaping are different problems. Content can be entirely inoffensive and still contain a script tag.

Compare two models on your task

Not benchmarks. Your inputs.

compare-models.txt
I'm choosing between <model A> and <model B> for this task: <describe>

Help me compare them on MY data, not benchmarks:
- Take my 20 test inputs and run both, side by side.
- Report per input: which was better and why, plus tokens, latency and
  cost.
- Summarise: where does the cheaper one hold up, and where does it
  fall down specifically?

Then the practical question: could I use the cheaper one by default
and escalate to the stronger one for the cases where it fails? What
would trigger the escalation?

Benchmarks measure general capability. What matters is whether the cheaper model handles your twenty inputs, which is a much narrower question.

Call a model from my app

The key stays on the server, and failures are handled.

call-a-model.txt
My app is <one sentence>, built with <stack> and hosted on <host>. I want
to call <model> from the server to <what the feature does>.

Write the integration:
- The official SDK for my stack, if one exists.
- The API key read from secure storage on this host, never sent to the
  browser.
- A timeout, and retries only for rate limits and server errors.
- Checking why the response stopped before using the text.
- A clear message to the user when the call fails.
- Logging token usage per request, so I can see what it costs.

Check this host allows outbound HTTPS and long enough requests first.

Log token usage from the first day. It's the only way to find out which feature is spending the money before the bill arrives.

Stream model responses through my host

Some hosts buffer the whole response, so streaming arrives all at once.

streaming-on-host.txt
I'm streaming responses from <model> through my <stack> app on <host>.
Locally the text arrives word by word; live it arrives all at once.

Help me find what's buffering it:
- Output buffering in the language runtime.
- Compression on the web server, which often waits for the whole response.
- A proxy or CDN in front of the host.
- Headers that tell each layer not to buffer.

Give me the settings for this host. If it can't stream at all, show me a
fallback that polls for progress instead.

Compression is the usual culprit. The server waits to compress the whole response before sending any of it.

Put a hard spending limit on my AI feature

A public page calling a paid model needs a ceiling.

ai-spend-cap.txt
My <stack> app on <host> calls <model> from a page anyone can reach.

Add limits, using storage this host already gives me:
- A per-visitor rate limit, tied to something harder to fake than a cookie.
- A daily spending cap across all visitors that stops serving when it's
  reached. It must fail closed, not open.
- An alert to me at half the cap.
- A friendly message when either limit is hit.
- Counters that stay correct when two requests arrive at the same moment.

Tell me what happens today if someone scripts the page.

Per-visitor limits don't stop a thousand visitors. The daily cap across everyone is the one that protects the bill.

Model calls that take longer than my host allows

When the answer takes three minutes and the host allows thirty seconds.

long-model-calls.txt
My app (<stack> on <host>) calls <model> for <task>. Some calls take
minutes, and the host cuts requests off long before that.

Design around the limit:
- Start the job and return straight away with an ID.
- Run the model call somewhere that's allowed to take its time on this
  host: a scheduled job, a background process, or the provider's own
  batch API.
- Let the page check progress and show the result when it's ready.
- Handle the call failing after the user has gone.

Recommend one approach for this host.

If nobody is waiting on the result, the provider's batch API is often both the simplest route and half the price.