Model APIs charge per token. A bug, a loop, or a stranger who finds your endpoint can produce a bill that bears no relation to your usage. This has ruined people's months, and the protections take an afternoon.

Know your numbers first

Work out what one typical request costs. Input tokens plus output tokens, at your model's rates. Then multiply by your expected volume — and by a volume you'd consider a disaster.

If "a thousand requests an hour" is a number you couldn't afford, you need a limit before you launch, not after.

The four protections

A hard spend cap that fails closed. Track spending, and when the daily limit is reached, stop serving. Not a warning — a stop. This is the one that prevents catastrophe.

Rate limits per user and per IP. Prevents both loops and abuse.

Input length caps. Someone pasting a novel into your summariser costs you real money. Cap it before it reaches the API.

Output length caps. max_tokens on every call. An unbounded response can be very long.

cost-controls.txt
My app calls <model> here: <describe the endpoint and expected volume>

Add cost controls:
- Track spend per request and accumulate per day.
- A hard daily cap that stops serving when reached — fails closed, not
  open. Tell me what users see when it triggers.
- An alert at 50% and 80% of the cap.
- Rate limit per user and per IP. Say what limits you chose and why.
- Input length cap before the call, and max_tokens on it.

Then estimate my monthly cost at 10x and 100x my expected volume, and
tell me which line grows fastest.

Reduce the cost per call

Once you're protected, optimise. In rough order of impact:

Cache identical requests. Often the biggest single win, and frequently overlooked.

Use provider prompt caching for the stable part of your prompt.

Use a smaller model where the task doesn't need the larger one. Split: cheap model for the common case, escalate only when needed.

Shorten the prompt. Every token in your system prompt is paid for on every single call.

Constrain the output. Shorter answers cost less and are often better anyway.

Watch it daily at first

Check the provider's usage dashboard every day for the first fortnight after launching anything public. You'll find out quickly whether anyone has found your endpoint and what they're doing with it.

Never put an uncapped model call behind an unauthenticated public form. It's the single most expensive mistake available in this category.