Half price for work that can wait
The Batch API is a 50% discount for anything not needed right now.
If a job doesn't need an answer immediately, you can submit it as a batch and pay half. Both input and output are discounted, on every model.
Most people building side projects never touch this, and a surprising amount of what they run would qualify.
What it costs
Exactly half the standard rate. Some examples, per million tokens:
- Claude Fable 5.1 — $5 in, $25 out
- Claude Opus 5 — $2.50 in, $12.50 out
- Claude Sonnet 5 — $1 in, $5 out
- Claude Haiku 4.5 — $0.50 in, $2.50 out
The discount stacks with prompt caching, so a batched job with a cached prefix is cheap on both axes.
What fits
Anything where nobody is watching a spinner. Processing yesterday's uploads. Tagging a backlog. Generating summaries overnight. Re-running a classification over your whole archive after you improved the prompt. Building an embeddings index. Enriching records on a schedule.
The test is simple: if the answer arriving in an hour instead of a second changes nothing, batch it.
What doesn't fit
Anything a person is waiting for. Chat. Anything interactive. Fast mode isn't available with batching either, which makes sense, since they optimise for opposite things.
How it works in outline
You submit a list of requests, each with an identifier you choose. You poll until the batch reports it has ended, then stream the results. Each result carries the identifier you gave it and either a message or an error.
I have <n> items to process with the model: <describe the work>. Nobody is waiting on the result. Set this up as a batch job: - Build the request list with a stable custom identifier per item. - Submit, then poll for completion with sensible backoff. - Read results by identifier, never by position — they come back in any order. - Handle per-item failures without losing the successful ones. - Store results as they arrive so a crash doesn't cost me the run. Then tell me what this costs batched versus one at a time.
That "never by position" rule is the one bug people hit. Results are not returned in submission order.
The mental shift
Once batching is available, the question changes from "can I afford to run this over everything?" to "which of my one-at-a-time jobs should have been a batch?" Re-running a better prompt across your entire archive stops being a big decision.
Anything on a schedule is a batch job you haven't converted yet.