If you're putting a model behind a feature other people use, you need a number before you ship. Not a precise one. An order of magnitude is enough to tell you whether the idea works.

The four quantities

For one typical use of the feature:

  1. Input tokens, including the system prompt, any documents, and the conversation so far
  2. Output tokens, including thinking
  3. How many model calls one use takes
  4. How many times a day it happens

Multiply out, apply the model's rates, and you have a monthly figure.

Do it at three volumes

Ten users, a thousand, a hundred thousand. The shape of the curve tells you more than any single number, and it shows you which line grows fastest.

cost-estimate.txt
I'm building: <describe the feature and what the model does in it>
Model: <model>, effort <level>.

Estimate the monthly cost at 10, 1,000 and 100,000 users. Show the
arithmetic so I can change the assumptions.

Specifically:
- Tokens in and out for one typical use, and how you got that figure.
- How many model calls one use takes.
- Which line grows fastest, and what makes it grow.
- The cheapest change that halves the biggest line.
- Where a bug or a bored stranger could run this up, and what limit
  stops that.

The parts people forget

Output is five times the input rate. A chatty feature costs far more than a terse one. Constraining response length is a cost control.

Thinking bills as output. Effort level moves this line directly.

The whole conversation is resent every turn. A ten-turn chat doesn't cost ten times one turn; it costs considerably more, because turn ten carries turns one through nine as input. Caching fixes most of this.

Tool definitions cost tokens on every request. See tools cost tokens too.

Retries. Failures still bill.

Server-side tools bill separately. Web search in particular is charged per search on top of tokens. See the tools that cost extra.

Then sanity-check it against the revenue

If the feature is part of something you charge for, cost per user needs to be comfortably below what a user pays. If it isn't, that's not a pricing problem to solve later — it's a design problem to solve now, usually with a cheaper model, lower effort, caching, or a usage limit.

The number that should worry you isn't the monthly estimate. It's the answer to "what happens if someone scripts this?" Get that one before launch.