If you're paying for API usage and the number is bigger than you expected, guessing is a poor way to find out why. The data is there.

The three questions worth asking

Which model? If you assumed you were on a cheap model and you're not, that's the whole answer. Coding tools sometimes default differently than you remember, and a model set once in a config six weeks ago is easy to forget.

Input or output? Output bills at five times input. If output dominates, the problem is verbose responses, high effort, or long thinking. If input dominates, it's context: files being re-pasted, long conversations, tool results piling up.

Cached or not? If your cache read count is zero across requests that should share a prefix, caching isn't working, and that's often the single biggest recoverable number on the page.

The two numbers that matter most

Cache read tokens and output tokens. Between them they explain most surprising bills.

A high output number with low effort settings usually means the model is being asked open-ended questions that invite long answers. A zero cache read number on a repetitive workload means something in your prefix is changing between requests. See caching economics.

Instrument your own side too

The dashboard tells you the total. It doesn't tell you which of your features spent it. If you're running more than one thing against the API, log the usage figures returned with each response alongside whatever identifies the feature.

usage-logging.txt
Add usage logging to my model calls: <paste the call site>

For each request, record: which feature made it, the model, effort
level, input tokens, output tokens, cache read tokens, cache write
tokens, and whether it succeeded.

Write it somewhere I can aggregate — a table or a log I can query.
Don't log prompt or response content, just the counts.

Then show me a query answering: what did each feature cost me last
week, and what was the cache hit rate per feature?

Excluding the content is deliberate. You want the numbers, not a copy of everything your users typed.

On a subscription

If you're using a coding tool on a subscription rather than the API, you won't have per-token figures. What you have is a usage allowance, and the same logic applies to it: model choice and effort level are what consume it fastest.

"The bill went up" is not a diagnosis. Ten minutes with the dashboard usually turns it into one specific answer.