Tool definitions cost tokens on every request
A big tool surface is a standing charge you pay per message.
When you give a model tools, the definitions travel with every single request. Names, descriptions, parameter schemas, all of it, resent each turn. Add a system prompt the API includes automatically to enable tool use, and you have a fixed cost on every message before anyone has said anything.
The standing charge
The automatic tool-use system prompt runs somewhere between roughly 280 and 800 tokens depending on the model, and it's only there if you provide at least one tool. On top of that come your own definitions, and those can be much larger.
Individual built-in tools add their own overhead. The bash tool adds a few hundred tokens. The text editor tool around 700. Computer use is in a different league, adding roughly 4,500 tokens just to declare, and browser use around 6,600.
Those last two are worth understanding before you reach for them casually. Declaring browser use costs more per request than many entire system prompts.
Why it compounds
It isn't the one-off. It's that this rides along on every turn of every conversation. A twenty-turn agent session pays it twenty times.
The fix is caching. Tool definitions sit at the very front of the request, ahead of the system prompt, which makes them the most reliably cacheable thing you have. Cached, that standing charge drops to a tenth, or a fortieth on Fable 5.1.
If you take one thing from this: if you use tools and you aren't caching, start there.
Keeping the surface small
Beyond caching, fewer tools is cheaper and usually works better. A model choosing between six well-described tools makes better choices than one choosing between thirty.
Here are my tool definitions: <paste> For each: how many tokens is it, and when was it last actually called? Then tell me: - Which could be removed entirely. - Which two or three could be merged into one with a mode parameter. - Which descriptions are longer than they need to be without losing the "when NOT to use this" guidance. - Whether my tool block is positioned to be cached. I care about both the token cost and the model choosing correctly.
The order that matters
Requests are assembled tools first, then system prompt, then messages. Since caching is a prefix match, anything unstable early in that order invalidates everything after it. A tool list built in a different order each run — iterating over a hash map, say — silently destroys your cache for the whole request.
Sort your tool list. It's a one-line fix for a bill that looks inexplicable.
Thirty tools is usually a sign the agent is doing too many jobs. Split the agent before you optimise the token count.