Current Claude models take up to a million tokens of context, and unlike some earlier arrangements there's no premium for using the top of that range. A 900,000-token request bills at the same rate per token as a 9,000-token one. Caching and batch discounts apply across the whole window too.

That's genuinely good, and it's also where people stop reading and reach a wrong conclusion.

Same rate, hundred times the tokens

No surcharge means no multiplier. It doesn't mean it's cheap. A hundred times the tokens costs a hundred times as much, at the same rate.

At Sonnet 5's input rate, a full million-token request costs about two dollars before the model has written a word. On Fable 5.1 it's ten. Once per day, unremarkable. On every turn of a conversation, a different conversation.

The part that actually bites

Conversations resend everything. If you fill the window early and then have a twenty-turn exchange, you pay for that context twenty times.

This is where large-context vibe coding gets expensive, and it's almost never the single big paste that does it. It's the paste plus the twenty turns afterwards.

Caching is the direct fix: put the big stable material at the front, cache it, and the repeat cost drops to a tenth, or a fortieth on Fable 5.1. Uncached, a long context in a long conversation is the most expensive pattern available.

Whether you need it at all

Usually not. Pasting an entire repository feels thorough and is usually worse than pasting the four relevant files, for quality as well as cost. The model has more to weigh, and more chance to anchor on something irrelevant.

context-trim.txt
I've been pasting <describe what — a whole repo, a long document> into
every session.

Here's the structure: <paste the file list>
Here's what I'm actually working on: <describe>

Which of these do you actually need to do this work? Which am I
including out of caution?

Give me the shortest set that doesn't lose anything important, and tell
me what you'd ask for if you hit something you needed.

That last line is the point. The model can ask for more. It cannot un-see something irrelevant you pasted at the start of a long session.

When the big window earns it

Genuinely large single documents. Whole-codebase questions asked once rather than iterated on. Long agent runs where the history is the work. In all of those, cache the stable part and keep the varying part small.

The million-token window is a capability, not a default. Using all of it because you can is the most expensive habit on this list.