Context windows, and what fills them
A million tokens sounds infinite. It isn't, and it isn't free.
The context window is how much the model can hold in view at once: your system prompt, the whole conversation so far, every file you pasted, every tool result, and the answer it's generating.
The sizes
Current Claude models from 4.6 onward have a 1M token context window. Haiku 4.5 has 200,000.
A million tokens is roughly 750,000 words, or a decent-sized codebase. It's a lot. It is not unlimited, and more importantly, filling it is not free.
Large context has no premium, but it still costs
On current models the full window is available at standard rates. A 900,000-token request bills at the same per-token rate as a 9,000-token one.
That's the good news. The bad news is arithmetic: it's a hundred times the tokens, so it costs a hundred times as much. "No premium" means no surcharge, not no cost.
What actually fills it
In a vibe coding session, in rough order of appetite:
Pasted files. Obviously, and it's usually the biggest item.
Tool results. A command that dumps a long log, a search that returns twenty documents, a fetched PDF. These accumulate turn after turn and nobody watches them.
The conversation itself. Every previous message is resent every turn. A long session carries all of its own history.
Thinking. It's generated, billed as output, and on models that keep it in context it takes up room.
What to do as it fills
Start a fresh conversation. The most effective move and the most underused. A new session with a good project brief beats a long one carrying every wrong turn you took.
Cache the stable part so the unchanging prefix costs a tenth.
Don't paste what you don't need. Three relevant files beat the whole repository, for cost and for quality.
Use compaction or context editing if you're building on the API. One summarises older history, the other clears old tool results outright.
The quality side
Cost isn't the only reason to keep context tight. A conversation stuffed with abandoned approaches and superseded versions gives the model more wrong material to weigh. Shorter, cleaner context tends to produce better answers as well as cheaper ones.
This conversation is getting long and expensive. Write me a handover brief for a fresh session: what we're building, the stack and constraints, decisions made and why — especially anything we tried and rejected — what works now, what doesn't, and the exact next step. Write it for someone who has never seen this conversation. Facts only.
Running out of context is rarely the real problem. Long before that, quality drops and the bill climbs. Start fresh sessions earlier than you think you need to.