What a vibe-coded project actually costs
Worked examples, from a weekend build to something with users.
Abstract pricing is hard to reason about. Here are concrete shapes, with the arithmetic visible so you can substitute your own numbers.
All figures use September 2026 first-party API rates. Check current pricing before relying on any of it.
A weekend project, built in a coding tool
You're on a subscription, so there's no per-token bill. What you're spending is your allowance, and the two things that consume it fastest are model choice and effort level.
The practical advice is the same as if you were paying: pick deliberately per task rather than leaving it at maximum all weekend. See subscription or API.
A tool that calls a model for its users
Say a summariser. Each use sends about 4,000 input tokens and generates about 600 output tokens, one call per use.
On Sonnet 5, at $2 in and $10 out per million:
- Input: 4,000 × $2 ÷ 1,000,000 = $0.008
- Output: 600 × $10 ÷ 1,000,000 = $0.006
- Total: about $0.014 per use
A hundred uses a day is roughly $42 a month. Ten thousand a day is roughly $4,200 a month, which is the number that tells you whether you have a business or an expensive hobby.
Note how close input and output are here despite output being five times the rate. That's because the summary is short. If it were as long as the input, output would dominate entirely.
The same tool with caching
If every request shares a large stable system prompt, cache it. Cached reads cost a tenth of the input rate on most models. On a workload where most of the input is that stable prefix, the input line falls by roughly 90%.
Whether that matters depends on your input-to-output ratio. In the example above, halving input saves less than a quarter of the total. In a retrieval feature sending 50,000 tokens of documents each time, it's the whole bill.
A chat feature
Chat is different because every turn resends the conversation. Ten turns doesn't cost ten times one turn; it costs roughly the sum of a growing number, which is considerably more.
This is why caching and conversation limits matter far more for chat than for one-shot features.
A long agent run
An hour of agent work might move a few hundred thousand tokens across many turns, much of it re-sent context. On Opus 5 that's low single-digit dollars for the hour. On Fable 5.1, roughly double.
Unattended overnight runs are where this stops being trivial, and where a task budget or spend cap earns its place.
Do yours
Work out what my project costs to run. What it does: <describe> Model and effort: <model>, <level> Per use it sends roughly: <n> input tokens, generates <n> output tokens Calls per use: <n> Expected volume: <n> per day Show the arithmetic. Then: monthly cost at 10x and 100x that volume, which line grows fastest, and the cheapest change that halves it.
For a personal project the answer is almost always "less than you feared". For anything public the answer is "it depends entirely on whether you capped it".