Measure cost per task, not per token
The cheaper model that needs four attempts is the expensive one.
Every pricing page is per token, so that's how people compare models. It's the wrong unit. What you actually care about is what it costs to get a working result.
The arithmetic people skip
Suppose a task takes one attempt on a model at $5 per million input tokens, and four attempts on one at $2. The cheaper model costs more, before counting your own time re-reading four wrong answers and the risk that you ship the fourth one without noticing it's still wrong.
The reverse happens too. If a task genuinely takes one attempt on either, the cheaper model is simply cheaper and using the expensive one is waste.
You cannot tell which case you're in without measuring.
How to measure it
Take five real tasks from your project. Run each on both candidates. Record for each: did it work, how many attempts, total tokens in and out.
Then divide total cost by number of tasks that actually came out right. That number is what you compare.
I'm comparing <model A> against <model B> for <the kind of work>. Here are five real tasks: <paste> For each, I'll run both and give you: whether the result was usable, how many attempts it took, and the token counts. Then work out cost per successful task for each model, show the arithmetic, and tell me which is actually cheaper for this workload. Include the ones that never came out right in the cost.
That last instruction matters. Failed attempts still cost money, and leaving them out of the total is how a weak model looks cheap.
The other costs in the denominator
Your time. Twenty minutes re-prompting is worth more than the token difference on any realistic volume.
Turns. An agent that needs more round trips to reach the same place burns tokens on every one, because the whole conversation is resent each time.
Undetected errors. The most expensive outcome is a wrong answer you accept. No token accounting captures that, which is exactly why it deserves the most weight.
The practical upshot
For most personal projects, the honest answer is that model cost is not your problem and you should use the model that gets it right. Cost per task starts mattering when you're running something at volume, or leaving agents unattended for hours.
If you can't say what a finished task costs, you can't say which model is cheaper. Per-token pricing is an input to that question, not an answer.