Why the better model is not the fix
A stronger model changes how often you get a bad run. It changes nothing about what a bad run can do.
Everyone reaches the same conclusion after their first bad run, and it is the wrong one: the model was not good enough. Buy a better one, and this stops happening.
It is worth taking seriously, because it is half true — and then looking at what the money actually buys.
| Tier | Model | Input | Output |
|---|---|---|---|
| fast | gpt-4.1-nano | 19 | 73 |
| balanced | gpt-4.1-mini | 73 | 292 |
| complex | GPT-5.4 | 455 | 2,730 |
fast
- Model
- gpt-4.1-nano
- Input
- 19
- Output
- 73
balanced
- Model
- gpt-4.1-mini
- Input
- 73
- Output
- 292
complex
- Model
- GPT-5.4
- Input
- 455
- Output
- 2,730
What the upgrade actually changes
The agent that moved ₹1,20,000 onto a stranger’s account was the middle tier, and its written reasoning was correct. It did not misunderstand the policy. Nothing about a stronger model addresses the failure that occurred, because the failure was not in the analysis.
And there is a harder piece of evidence than that argument. The same model ran the same ticket twice and produced two different failures. Upgrading moves the distribution of outcomes. It does not truncate it — there is no tier at which the bad tail becomes empty, and you cannot buy your way to a guarantee that is not on sale.
| What a stronger model gives you | What it leaves exactly where it was | |
|---|---|---|
| Quality of reasoning | better analysis, fewer misreadings, more reliable structure | — |
| Authority | nothing | it still holds every tool you attached, with the same permissions |
| Exposure | nothing | the worst thing one bad run can do is unchanged, because that is set by the tool list |
| Auditability | nothing | a better explanation is still an explanation, and still not evidence |
Quality of reasoning
- What a stronger model gives you
- better analysis, fewer misreadings, more reliable structure
- What it leaves exactly where it was
- —
Authority
- What a stronger model gives you
- nothing
- What it leaves exactly where it was
- it still holds every tool you attached, with the same permissions
Exposure
- What a stronger model gives you
- nothing
- What it leaves exactly where it was
- the worst thing one bad run can do is unchanged, because that is set by the tool list
Auditability
- What a stronger model gives you
- nothing
- What it leaves exactly where it was
- a better explanation is still an explanation, and still not evidence
The limit case, from this build. The reviewer in module 4 runs at temperature zero and it still approved a proposal that named a colleague who does not exist. No model could have caught that one — the reviewer never sees the tool results, only the proposal, so it is structurally incapable of checking a fact. That is an information problem, and capability does not fix information problems.
Where to spend instead
The comparison people actually face is not nano versus the flagship. It is: one expensive agent acting alone, or two cheaper ones with a gate between them. The second is usually better on both axes at once — a proposer and an independent reviewer at the middle tier cost a fraction of one top-tier pass, and the thing you were trying to buy with the upgrade was never in the model.
A defensible policy for choosing a tier. Use the cheapest model whose reasoning you have tested on your own hard cases, and put the money you saved into the control layer. Upgrade when the analysis is wrong, not when the outcome was — those are different diagnoses and only one of them is a model problem.
Prices move. The table above was read from the platform’s rate table on 3 August 2026 and is here to show the shape of the gap between tiers, not to be quoted. Check the current figures before you build a business case on them.
Related lessons
Base rates — what a piece of evidence is actually worth
A face-recognition system that is 99.9% accurate and almost entirely wrong, and a number that sent an innocent woman to prison. Both are the same arithmetic, and it is the arithmetic that decides what any piece of evidence is worth.
ReadConfirmation and survivorship — what you never looked for
Two questions about evidence you did not go looking for. One is a rule you have to discover, and one is a pattern in five famous people — and in both, the thing that would have told you the truth is the thing nobody checks.
ReadLoss aversion, sunk cost and regression — what it costs you
Four questions you answer about yourself rather than about a scenario, and your own answers are the finding. Then the pattern that makes praise look useless and criticism look like it works, whatever you actually do.
Read
