On 24 July Anthropic shipped Claude Opus 5. Going by the version number alone this is Opus 4.8 nudged along one notch, and there is not much to get excited about. The real change is not in the score column, it is in the price column — close-to-flagship capability dropped into the slot Opus already occupied. Against 4.8 at the same money it is a straight upgrade; against Fable 5 it is half the price for most of the work. The cost is a quiet one, and anyone who swaps the model ID and nothing else will walk into it.
The fourth model in under two months, and this one is worth a pause
Ground rules first. Every score below comes from the official release and public roundups. I have not run these benchmarks myself — ARC-AGI and GDPval are not the sort of thing you reproduce on a laptop. Benchmarks also shift between versions, so treat whatever the official page shows at the time as the number that counts. Read the direction, not the decimal.
Anthropic's release pace lately has been a little alarming: Mythos 5, Fable 5, Sonnet 5, then Opus 5 on 24 July — four models in under two months. The side effect is that people have started skipping release notes entirely, on the theory that there will be another one next month anyway.
This one is worth ten minutes, though, and not because it is stronger. It is because of which price tier it is strong at. Anthropic positions Opus 5 as frontier intelligence close to Fable 5 while staying at Opus pricing. In plain terms: the capability level that used to cost you the tier above is now mostly available one tier down. It is the new default on Claude Max, and the strongest thing you can reach on Pro.
A few specs worth keeping in your head: $5 per million input tokens and $25 per million output tokens; a 1M-token context window; three switchable thinking levels — low, medium and high — plus an xhigh setting. And one line that sits nowhere near the top of the spec sheet yet moves your bill more than any of them: thinking is on by default. That earns its own section, and it gets one — section four.
Against Opus 4.8: not a cent more, not one thing worse
Conclusion first: this is a same-price upgrade, and the kind you can move to without agonising over it.
Opus 4.8 landed on 28 May at $5 / $25. Opus 5 is identical. That makes this the cleanest of the two comparisons — no price-performance arithmetic to do, just the capability gap.
| Item | Opus 4.8 | Opus 5 |
|---|---|---|
| Pricing (per million tokens) | $5 in · $25 out | $5 in · $25 out |
| SWE-bench Pro | 69.2% | 79.2% |
| Reasoning benchmark average | 72.1 | 90.4 |
| Frontier-Bench v0.1 | Anthropic says Opus 5 delivers roughly twice the performance of 4.8, at a lower cost per task | |
| Benchmarks won in the public comparison | 0 | 6 |
SWE-bench Pro going from 69.2% to 79.2% is a full ten points, and that is the row anyone who writes code will feel first. The reasoning average, 90.4 against 72.1, opens up wider still. On Frontier-Bench v0.1 the official line is that Opus 5 delivers about twice the performance of 4.8, and spends less per task doing it.
The number I find most informative, though, is not any single score. It is the six-to-zero: in the public comparison Opus 5 comes out ahead on six benchmarks — BrowseComp, DeepSWE 1.1, FrontierCode 1.1, GDPval-AA, HealthBench Professional and Humanity's Last Exam — and Opus 4.8 comes out ahead on none.
Of those six you probably care about one or two. BrowseComp is web retrieval, DeepSWE and FrontierCode lean coding, GDPval-AA leans real occupational tasks, HealthBench Professional is medical, and Humanity's Last Exam is a deliberately brutal question set. So the useful part is not "won six". It is "lost zero".
The annoying part of a model swap was never that the new one is not strong enough. It is that some capability you quietly depend on slips half a step. Your prompts were tuned to the old model's temperament; the new one gains three points on a leaderboard you never look at and drops two on the step you run every day, and you find out when something breaks in production. Zero regressions settles the switch question better than any individual score does.
So if you are already on 4.8: the capability-side risk of moving is low. The cost side is not — keep reading.
Against Fable 5: half the money, and the gap is down to what?
This is the actual story of the release. Fable 5 is the tier above, at $10 per million input and $50 per million output. Opus 5 is $5 and $25 — exactly half.
| Comparison | Opus 5 | Fable 5 |
|---|---|---|
| Pricing (per million tokens) | $5 in · $25 out | $10 in · $50 out |
| Frontier-Bench | 43.3% | 33.7% |
| GDPval-AA v2 | 1,861 | 1,747 |
| Agentic task average | 90.8 | 84.6 |
| SWE-bench Pro | Within a point of each other, effectively level | |
| ARC-AGI 3 | Opus 5 reportedly around three times the next-best model — the widest single-item gap in this release | |
Every one of those rows says the same thing: on these measures Opus 5 has caught up with or passed the tier above, at half the price. Frontier-Bench 43.3% against 33.7%, GDPval-AA v2 at 1,861 against 1,747, agentic tasks averaging 90.8 against 84.6, and SWE-bench Pro inside a single point.
ARC-AGI 3 deserves its own paragraph. Opus 5 is reported at roughly three times the next-best model there, the widest single-item gap in this release. ARC-AGI measures not how much a model has learned but whether it can work something out cold: hand it a set of rules it has never seen and watch whether it finds the pattern. A high score does not mean it knows more, it means it gets stuck less often on unfamiliar problems.
Why does that matter to a normal user? What actually stops you day to day is rarely a question with a settled answer — current models handle those fine. It is the work nobody wrote a tutorial for, the kind you cannot fully describe yourself. That said, a benchmark sits a long way from whatever is on your desk; do not read three times on a chart as three times better in daily use.
Fable 5 has not been retired either. For the jobs that need the longest autonomous runs — let the model go for hours, plan and correct itself, no hands on the wheel — Anthropic still points at Fable 5. If your pattern is to throw a big task over the wall and go to sleep, that line decides it for you, and halving the unit price is not a reason to move.
One more that never shows up on a leaderboard but matters in practice: Opus 5's safety classifier is said to fire falsely around 85% less often than Fable 5's. For anyone doing security research, writing penetration-testing scripts, or handling financial and medical text, that can be worth more than any benchmark score — every time legitimate work gets refused you spend time rewriting the prompt to get around it, and that friction lands on no eval sheet anywhere.
Short version: against 4.8 it is the same price and nothing is worse. Against Fable 5 it is half the price and close enough for most work, except the very long autonomous runs.
Now the money trap: thinking is on by default
Thinking is on by default in Opus 5. The same prompt produces more output tokens than it did with thinking off, and output costs five times input ($25 against $5). Which gives you this: you changed one model ID, the price sheet did not move a cent, and the month-end bill went up.
This is not a defect. A default changed. Default changes never show up on the pricing page, only on the invoice, which is exactly why they slip past people.
The person most likely to get caught looks like this: they were running a batch of cheap work on Opus 4.8 with thinking switched off — classification, routing, squeezing long documents into summaries, format conversion. None of that needs the model to think much, and turning thinking off makes it fast and cheap. Then Opus 5 shows up, same price and stronger, so one line of code changes the model ID and nothing else moves. Capability did go up. So did output volume — on a unit price multiplied by five.
Four defences, in order of payoff:
- Compare output volume on a small sample before you switch. Take twenty or thirty of your real requests, run them through both models, write down the total output tokens. It takes ten minutes and it is the only way to see the bill change before it arrives.
- Match the effort level to the task instead of picking one for everything. Low, medium and high are there for you to use, with xhigh on top. Give bulk simple work the low end; move up when you need cross-file reasoning or you are chasing a hard bug. The test is not how important the task is, it is whether it needs the model to think a few more steps — for work that does not, extra thinking budget is just extra money.
- Watch output volume, not unit price. A request costs input volume × input price + output volume × output price. What changed here is the output-volume term; the price columns did not move at all, so the pricing page will never show you the problem.
- For genuinely high-volume work, look down the range. Pushing Opus 5 to its lowest thinking level to write summaries is usually worse value than moving to a cheaper model outright. A flagship earns its keep on hard problems, not on throughput.
A word on the 1M context window while we are here: a window is a ceiling, not a target. Dumping a whole repo or a year of logs into it costs real money on the input side, and the model will not necessarily be more accurate for it.
Three kinds of work, three choices
Compressing all of that into something you can act on:
Day-to-day coding, analysis, knowledge work → Opus 5. This tier now wins in both directions: same price as 4.8 with no regressions, half the price of Fable 5 with most of the capability. If you are on 4.8 already, switch — just run the output-volume check from the last section first.
Jobs where the model has to run on its own for a long time → Fable 5 is still that tier. Long chains, hours without a break, self-correction mid-run: still its home ground. Not many people work this way, but if you do, do not shop on unit price alone.
Bulk simple tasks → go down the range, not up. Classification, tagging, routing, templated summaries — those compete on unit cost and throughput, not on intelligence. Using a flagship here is pure waste, even dialled down to the lowest thinking level.
One general rule on top: add up total spend, not unit price. Two models at the same price, one talkative and one terse, can produce noticeably different invoices; the other way round, a model at twice the price that gets it right first time and saves you three rounds of rework may end up cheaper. What to watch is what the whole job cost, not the number on the price list.
For the crypto readers using AI as a work tool: reviewing order-placement and position-sizing code, checking API permissions and error handling — work where one mistake costs real money — is worth Opus 5. Compressing a day of announcements and posts into a summary belongs on a cheaper tier. Letting an agent watch data on its own all night is Fable 5's scene. The boundary has not moved, though: all of this is information handling, not judgement made on your behalf. A model can walk through how to read a candle chart with you and pick holes in your logic, but it does not know whether tomorrow is up or down; and the news it assembles for you can be invented, especially the parts that read most smoothly. Old scams recycled in an AI wrapper are something I wrote up in 8 crypto scams I have seen.
One last thing, and it may matter more than which model you pick: do not weld your workflow to a model ID. Four models in under two months means whatever you settle on today is unlikely to be the strongest or the cheapest a couple of months from now. Put the model ID, the thinking level and the cost ceiling in config so a switch is a one-line change; keep a small set of your own real tasks as a regression test and run it whenever something new ships — fifteen minutes tells you whether the switch is worth making. The last time I wrote about picking models it was the three GPT-5.6 tiers, and that piece's conclusion still mostly holds — not because the scores stayed put, but because tier the work and use the level that is just enough does not date much.
Questions you're probably about to ask
How much better is Claude Opus 5 than Opus 4.8?
Going by the official release and public roundups, SWE-bench Pro rises from 69.2% to 79.2%, the reasoning benchmark average is 90.4 against 72.1, and on Frontier-Bench v0.1 Anthropic says Opus 5 delivers roughly twice the performance of 4.8 at a lower cost per task. In the public comparison Opus 5 is ahead on six benchmarks and Opus 4.8 on none. Pricing is identical: $5 per million input tokens and $25 per million output tokens. I have not run these benchmarks myself, and benchmarks change between versions, so go by whatever the official page shows at the time.
Can Opus 5 replace Fable 5?
For most work yes, at half the price — Fable 5 is $10 per million input and $50 per million output, Opus 5 is $5 and $25. In the published figures Opus 5 scores 43.3% against 33.7% on Frontier-Bench, 1,861 against 1,747 on GDPval-AA v2, within a point on SWE-bench Pro, and 90.8 against 84.6 on agentic tasks. For the jobs that need the longest autonomous runs, though, Fable 5 is still the tier being pointed at.
Why did my bill go up after switching to Opus 5?
Most likely because thinking is on by default. With thinking on, the same prompt produces more output tokens, and output costs five times input. If you were running batch work on Opus 4.8 with thinking off and you changed only the model ID, the unit price stayed the same but the output volume did not, so the bill can rise instead of fall. Compare output token counts on a small sample before you switch, then decide which effort level to use.
Which effort level should I use for everyday coding?
Opus 5 offers low, medium and high thinking levels, plus xhigh. For everyday edits and small scripts the low end is usually enough; move up when you need cross-file reasoning, interface design or a hard bug hunt. The test is not how important the task is, it is whether it needs the model to think a few more steps — where it does not, a bigger thinking budget only costs more.
Read next: The three GPT-5.6 tiers · 8 crypto scams I have seen · What is Bitcoin
