- Genial
- AI Automation Tutorials
- Is Claude Sonnet 5 Actually Worth It?
Is Claude Sonnet 5 Actually Worth It?
Anthropic just released Sonnet 5, the most powerful version of Sonnet yet, and on paper it looks like a clean win. In this walkthrough I go through the benchmarks Anthropic published, then the pricing charts that tell a very different story, and I finish with a clear recommendation on what you should actually run today. By the end you will understand why a model can beat its predecessor on every metric and still be hard to justify for business use.
1. The benchmarks: Sonnet 5 vs 4.6 and Opus 4.8 (0:16)
Starting at 0:16, I walk through the first set of benchmarks Anthropic released with the model. The short version: Sonnet 5 shows an increase across every single metric possible. Computer usage, agentic coding, general reasoning, all of it goes up. Sonnet 5 beats Sonnet 4.6 across the board, flat out.
The more interesting detail is where it lands relative to the bigger models. Everywhere except agentic coding, Sonnet 5 gets near the level of Opus 4.8, which used to be the best model before Fable 5 arrived. In theory that should be perfect. You would have a model that is just about as intelligent as Opus 4.8 while costing less per token. That is the whole promise of the Sonnet family: most of the intelligence at a fraction of the price.
But that is not really the case, and the next chart is where the story falls apart.
2. The pricing problem (0:58)
At 0:58 I move to another graph released by Anthropic itself, and this is the part you need to look at before you switch anything over. It compares the performance of Sonnet 5 and Opus 4.8, and the problem is one specific region of that chart.
In theory you are paying for a cheaper model, because you pay less per token. But when you look at the cost per task, you end up paying about the same price as Opus 4.8. That makes Sonnet 5 a very weird model, because you do not really know where to put it. It is meant to be cheaper than Opus 4.8, but in practice it is just as expensive.
A graph shared on X shows exactly the same phenomenon from a different angle. Sonnet 5 cost two dollars and thirty cents for a given task, while GPT 5.5 cost only one dollar for the same thing. That is more than a two times increase in pricing against the competition.
Why does this happen? Because per-token pricing and per-task pricing are not the same thing. Sonnet 5 is a less intelligent model than Opus 4.8, so it has to think for longer before it actually does the task. It spends more tokens getting to the same place. To be fair, going back to Anthropic's own benchmarks, the model also misbehaves less than the previous one, and that is genuinely great. But you are paying for that in performance, and the token burn erases the discount that makes Sonnet worth choosing in the first place.
3. The verdict: an identity crisis (2:42)
At 2:42 I give you my honest read. I cannot say this is a good model. I would even go as far as saying it is a bad model, or at least one with an identity crisis, because you cannot use it for business. If a task through Sonnet 5 costs roughly what the same task costs through Opus 4.8, you might as well just use Opus 4.8 and get the extra intelligence for free.
And that leads to the bigger question. If this does not get fixed within the next few days, Sonnet 5 and the Sonnet family do not really have a reason to exist. You might as well have a lineup of Opus 4.8, Fable 5, and Haiku 5 and cover every use case cleanly: a top model, a frontier model, and a cheap fast one. The middle tier only makes sense when it is actually cheaper in practice, not just on the pricing page.
4. What to use instead (3:13)
From 3:13, my recommendation is simple. If you are someone who runs on Sonnet today, stay with Sonnet 4.6 for the time being, until this situation gets fixed. Maybe the fix comes with an Opus 5.1 or something along those lines, but I would not migrate my workloads onto Sonnet 5 as it stands.
Pitfalls and tips
- Always take launch metrics with a grain of salt. It is marketing as usual, and every provider publishes charts that make their model look best. What makes this case unusual is that even in the marketing material, the pricing problem is visible.
- Compare cost per task, not cost per token. A cheaper per-token price means nothing if the model thinks longer and burns more tokens to finish the same job.
- Do not upgrade just because the version number went up. Benchmarks going up across the board did not make this model a better deal for business use.
Where to go next
If you follow model releases, watch my breakdown of OpenAI's answer in GPT 5.6 Just Dropped: Sol, Terra & Luna Explained. And if you want to get more out of whichever Claude model you run, learn to package your daily workflows in How to Build Your First Claude Skill.


