- Genial
- Tutorial su IA e automazione
- Claude Opus 5: Fable 5 Killer or Benchmark Hype?
Claude Opus 5: Fable 5 Killer or Benchmark Hype?
Claude Opus 5 just released, and according to the benchmarks it is supposed to be a Fable 5 killer. Instead of taking the charts at face value, I ran my own tests: a 3D solar system build and a security audit of a real open-source app, comparing Opus 5 and Fable 5 head to head. The results say a lot about guardrails, benchmarks, and why these two models actually complement each other.
1. What the benchmarks say (0:12)
On the Anthropic release page, Opus 5 sits all the way to the left, then Fable 5, then GPT 5.6. Right now, Kimi K3 and Fable 5 are the best models out there, with GPT 5.6 Sol trailing slightly behind them. And the charts show an increase across the entire board for Opus 5, both on coding tasks and business tasks.
2. Why benchmarks are not enough (0:43)
A model being higher on a benchmark does not make it better; you have to test capacities with real use cases. And frankly, it does not make a lot of sense for Opus 5 to beat Fable 5. Fable 5, or Mythos as it was hyped up for such a long time, was positioned as the end all be all for cybersecurity and coding. Releasing a better model that is also cheaper would mean Anthropic killing their own flagship. So the charts describe a world where Opus 5 sits above Fable 5, and my job was to see if that holds up. The model is very young, so we will know its true performance in the coming days, but there is already enough to compare.
3. Test 1: 3D solar system (1:43)
The first test is a simple 3D representation I always like to run, because it gives me a grasp of how a model performs on 3D tasks:
Create a 3D representation of the entire solar system with interactions.
The result: I can click on a planet, pan the view, click on another planet. The design is still very good. But compared to the exact same use case on Fable 5, which produced a lot more interactions and a lot more elements inside the actual website, Opus 5 is not quite at that level. It is a definite improvement over Opus 4.8, though.
4. Test 2: security audit of a real app (2:35)
The second test is coding and cybersecurity oriented. I am currently building an open-source application that is pretty much a WisprFlow alternative, so I ran an audit of it with both models, Opus 5 and Fable 5, using nearly the exact same prompt asking each to find vulnerabilities in the code.
This is where the two models start working together, because something interesting happened.
5. Guardrails: where Fable 5 refuses (3:02)
Fable 5 got triggered by my prompt three different times. Asking it to find vulnerabilities in my own application made it fall back to Opus 4.8, as it usually does. That is a real problem with Fable 5: for anything related to biosciences, or even finding cybersecurity risks in your own app, it is very hard to let Fable 5 do the work by itself.
Opus 5, with the exact same prompt, pinpointed bugs that Fable 5 could not surface. In the report, Opus 5 found two cybersecurity bugs, with another eight in total that were not very urgent. Fable 5 actually managed to find more bugs overall, but in parts of the application where Opus found none. The explanation is that Opus 5 ships with a lower security system than Fable 5, what we call guardrails. Opus 5 finds things Fable 5 does not, not because Fable 5 is not good enough, but because we are limiting the capacities of the model.
As of today, that makes Opus 5 the interesting choice for finding threats in your applications, or for any work where medical data is involved, even simple things like the weight you log every morning.
6. Two complementary models (4:52)
Put the results together and the picture is clear: Fable 5 found more bugs in UX and UI, Opus 5 found the security bugs. These models complement one another. Fable 5 is still the undisputed king for UI, UX, and simply breaking down complex tasks.
7. My verdict on the benchmarks (5:26)
I firmly believe Anthropic's benchmarks should be taken with a grain of salt, because shipping a model that beats Fable 5 at half the cost literally makes no sense for them as a company. Based on what I could run so far, Opus 5 feels like a decent model to use in combination with Fable 5, not a replacement for it.
What's next (5:42)
I am running more tests overnight and will show you how to actually put this model to work. In the meantime, treat Opus 5 as a partner for Fable 5, not a killer.
Pitfalls and tips
- Do not trust launch-day charts. Run the model on your own real use cases before switching anything over.
- If Fable 5 keeps refusing a legitimate audit of your own code, hand that task to Opus 5 instead of fighting the guardrails.
- Remember the model is days old. Where it stands against GPT 5.6 Sol, Grok 4.5, and Fable 5 on long-running tasks is still unknown.
Where to go next
- Get the most out of both models with the delegation setup in Claude Fable 5 Is Here to Stay - Maximize Your Usage.
- For another case of benchmarks vs reality, watch Gemini 3.6 Flash: Don't Believe the Benchmarks.


