- Genial
- Tutorial su IA e automazione
- Why I Went Back to Opus 4.6
Why I Went Back to Opus 4.6
Anthropic shipped Opus 4.7 and everyone assumed it was an automatic upgrade. I tested it, ran into a wall with the new adaptive thinking mode, and rolled my daily workflow back to Opus 4.6. In this walkthrough I explain exactly what changed under the hood, why a smarter model does not guarantee better results, and the testing habits I now apply to every new model release.
What you need
- A Claude account with access to the model selector, so you can switch between Opus 4.7 and Opus 4.6
- Optionally Claude Code, where reasoning is still more controllable than in the chat application
1. What is Opus 4.7? (0:09)
Claude Opus 4.7 was released by Anthropic on April 16th, not even a week before I recorded this. On paper it is their best model yet. If you scroll through the press release, their benchmarks show it performing better overall than any model out there, whether that is Opus 4.6 or the competitors, Gemini by Google and GPT by OpenAI.
Here is the first thing to keep in mind: always take it with a grain of salt when a provider publishes its own benchmark. You do not know what is happening behind the scenes, and you get no explanation of how things were measured. I am not saying Opus 4.7 is a bad model. It is a fantastic model, and honestly all the frontier models are great at most cognitive tasks. My point is different: it is not because a model is the latest that you should adopt it immediately.
2. A smarter model does not mean better results (1:06)
A lot of people assume that a smarter model automatically produces better outcomes. That is false. Right now, any of these models can take on most cognitive tasks. They are all smart enough to run analyses, work through data, build dashboards, you name it.
What really makes the difference is being able to connect your model to the world. Most of the actual productivity increases coming from AI come from the features that surround the models, not from the raw model itself. Keep that frame in mind for the rest of this walkthrough, because it explains why a technically stronger model can still feel worse in practice.
3. The backlash on Reddit (1:52)
There has been huge backlash against 4.7. It is all over Reddit right now, with a lot of people complaining that this release is an actual regression, not an upgrade. Threads keep asking the same question: 4.6 versus 4.7. Adding fuel to the fire, 4.6 got what users called a quote-unquote downgrade right before the release of 4.7, which raised a lot of questions on its own.
The interesting part is that the problem is not the model's brains. It is a feature that shipped alongside it, called adaptive mode, and it applies both to Opus 4.7 in the chat application and in Claude Code.
4. How a conversation actually works (2:52)
To understand why adaptive thinking matters, picture your conversation as a white sheet of paper. Inside it there are things you never see as a user. First there is the system prompt: the way Anthropic tells Claude that he is Claude, that he is a helpful assistant, that he should be friendly and warm. It is invisible, but it is there.
Then you start asking questions. Depending on their difficulty, Claude reasons through the problem before answering. Ask him how the weather is today and there is nothing to think about. Ask him to analyze five books and highlight the common denominators between all of them, and there is real cognitive work involved. That internal deliberation is what we call reasoning tokens. You do not see them, but they exist, the same way you talk to yourself when you work through a problem: should I do A, B, C, or maybe a combination of both? And crucially, those reasoning tokens used to be something you could control.
5. The problem: you cannot control reasoning anymore (4:29)
Before, you could tell Claude: I am going to ask you some really difficult questions, so take those reasoning tokens and put as much reasoning as you can into the task. With adaptive thinking, you can no longer do that. Claude now classifies your question himself into easy, medium, or hard, and he does not really follow your instructions on it. Sometimes you want him to reason deeply, and he takes the easy route anyway. That hinders the quality of the answer you get.
You can still change the adaptive thinking level, but in the chat application you cannot force him to think a lot. In Claude Code you can, which is one more reason the chat experience is where the regression is felt most.
6. Extended thinking on 4.6 vs adaptive thinking on 4.7 (5:38)
Now switch back to the older models and select Opus 4.6. You will notice the feature is not called adaptive thinking there. It is called extended thinking, and the difference is exactly the point:
Opus 4.7: Adaptive thinking (the model decides how much to reason)
Opus 4.6: Extended thinking (you decide when it reasons more or less)
With 4.6 you choose whenever you want the model to reason more or less depending on the task. That is one of the main reasons 4.7 feels like a downgrade to so many people: you can no longer steer it the way you could before.
7. Why I went back, and how to track model performance (6:24)
I tried 4.7 and I regressed back to 4.6, simply because the adaptive behavior kept causing trouble in my usual workflow. The key lesson here goes beyond this one release: keep track of when you were using your models. If 4.6 felt smarter on Friday and dumber on Tuesday, there can be a real reason behind it. Models carry timestamps. A 4.6 snapshot created on July 1st and one created on September 3rd can behave differently, even inside the same family. If you ever program with these tools, you can select the precise version that gave you your best performance and pin it.
8. How to test new models properly (7:48)
The reflex to build: as soon as a new model drops, run your own tests, and do not draw conclusions without at least two to three days of usage. Do not fall into the AI hype. The jumps between recent versions are really insignificant. These models are already way smarter than we are. Most of the improved performance will not come from a slightly bigger brain; it will come from how well you connect that brain to your world.
Pitfalls and tips
- Provider benchmarks are marketing until proven otherwise. You cannot see how they were measured, so test on your own workflow.
- Do not switch your production workflow to a new model on day one. Give it two to three days of real testing first.
- Write down when your model performs best. Version timestamps matter, and pinning a specific snapshot is possible when you build with the API.
- If you need heavy reasoning on demand today, Opus 4.6 with extended thinking still gives you that control in the chat application.
Where to go next
If you are weighing up which assistant deserves your subscription, watch my full comparison in ChatGPT vs Claude: Which AI Tool Should You Pick?. And if you want the connectivity gains I keep talking about, start with Claude Code for Beginners (2026 Guide).


