- Genial
- Tutorial su IA e automazione
- Gemini 3.6 Flash: Don't Believe the Benchmarks
Gemini 3.6 Flash: Don't Believe the Benchmarks
Google just announced three new Gemini models in a single drop, and it seems they are in trouble. In this breakdown I walk through what actually shipped, read the benchmark data the way you should read it, and tell you what I would use instead. By the end you will know whether any of these releases deserve a place in your stack.
1. Google's new Gemini drop (0:10)
The newest models, released just today, are Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber. That last one is not really a model per se, just a version of Gemini Flash oriented towards cybersecurity. Of course the launch benchmarks say they perform better and faster than the last generation; you should expect that from every model. The true question is whether these models keep up with the level of performance we have now, and the costs. That is where it gets tricky.
2. Gemini's model lineup explained (0:47)
Gemini ships three sizes. The smallest is Flash Lite, and it is probably my favorite of all of them because it is extremely affordable and you can use it for a lot of analysis at scale. The one superpower left to the Gemini family is that it is still the only AI that can natively analyze videos, documents, and anything you throw at it. It really sees, with incredibly powerful OCR capacities, meaning it can visualize what is in a document perfectly. Then there is Flash in the middle, and Pro at the top.
3. Where is Gemini Pro? (1:28)
Pro, the biggest of the three, was a contestant back when we still had Opus 4.5. Gemini 3.1 Pro was an excellent model, but there has been no 3.5 Pro and no 3.6 Pro in a while now. We do not even know where they are, really. Which raises the question: how do the models we did get compare to the rest of the market? The news is not good.
4. Benchmarks: 3.5 Flash vs 3.6 Flash (2:02)
On the Artificial Analysis chart, 3.5 Flash and 3.6 Flash sit at exactly the same level of performance. According to that source, the new release brings no big improvement. Gemini 3.5 Flash Lite sits way behind in the charts, which is normal: it is a smaller, cheaper model, so that position makes sense. The flat line at the top is the worrying part.
5. How Gemini compares to state of the art (2:52)
Things get worse when you look at the state-of-the-art benchmarks: Kimi K3, GLM, Fable, Sol. On the intelligence level, 3.6 Flash hits a score of fifty. The top five, Fable, Sol, Kimi K3, Grok, and GLM, sit ten to twenty percent higher. Yes, Gemini 3.6 Flash is still the fastest model of them all. But if you are building applications, you do not really care about having the fastest model. You care about the best quality at around sixty tokens per second.
6. Cost per task: Grok 4.5 wins (3:57)
The cost per task is what dictates how much it costs you to get stuff done with your AI, and this is where it stings. Grok 4.5 costs about forty percent less per task and has a higher level of intelligence; for reference, Grok 4.5 sits at a level similar to Opus 4.8. So there is literally a model that performs better and is also cheaper. Google has fallen severely behind in the AI race, despite being the company with access to all the resources needed to keep going: the hardware, the software, and the data.
7. What should you use instead? (4:33)
I am not going to recommend the 3.6 model, because you have better options elsewhere via Grok. If you want an open-source model, GLM is perfectly fine. Using Flash Lite is still good for cheap analysis at scale, but even there, nothing has really changed with this release. The real hope is that Gemini 3.5 Pro or 3.6 Pro change this picture, but things are not looking very good for Google right now.
Final verdict (5:36)
Stay up on the AI news, but know how to filter it. In this case, it was just a bunch of noise.
Pitfalls and tips
- Launch benchmarks always show an improvement over the previous generation. Compare against the current market, not against the vendor's own last release.
- Fastest is not best. Judge models on quality at a usable speed, around sixty tokens per second, and on cost per task.
- Do not throw away Flash Lite for bulk analysis work; the price is still the point. Just do not expect frontier intelligence from it.
Where to go next
- See how I pressure-test a launch instead of trusting charts in Claude Opus 5: Fable 5 Killer or Benchmark Hype?.
- Gemini's one real superpower in action: Stop Wasting Hours: Analyze Any Video with AI.


