Gemini 3.6 Flash: Don't Believe the Benchmarks — Transcript
Full transcript
- 0:00Hey, what's up, everyone? My name's
- 0:02Miguel. Now, Google just announced the
- 0:04dropping of three new Gemini models.
- 0:07However, it seems that they're in
- 0:08trouble.
- 0:09Let me break it down for you.
- 0:10We've got the newest models released
- 0:12just today, Gemini 3.6 Flash, 3.5 Flash
- 0:16Light, and 3.5 Flash Cyber. Now, this
- 0:19last one is not really a model per se,
- 0:21it's just a version of Gemini Flash more
- 0:25oriented towards cybersecurity.
- 0:27Of course, if we look at the benchmarks,
- 0:29they're going to tell you that they're
- 0:30performing better and faster than the
- 0:33the last generations. That's something
- 0:34that you should expect from every model,
- 0:36but the true question is, are these
- 0:38models really keeping up with the level
- 0:41of performance that we have now, as well
- 0:43as the costs?
- 0:44And that's where it actually gets
- 0:46tricky. So, for information, Gemini has
- 0:48three different versions of the models.
- 0:51You have the smallest one, which is
- 0:52called Flash Light. This model is
- 0:55probably my favorite model out of all of
- 0:57them because it's extremely affordable,
- 0:59and you can use it for a lot of analysis
- 1:01at scale.
- 1:03You see, the one superpower that's left
- 1:05for all of the Gemini models is that
- 1:07it's still the only AI that can really
- 1:10analyze videos, documents, and anything
- 1:13that you throw at it
- 1:15at a native level. So, it really sees
- 1:18and has
- 1:19incredibly powerful OCR capacities. That
- 1:23means that it can actually visualize the
- 1:24things that you that you have in a
- 1:26document perfectly.
- 1:28Now, Pro, the biggest one out of the
- 1:30three
- 1:31was a contestant when we still had Opus
- 1:344.5, I believe.
- 1:37Gemini 3.1 Pro was an excellent model,
- 1:41but we haven't had a release of 3.1 3.5
- 1:45Pro nor 3.6 Pro in a while now. We don't
- 1:49even know where they are, really.
- 1:51And that makes the question,
- 1:54well, how are these two performing
- 1:56compared to the rest of the market? And
- 1:59well, the news are not really that good.
- 2:02So, as you can see right here on this
- 2:03chart, we have the level of performance
- 2:06of 3.5
- 2:09flash and 3.6
- 2:11flash
- 2:12at exactly the same level.
- 2:16So, that means that these model,
- 2:17according to artificial analysis, the
- 2:20source where this graph comes from,
- 2:23doesn't really have a big improvement in
- 2:25terms of performance.
- 2:26Now, not only that, but of course,
- 2:28Gemini 3.5 flashlight, which is the
- 2:31lower end of the models, is still
- 2:34sitting way, way behind in the charts.
- 2:36That's normal. It's a slower model. It's
- 2:39a smaller model. So, it actually makes
- 2:42sense for it to be there.
- 2:43But, where things get actually worrisome
- 2:46is whenever you actually look in detail
- 2:49at what's happening in the back. So, as
- 2:51you can see right here, we have the
- 2:53benchmarks of the state-of-the-art
- 2:55models, Chimi K3, GLM,
- 2:58Fable, Soul, etc. And this is where we
- 3:02see
- 3:03where Gemini actually stands. So, on the
- 3:06intelligence level, we see that 3.6
- 3:09flash hits a 50,
- 3:11score of 50.
- 3:13However, whenever you actually reach
- 3:16the higher end
- 3:18of the models, so we're going to call
- 3:20this the top five, Fable, say Fable
- 3:23Soul,
- 3:24uh Chimi K3, Grok, and GLM,
- 3:29we see that we have a 10% to 20%
- 3:33increase in levels of intelligence. Of
- 3:36course, Gemini 3.6 flash is still the
- 3:39fastest one out of all of the models,
- 3:41but if you're building applications,
- 3:43really you don't really care about
- 3:45having the fastest model. You care about
- 3:47the best quality at around 60 tokens per
- 3:51second.
- 3:52Now,
- 3:53something that also is quite interesting
- 3:55to see is the cost per task. So, the
- 3:57cost per task is what dictates how much
- 4:00it costs you to actually get stuff done
- 4:03using your AI.
- 4:04And
- 4:05this is where we can actually see that
- 4:07we have models such as Grok 4.5, which
- 4:11actually cost about 40% less
- 4:14and
- 4:16have a higher level of of intelligence.
- 4:18So, for reference, the level of
- 4:20intelligence of Grok 4.5 is very similar
- 4:23to that of Opus 4.8. So, we literally
- 4:26have a model that only performs better,
- 4:30but it's actually also cheaper.
- 4:33So, that begs the question,
- 4:34what do we really do with the Gemini
- 4:36models? Right now, Google has fallen
- 4:38severely behind in the AI race, despite
- 4:42being the company that actually has
- 4:43access to all of the resources that are
- 4:46necessary
- 4:47to keep going. So, the hardware, the
- 4:50software, and the data.
- 4:54We really hope that Gemini 3.5 Pro or
- 4:573.6 Pro are going to change this, but
- 5:01things are not looking very good for
- 5:03Google right now.
- 5:05Now, I'm going to say, I hope that you
- 5:06enjoyed this video.
- 5:08I'm not going to recommend the 3.6 model
- 5:11because I firmly believe that
- 5:14you have better options elsewhere via
- 5:16Grok. And if you want to just have your
- 5:19own open-source model, you can also use
- 5:21GLM. It's perfectly fine.
- 5:24Uh
- 5:25using flashlight is good, but even then,
- 5:29nothing has really changed.
- 5:32I hope that you enjoyed this video. I'm
- 5:34going to say, I'll catch you in the next
- 5:35one. Remember stay up to all of the
- 5:38latest AI news. In this case, it was
- 5:40just a bunch of noise.
- 5:42And I'll see you in the next one. See
- 5:44you.
About this transcript
This page contains the full transcript of Gemini 3.6 Flash: Don't Believe the Benchmarks by Miguel Torrez, generated from the public captions YouTube serves with the video. The transcript has 871 words across 146 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.