YouTube2Text

Gemini 3.6 Flash: Don't Believe the Benchmarks — Transcript

by Miguel Torrez · 871 words · 146 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hey, what's up, everyone? My name's
  2. 0:02Miguel. Now, Google just announced the
  3. 0:04dropping of three new Gemini models.
  4. 0:07However, it seems that they're in
  5. 0:08trouble.
  6. 0:09Let me break it down for you.
  7. 0:10We've got the newest models released
  8. 0:12just today, Gemini 3.6 Flash, 3.5 Flash
  9. 0:16Light, and 3.5 Flash Cyber. Now, this
  10. 0:19last one is not really a model per se,
  11. 0:21it's just a version of Gemini Flash more
  12. 0:25oriented towards cybersecurity.
  13. 0:27Of course, if we look at the benchmarks,
  14. 0:29they're going to tell you that they're
  15. 0:30performing better and faster than the
  16. 0:33the last generations. That's something
  17. 0:34that you should expect from every model,
  18. 0:36but the true question is, are these
  19. 0:38models really keeping up with the level
  20. 0:41of performance that we have now, as well
  21. 0:43as the costs?
  22. 0:44And that's where it actually gets
  23. 0:46tricky. So, for information, Gemini has
  24. 0:48three different versions of the models.
  25. 0:51You have the smallest one, which is
  26. 0:52called Flash Light. This model is
  27. 0:55probably my favorite model out of all of
  28. 0:57them because it's extremely affordable,
  29. 0:59and you can use it for a lot of analysis
  30. 1:01at scale.
  31. 1:03You see, the one superpower that's left
  32. 1:05for all of the Gemini models is that
  33. 1:07it's still the only AI that can really
  34. 1:10analyze videos, documents, and anything
  35. 1:13that you throw at it
  36. 1:15at a native level. So, it really sees
  37. 1:18and has
  38. 1:19incredibly powerful OCR capacities. That
  39. 1:23means that it can actually visualize the
  40. 1:24things that you that you have in a
  41. 1:26document perfectly.
  42. 1:28Now, Pro, the biggest one out of the
  43. 1:30three
  44. 1:31was a contestant when we still had Opus
  45. 1:344.5, I believe.
  46. 1:37Gemini 3.1 Pro was an excellent model,
  47. 1:41but we haven't had a release of 3.1 3.5
  48. 1:45Pro nor 3.6 Pro in a while now. We don't
  49. 1:49even know where they are, really.
  50. 1:51And that makes the question,
  51. 1:54well, how are these two performing
  52. 1:56compared to the rest of the market? And
  53. 1:59well, the news are not really that good.
  54. 2:02So, as you can see right here on this
  55. 2:03chart, we have the level of performance
  56. 2:06of 3.5
  57. 2:09flash and 3.6
  58. 2:11flash
  59. 2:12at exactly the same level.
  60. 2:16So, that means that these model,
  61. 2:17according to artificial analysis, the
  62. 2:20source where this graph comes from,
  63. 2:23doesn't really have a big improvement in
  64. 2:25terms of performance.
  65. 2:26Now, not only that, but of course,
  66. 2:28Gemini 3.5 flashlight, which is the
  67. 2:31lower end of the models, is still
  68. 2:34sitting way, way behind in the charts.
  69. 2:36That's normal. It's a slower model. It's
  70. 2:39a smaller model. So, it actually makes
  71. 2:42sense for it to be there.
  72. 2:43But, where things get actually worrisome
  73. 2:46is whenever you actually look in detail
  74. 2:49at what's happening in the back. So, as
  75. 2:51you can see right here, we have the
  76. 2:53benchmarks of the state-of-the-art
  77. 2:55models, Chimi K3, GLM,
  78. 2:58Fable, Soul, etc. And this is where we
  79. 3:02see
  80. 3:03where Gemini actually stands. So, on the
  81. 3:06intelligence level, we see that 3.6
  82. 3:09flash hits a 50,
  83. 3:11score of 50.
  84. 3:13However, whenever you actually reach
  85. 3:16the higher end
  86. 3:18of the models, so we're going to call
  87. 3:20this the top five, Fable, say Fable
  88. 3:23Soul,
  89. 3:24uh Chimi K3, Grok, and GLM,
  90. 3:29we see that we have a 10% to 20%
  91. 3:33increase in levels of intelligence. Of
  92. 3:36course, Gemini 3.6 flash is still the
  93. 3:39fastest one out of all of the models,
  94. 3:41but if you're building applications,
  95. 3:43really you don't really care about
  96. 3:45having the fastest model. You care about
  97. 3:47the best quality at around 60 tokens per
  98. 3:51second.
  99. 3:52Now,
  100. 3:53something that also is quite interesting
  101. 3:55to see is the cost per task. So, the
  102. 3:57cost per task is what dictates how much
  103. 4:00it costs you to actually get stuff done
  104. 4:03using your AI.
  105. 4:04And
  106. 4:05this is where we can actually see that
  107. 4:07we have models such as Grok 4.5, which
  108. 4:11actually cost about 40% less
  109. 4:14and
  110. 4:16have a higher level of of intelligence.
  111. 4:18So, for reference, the level of
  112. 4:20intelligence of Grok 4.5 is very similar
  113. 4:23to that of Opus 4.8. So, we literally
  114. 4:26have a model that only performs better,
  115. 4:30but it's actually also cheaper.
  116. 4:33So, that begs the question,
  117. 4:34what do we really do with the Gemini
  118. 4:36models? Right now, Google has fallen
  119. 4:38severely behind in the AI race, despite
  120. 4:42being the company that actually has
  121. 4:43access to all of the resources that are
  122. 4:46necessary
  123. 4:47to keep going. So, the hardware, the
  124. 4:50software, and the data.
  125. 4:54We really hope that Gemini 3.5 Pro or
  126. 4:573.6 Pro are going to change this, but
  127. 5:01things are not looking very good for
  128. 5:03Google right now.
  129. 5:05Now, I'm going to say, I hope that you
  130. 5:06enjoyed this video.
  131. 5:08I'm not going to recommend the 3.6 model
  132. 5:11because I firmly believe that
  133. 5:14you have better options elsewhere via
  134. 5:16Grok. And if you want to just have your
  135. 5:19own open-source model, you can also use
  136. 5:21GLM. It's perfectly fine.
  137. 5:24Uh
  138. 5:25using flashlight is good, but even then,
  139. 5:29nothing has really changed.
  140. 5:32I hope that you enjoyed this video. I'm
  141. 5:34going to say, I'll catch you in the next
  142. 5:35one. Remember stay up to all of the
  143. 5:38latest AI news. In this case, it was
  144. 5:40just a bunch of noise.
  145. 5:42And I'll see you in the next one. See
  146. 5:44you.

About this transcript

This page contains the full transcript of Gemini 3.6 Flash: Don't Believe the Benchmarks by Miguel Torrez, generated from the public captions YouTube serves with the video. The transcript has 871 words across 146 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.