YouTube2Text

Gemma 4 on RTX 3090 vs 4090 vs 5090 vs Mac: Benchmarks That Will Surprise You (31B and 26B-A4B) — Transcript

by Tech-Practice · 753 words · 123 segments · language en · Watch on YouTube

Full transcript

  1. 0:05Hi everyone, welcome to the channel. In
  2. 0:07this video, we will do a comparison of
  3. 0:09GPUs running the latest model from
  4. 0:12Google, the Gemma 4 model.
  5. 0:15It's a claimed to be one of the best
  6. 0:17open-source model, so we'll see how that
  7. 0:20runs.
  8. 0:21Gemma 4 comes in four size. It
  9. 0:24comes with 2.3 billion, 4.5 billion, 31
  10. 0:30billion, and the 26 billion. The 26
  11. 0:33billion is a mixture of experts model,
  12. 0:36so which makes it running very fast.
  13. 0:39The GPUs we'll be testing including the
  14. 0:413090, 4090, and the 5090. They all have
  15. 0:45at least 24 GB of VRAM, so we will be
  16. 0:49focusing on the two largest model, the
  17. 0:5331B model and the 26B model.
  18. 0:56For the inference engine, we'll be using
  19. 0:59the llama.cpp.
  20. 1:00The operating systems are Linux. We will
  21. 1:02build the llama.cpp from scratch, so
  22. 1:05everything will have the same build. Now
  23. 1:08we put the three GPU side by side, and
  24. 1:11let's start the llama.cpp command line.
  25. 1:15We'll be using the 31B model.
  26. 1:21So now So now let's give it a prompt.
  27. 1:26We'll be using the same prompt on 3090,
  28. 1:294090, and the 5090.
  29. 1:37And let's press enter to start it.
  30. 1:41So we now can see that the direct
  31. 1:44comparison of the speed. We do see that
  32. 1:47the 5090 is just a outlier. It performs
  33. 1:51much faster than both 4090 and the 3090.
  34. 2:075090 is able to complete first.
  35. 2:11The generating speed is at 64.8
  36. 2:14tokens per second.
  37. 2:15And the 4090 comes second at 42.3 tokens
  38. 2:21per second.
  39. 2:22And the 3090 is the last one at 35.7
  40. 2:27tokens per second.
  41. 2:30So we can now do a direct comparison in
  42. 2:33terms of the generating speed.
  43. 2:36Yeah, let's uh
  44. 2:38see more details about the generation
  45. 2:40here. We see that the 3090 and the 4090
  46. 2:43is much closer, but the 5090 is much
  47. 2:46faster than both 3090 and the 4090.
  48. 2:50Now let's switch gear to the smaller
  49. 2:53model, the 26 billion A4B model.
  50. 2:57So this is a mixture of experts model.
  51. 3:00So this is a mixture of experts model.
  52. 3:02Each time, only 4 billion parameters are
  53. 3:06activated.
  54. 3:07Similarly, let's
  55. 3:09do the same prompt for all three of
  56. 3:12them.
  57. 3:15Press enter.
  58. 3:21I think compared to the previously 31B
  59. 3:25dense model, this is much faster.
  60. 3:28The 5090 comes at 182 tokens per second.
  61. 3:33Wow, that's a lot. And the 4090 is 147.
  62. 3:38And the 3090 is 120
  63. 3:41tokens per second. All three of them are
  64. 3:44above 100 tokens per second.
  65. 3:50Yeah, because it's generating so fast,
  66. 3:52let's do another one.
  67. 4:05The generate The generating speeds are
  68. 4:07pretty similar to previous run.
  69. 4:11Lastly, let's do a benchmarking for
  70. 4:14running this model on 3090, 4090, and
  71. 4:17the 5090. So this comes with the
  72. 4:19llama.cpp. There is llama-bench
  73. 4:24command. You can use it to test
  74. 4:26different lengths of tokens, for
  75. 4:28example, for the prompt evaluation and
  76. 4:30also for the token generating.
  77. 4:36I'm also opening a GPU monitoring for
  78. 4:38the RTX 5090. It has 32 GB of VRAM. It
  79. 4:42used around the 17
  80. 4:45GB of the VRAM. There's still lots of
  81. 4:48room left. Utilization for the GPU is at
  82. 4:5190%.
  83. 4:54Here Here comes the final benchmarking
  84. 4:57results.
  85. 4:58We can do some direct comparison. We see
  86. 5:01something very interesting. So for the
  87. 5:043090, the prompt evaluating is the
  88. 5:07slowest. It's around 4,000, while the
  89. 5:104090 is at
  90. 5:138,600.
  91. 5:15It's a much faster. And the 5090 is at
  92. 5:18uh
  93. 5:19almost 10,000.
  94. 5:23For the prompt
  95. 5:27For the token generating
  96. 5:29the 3090 is at around 129
  97. 5:33tokens per second. While 4090 is at 160
  98. 5:39tokens per second.
  99. 5:40The 5090 here is the about 190 tokens
  100. 5:44per second. So I would say that because
  101. 5:47the model is a mixture of experts model,
  102. 5:50so it doesn't depends a lot on the VRAM.
  103. 5:54All three of them have enough VRAM, so
  104. 5:56the generating speeds are much closer
  105. 5:58than the previous the dense model.
  106. 6:02Overall, I think all three card are
  107. 6:04great running the Gemma 4 model. I hope
  108. 6:07this video is useful for you. Please
  109. 6:10give it a trial and a post your
  110. 6:12generating numbers in the comments.
  111. 6:14Additionally, if you are interested in
  112. 6:17using MacBook to running the models, so
  113. 6:19here are some numbers I tested using my
  114. 6:23MacBook, the M3 Max 36 GB unified RAM.
  115. 6:28If you want to more details, please
  116. 6:30check out my previous video about using
  117. 6:33MacBook to run all four Gemma 4 models.
  118. 6:36Thank you for watching.
  119. 6:39Please give it a sum up and share it.
  120. 6:42Please
  121. 6:43subscribe to the channel for future
  122. 6:45content. Thank you for your support.
  123. 6:48Goodbye.

About this transcript

This page contains the full transcript of Gemma 4 on RTX 3090 vs 4090 vs 5090 vs Mac: Benchmarks That Will Surprise You (31B and 26B-A4B) by Tech-Practice, generated from the public captions YouTube serves with the video. The transcript has 753 words across 123 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.