Gemma 4 on RTX 3090 vs 4090 vs 5090 vs Mac: Benchmarks That Will Surprise You (31B and 26B-A4B) — Transcript
Full transcript
- 0:05Hi everyone, welcome to the channel. In
- 0:07this video, we will do a comparison of
- 0:09GPUs running the latest model from
- 0:12Google, the Gemma 4 model.
- 0:15It's a claimed to be one of the best
- 0:17open-source model, so we'll see how that
- 0:20runs.
- 0:21Gemma 4 comes in four size. It
- 0:24comes with 2.3 billion, 4.5 billion, 31
- 0:30billion, and the 26 billion. The 26
- 0:33billion is a mixture of experts model,
- 0:36so which makes it running very fast.
- 0:39The GPUs we'll be testing including the
- 0:413090, 4090, and the 5090. They all have
- 0:45at least 24 GB of VRAM, so we will be
- 0:49focusing on the two largest model, the
- 0:5331B model and the 26B model.
- 0:56For the inference engine, we'll be using
- 0:59the llama.cpp.
- 1:00The operating systems are Linux. We will
- 1:02build the llama.cpp from scratch, so
- 1:05everything will have the same build. Now
- 1:08we put the three GPU side by side, and
- 1:11let's start the llama.cpp command line.
- 1:15We'll be using the 31B model.
- 1:21So now So now let's give it a prompt.
- 1:26We'll be using the same prompt on 3090,
- 1:294090, and the 5090.
- 1:37And let's press enter to start it.
- 1:41So we now can see that the direct
- 1:44comparison of the speed. We do see that
- 1:47the 5090 is just a outlier. It performs
- 1:51much faster than both 4090 and the 3090.
- 2:075090 is able to complete first.
- 2:11The generating speed is at 64.8
- 2:14tokens per second.
- 2:15And the 4090 comes second at 42.3 tokens
- 2:21per second.
- 2:22And the 3090 is the last one at 35.7
- 2:27tokens per second.
- 2:30So we can now do a direct comparison in
- 2:33terms of the generating speed.
- 2:36Yeah, let's uh
- 2:38see more details about the generation
- 2:40here. We see that the 3090 and the 4090
- 2:43is much closer, but the 5090 is much
- 2:46faster than both 3090 and the 4090.
- 2:50Now let's switch gear to the smaller
- 2:53model, the 26 billion A4B model.
- 2:57So this is a mixture of experts model.
- 3:00So this is a mixture of experts model.
- 3:02Each time, only 4 billion parameters are
- 3:06activated.
- 3:07Similarly, let's
- 3:09do the same prompt for all three of
- 3:12them.
- 3:15Press enter.
- 3:21I think compared to the previously 31B
- 3:25dense model, this is much faster.
- 3:28The 5090 comes at 182 tokens per second.
- 3:33Wow, that's a lot. And the 4090 is 147.
- 3:38And the 3090 is 120
- 3:41tokens per second. All three of them are
- 3:44above 100 tokens per second.
- 3:50Yeah, because it's generating so fast,
- 3:52let's do another one.
- 4:05The generate The generating speeds are
- 4:07pretty similar to previous run.
- 4:11Lastly, let's do a benchmarking for
- 4:14running this model on 3090, 4090, and
- 4:17the 5090. So this comes with the
- 4:19llama.cpp. There is llama-bench
- 4:24command. You can use it to test
- 4:26different lengths of tokens, for
- 4:28example, for the prompt evaluation and
- 4:30also for the token generating.
- 4:36I'm also opening a GPU monitoring for
- 4:38the RTX 5090. It has 32 GB of VRAM. It
- 4:42used around the 17
- 4:45GB of the VRAM. There's still lots of
- 4:48room left. Utilization for the GPU is at
- 4:5190%.
- 4:54Here Here comes the final benchmarking
- 4:57results.
- 4:58We can do some direct comparison. We see
- 5:01something very interesting. So for the
- 5:043090, the prompt evaluating is the
- 5:07slowest. It's around 4,000, while the
- 5:104090 is at
- 5:138,600.
- 5:15It's a much faster. And the 5090 is at
- 5:18uh
- 5:19almost 10,000.
- 5:23For the prompt
- 5:27For the token generating
- 5:29the 3090 is at around 129
- 5:33tokens per second. While 4090 is at 160
- 5:39tokens per second.
- 5:40The 5090 here is the about 190 tokens
- 5:44per second. So I would say that because
- 5:47the model is a mixture of experts model,
- 5:50so it doesn't depends a lot on the VRAM.
- 5:54All three of them have enough VRAM, so
- 5:56the generating speeds are much closer
- 5:58than the previous the dense model.
- 6:02Overall, I think all three card are
- 6:04great running the Gemma 4 model. I hope
- 6:07this video is useful for you. Please
- 6:10give it a trial and a post your
- 6:12generating numbers in the comments.
- 6:14Additionally, if you are interested in
- 6:17using MacBook to running the models, so
- 6:19here are some numbers I tested using my
- 6:23MacBook, the M3 Max 36 GB unified RAM.
- 6:28If you want to more details, please
- 6:30check out my previous video about using
- 6:33MacBook to run all four Gemma 4 models.
- 6:36Thank you for watching.
- 6:39Please give it a sum up and share it.
- 6:42Please
- 6:43subscribe to the channel for future
- 6:45content. Thank you for your support.
- 6:48Goodbye.
About this transcript
This page contains the full transcript of Gemma 4 on RTX 3090 vs 4090 vs 5090 vs Mac: Benchmarks That Will Surprise You (31B and 26B-A4B) by Tech-Practice, generated from the public captions YouTube serves with the video. The transcript has 753 words across 123 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.