Is Qwen 3.8 27b The New King of the Local LLMs? — Transcript
Full transcript
- 0:00So Quen 3.8 27 billion parameter local
- 0:03AI model has been out for a few days now
- 0:06and I've been testing it to see what it
- 0:08is capable of and I've been surprised on
- 0:10how well it is performing especially
- 0:12compared to the models that we had what
- 0:14just a few weeks ago. This is another
- 0:17major step change in what you can
- 0:19achieve locally on local hardware
- 0:22without having to send everything up
- 0:24into the cloud. That's the topic of
- 0:26today's video. So if you want to find
- 0:27out more, please let me explain.
- 0:32So Quen 3.8 is a 27 billion parameter
- 0:35model. Now it's full size at 16 bit
- 0:38floating point numbers. It's a 55 GB
- 0:41download and you're probably going to
- 0:43need 128 GB of RAM to run it, which is
- 0:46beyond what most people have. That's
- 0:48VRAM. Beyond what most people have kind
- 0:51of at home, even with a very good gaming
- 0:53setup. However, if you do have something
- 0:56like an RTX 5090 or something else with
- 0:5932 GB of VRAM, then the [snorts] 4bit
- 1:03quantized version of Quen 3.827 billion
- 1:06is runnable on your local PC and
- 1:10actually at a fairly decent rate. I was
- 1:12getting 96 tokens a second uh running it
- 1:15on an RTX5090. That means you can run it
- 1:17locally on your PC and do everything
- 1:20you'd want to do with a normal large
- 1:22language model out in the cloud, but now
- 1:24you're running it locally. And I've been
- 1:26testing it specifically for how good is
- 1:28it at you doing coding using open code
- 1:31as the coding harness. But first, let's
- 1:34start with a few simple things. Does it
- 1:36answer the Alice question? Yes, it does.
- 1:39No problem, as you would expect. Does it
- 1:41handle the 7 minute egg timer and the 11
- 1:45minute egg timer question? Yes, it does.
- 1:48No problem. So, now those are out the
- 1:50way, we can kind of get a feel for the
- 1:52kind of capabilities of this model.
- 1:54Let's push it a little further. The
- 1:56first thing I did cuz I wanted a bit of
- 1:58fun was to get it to create a Space
- 2:00Invaders game using a single HTML file
- 2:03using inline JavaScript and no external
- 2:06libraries. So this is just had to create
- 2:08everything itself inside of the HTML
- 2:10file. It did that. It produced me
- 2:13index.html. You can doubleclick it and
- 2:16you get a nice space invaders game. So I
- 2:18wanted to push it a bit more. So I use
- 2:19my protocol encoding type length value
- 2:22encoding question. Asked it to create a
- 2:25specification similar to ASN1 similar to
- 2:29Google's Protobuff. Now it created that
- 2:31specification. No trouble. There were a
- 2:33few tiny mistakes in it. It did all the
- 2:35specification defining how the protocol
- 2:37should be laid out. But then at the very
- 2:39end, it gives some working examples.
- 2:40This is what happens if you want to
- 2:41encode an integer. This is what happens
- 2:43if you want to encode a string. And
- 2:45actually one of those or two of those
- 2:46had a minor error in the numbers it was
- 2:49telling you that you should use. So the
- 2:51examples it was giving uh didn't
- 2:53actually quite work out. But if you fix
- 2:54those, no problem whatsoever. So I then
- 2:56went on to say, well now implemented
- 2:58including testing, including end toend
- 3:00testing, unit testing, and so on. And I
- 3:02did open code, let it run. A few times I
- 3:05had to interact with it, telling it, you
- 3:07know, here's a plan. This is what you
- 3:09should do now. Go ahead and do it. And
- 3:12just as you would do with normal coding
- 3:14with a harness and it worked perfectly.
- 3:16It gave me a nice little library that
- 3:19could implement the specification,
- 3:20including some end to-end testing,
- 3:21including some unit tests. Worked
- 3:24absolutely fine. So my big test after
- 3:26that as I've done in several other
- 3:27videos here on this channel is my new
- 3:30scrippy uh programming language. So the
- 3:32idea is I just give it a example of the
- 3:35of the new scripping language and even
- 3:37define it formally. I tell it a little
- 3:39bit about it like it doesn't use
- 3:41semicolons. I tell it it's an
- 3:43interpreter and I tell it it's written
- 3:45in C and then I say go ahead knock
- 3:47yourself out go and write the
- 3:49interpreter for this language. Now
- 3:51that's quite a tall order because
- 3:52there's no formal specification for the
- 3:54language just an example code and a lot
- 3:57of the stuff it has to kind of infer
- 3:59from the example code infer from what
- 4:01I've said uh in the bit of text. Now
- 4:04I've tried this before in other videos
- 4:05and really local models weren't able to
- 4:08do it even when uh 3.6 they weren't able
- 4:11to do it. They kind of either got stuck
- 4:13in a loop just not being able to finish
- 4:15it or what they produced crashed all the
- 4:17time had severe memory leaks. It was bit
- 4:20of a disaster. However, the amazing news
- 4:22is here is that in a relatively short
- 4:25amount of time certainly wasn't even
- 4:27overnight just a few hours again few
- 4:30prompting there was a plan and then we
- 4:33went through the different steps on how
- 4:34to implement it and then you know I said
- 4:36also write some documentation and I told
- 4:38it all the stuff back and forth as you
- 4:40would with a coding harness. It worked
- 4:43absolutely worked. You can run new
- 4:45scrippy. You It created its own set of
- 4:47test scripts to run as well. Uh and the
- 4:51example script that I gave it for the
- 4:53readme file that runs. And so there you
- 4:56go. This local LLM running on a PC.
- 5:00Quinn 3.8 was able to do all of that
- 5:02coding using open code. You could use
- 5:04other harnesses as well, I'm sure.
- 5:06Haven't tested them specifically, but
- 5:07I'm sure it'll be just as good. And it
- 5:09did it all for me. And it actually
- 5:11works. So, this is really amazing that
- 5:12we've reached this point. Uh, and I'm
- 5:15sure things are going to get better, but
- 5:17this is amazing that we've reached this
- 5:18point for local LLMs, just the power and
- 5:21the capabilities that they have. Okay,
- 5:23I'd love to hear your thoughts in the
- 5:24comments. Have you tried Quen 3.8? Do
- 5:27you use uh local uh AI for coding or do
- 5:31you only rely on the big models? Do you
- 5:33not do any coding on it? Love to hear
- 5:35your thoughts in the comments below.
- 5:37Okay, that's it. I'll see you in the
- 5:39next one.
- 5:43>> [music]
About this transcript
This page contains the full transcript of Is Qwen 3.8 27b The New King of the Local LLMs? by Gary Explains, generated from the public captions YouTube serves with the video. The transcript has 1,103 words across 149 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.