I Don't Need Frontier Models Anymore (Qwen 3.8 27B + DeepSeek Harness) — Transcript
Full transcript
- 0:01When 3.8 27 billion is out, okay, this
- 0:04is the model that I was waiting for.
- 0:07I've been using the previous version,
- 0:09the 3.6 27 billion, for quite a while
- 0:12and it's become my daily drive. I do 80%
- 0:15of the work with that model.
- 0:17I like it a lot. I bought a new computer
- 0:20just to run that model.
- 0:22I now run it on the 5090 and it's
- 0:24performing amazing. So, you can imagine
- 0:27how much I was waiting for this new
- 0:29version.
- 0:31And unfortunately, I have mixed feeling
- 0:34about it.
- 0:35Now,
- 0:36I want to explain because this is really
- 0:38important. There is something extremely
- 0:40important to understand, but also I have
- 0:42an insane good news about this model.
- 0:45This is what when released and this is
- 0:48the reason why a lot of people are
- 0:49comparing it to Opus 4.6. I personally
- 0:53I'm not much interested in this model.
- 0:55I'm more interested to understand what
- 0:57I'm gaining from what I had before to
- 0:59what I can run today.
- 1:01And I can see that in some scenario, the
- 1:04gap is not that big,
- 1:06but other scenarios like here, agent
- 1:08decoding, the gap is huge from 13.3 to
- 1:1242.2, it's massive. And software
- 1:15engineer for 49.3 to 79, this is another
- 1:20massive step forward. On frontier agent
- 1:23task,
- 1:25from 10.6 to 20, this is 100%
- 1:28improvement. The gap sometimes are
- 1:32pretty noticeable. So, why I felt this
- 1:35mixed feeling when I I tested myself?
- 1:38Because
- 1:39I tested the Qwen 3.6 27 billion
- 1:44Q6
- 1:45with the Qwen 3.8 27 billion Q5. In my
- 1:50test, bug hunting, the Q6 found one more
- 1:54bug than and Q5, even though the Q5 was
- 1:57the newer version. So, I stopped the
- 2:00test and I said, "Okay, this is
- 2:02This is disappointing for
- 2:04us, but let's compare them with the same
- 2:06quantization." So, I also installed the
- 2:08Qwen 3.8 27 billion Q6, then I could see
- 2:12the difference. But, the difference was
- 2:14still not that big as much as I was
- 2:17hoping for.
- 2:19So, there is probably it's hard to tell,
- 2:21but let's say a 30% improvement, which
- 2:24is noticeable, but not another level.
- 2:27So, I was expecting much more. This is
- 2:29because Luna itself could do such a
- 2:33great job when improving reasoning. So,
- 2:36here is GPT 5.6 Luna non-reasoning 27.
- 2:41When we start increase reasoning,
- 2:44go to 34, then
- 2:4639,
- 2:48then
- 2:4947, and then GPT 5.6 Luna max 52. The
- 2:55jump is insane. It's more than double
- 2:58the capability.
- 3:00So, let's look at
- 3:02Qwen 3.6 27 billion non-reasoning 31.
- 3:06When does reasoning go to 38? I was
- 3:09expecting, okay, the new model is going
- 3:11to be better as a model, but plus then
- 3:13going to increase the reasoning, so it
- 3:15should push us towards this kind of
- 3:18level. We are probably in the 40, but I
- 3:21was hoping to get actually up to in the
- 3:2350s. That is reality. Considering that
- 3:27when non-reasoning 31 and Luna
- 3:30non-reasoning is 27. So, potentially, we
- 3:34could have done much more than 50 and
- 3:37reach, you know, this kind of level.
- 3:40The DeepSeek version 4 Pro.
- 3:43That would be amazing. The GLM 5.2 max.
- 3:46Okay, that would be an amazing
- 3:48achievement, but we didn't. So, probably
- 3:51I was expecting too much.
- 3:54But let me tell you the good news here.
- 3:58The good news is this model is capable.
- 4:02But
- 4:04intelligence is not just about the
- 4:06model.
- 4:07Is the harness around it.
- 4:10Before I explain why
- 4:12I'm talking about harness now, it
- 4:13because I found that the killer combo
- 4:16between this model and a harness. Until
- 4:20now I was using Hermes and Open Code.
- 4:24And while Hermes I find it that is a
- 4:27fantastic system.
- 4:30Recently, when they start pushing their
- 4:33own subscription, I felt like
- 4:36yeah, that is not cool anymore. I felt
- 4:39like a client.
- 4:40I'm a guest within their system, and
- 4:43this is the reason why I'm building with
- 4:45my community Resonate OS, so we can
- 4:48co-own it and it's ours.
- 4:50The moment they started pushing their
- 4:53subscription, the moment that people are
- 4:54starting asking me is do I need to
- 4:57subscribe to use Hermes?
- 4:59I could see that
- 5:01Hermes just took the wrong path.
- 5:05Okay, and therefore I was already open
- 5:07with the idea I need to look to
- 5:09something different.
- 5:11With Resonate OS, I was thinking, okay,
- 5:13for us everything needs to be an add-on,
- 5:16okay? All the elements need to be
- 5:19possible for the community to add to the
- 5:21system, so they can customize, they can
- 5:22change it, they can replace part, they
- 5:24can replace the memory, they can replace
- 5:26the main agent, they can replace all
- 5:28those parts, but at least we have a
- 5:30system that we can work together and
- 5:32build it and customize together.
- 5:35So, at one point I came across randomly
- 5:38to this new harness that came out
- 5:41recently. I think less than a a week ago
- 5:44Deep Seek released Deep Seek harness.
- 5:48Now, I believe that this is the best
- 5:51combo for when 3.8 27 billion.
- 5:56I started using this harness. I'm going
- 5:58to explain a little bit, even though I
- 5:59will probably need to do a dedicated
- 6:02video to this harness because it does
- 6:04something that no other harnesses do.
- 6:07What happened? I started using it
- 6:08together with this new version of Quen
- 6:11and everything I wanted to build so far,
- 6:14it built it. So, in the past, I used to
- 6:17use Quen 3.6 27 billion, let's say 80%
- 6:21of the time. Sometimes, it would find a
- 6:24problem that was not solvable. I had to
- 6:26change model and use one of those
- 6:29frontier model. Now, with this Deep Seek
- 6:32harness and the new 3.8 27 billion, I
- 6:35never
- 6:37had a problem to swap to a different
- 6:39model.
- 6:40So, what is happening here? When I went
- 6:43to this website, I read everything is a
- 6:45plugin. So, in my world, everything is
- 6:48an add-on, for them, everything is a
- 6:50plugin. It resonated and said, "Hmm, I
- 6:53need to look into this thing."
- 6:55I was coming from
- 6:57not totally sure about Hermes anymore.
- 7:00We are still building with our community
- 7:02resident of us.
- 7:03I see this that says everything is a
- 7:06plugin.
- 7:07Total alignment. I installed it.
- 7:09Actually, I didn't even install, I asked
- 7:11Hermes to install it for me and
- 7:13configure it with the new model
- 7:16and it's amazing. So, now, this is
- 7:20how it looks like, okay?
- 7:22And I can't show you while it's running
- 7:25because while it's running, I can't
- 7:27record the screen because memory
- 7:29management of the file format that I'm
- 7:32recording. What I can tell you here is
- 7:36that is just amazing. It's not a
- 7:39finished product, okay? Okay, so they
- 7:41call it
- 7:42developer preview. So it's not finished.
- 7:45It's there is a lot of things maybe I'm
- 7:47missing things that are not perfect. So
- 7:49sometimes it doesn't do compaction in
- 7:52the right moment.
- 7:54And therefore it's a kind of a crash the
- 7:56system. Nothing serious because you just
- 7:58say resume working and it start working
- 8:01without any problems. But nevertheless
- 8:03it's still a preview. But what it did
- 8:06they change how the context window is
- 8:09managed.
- 8:10What is happening here? It's
- 8:12unbelievable.
- 8:13There is a lot of new elements inside.
- 8:17So it takes track of everything that has
- 8:20been done.
- 8:21And require for whatever system that
- 8:24they use less context window. I can't
- 8:27explain because it's confusing to me so
- 8:29I need to read and understand it better.
- 8:31But what I know it's this 131,000 tokens
- 8:35that I have available for this model on
- 8:37my 5090.
- 8:39They just last longer. The compaction is
- 8:43so much more solid. And this system is
- 8:46being chatting for so long building for
- 8:49so long that is incredible. So if you
- 8:52see I don't know if you can see but here
- 8:54it had an input of 38 million tokens.
- 8:59Still going strong. What did I build
- 9:02with this? I build this the augmentor.
- 9:06This is my agent. Okay, this is how I
- 9:07call my agent. This is part of resonant
- 9:10OS is also included in resonant OS. The
- 9:12ability augmentor that is powered by
- 9:15this deep seek harness.
- 9:17What it does it's it takes total control
- 9:20of the browser and and just go around
- 9:24and click and do things for me. So what
- 9:27I ask here was to check out of my ladies
- 9:31videos to see the comments and suggest a
- 9:35few videos idea.
- 9:37So that was a way to send
- 9:39these agents into the browser, click
- 9:42around, and grab information.
- 9:46So, here is what came out. It checked
- 9:49four five different videos.
- 9:52It saw some of the comments, and then
- 9:54come out with some ideas.
- 9:56Now, it doesn't matter the ideas because
- 9:58it's honestly I don't from what I see
- 10:01maybe are valuable, maybe not, but
- 10:03doesn't matter because it didn't have
- 10:05the instruction of what is this channel
- 10:07is about and so on. So, at the moment it
- 10:08doesn't have much context. It just went
- 10:11there and grabbed information. But, what
- 10:13is important is what was capable of
- 10:16doing. I built the all thing with this
- 10:19Qwen 3.8.
- 10:21It works completely fine. Yeah, there is
- 10:24a a little things that need to be
- 10:25improved.
- 10:26And the most important part, I never had
- 10:29the need to use a bigger model.
- 10:32Something that always happened before
- 10:34when was doing something as complex as
- 10:37this kind of tasks. So, we have another
- 10:40deep seek moment, but this time the deep
- 10:42seek moment is not on the model, is on
- 10:44the harness, and is done by deep seek.
- 10:47Now, I'm extremely curious to see what
- 10:49this harness can do even with a more
- 10:52powerful model. If you've been following
- 10:54this channel, you know that I'm really
- 10:56pro local AI and so on.
- 10:59I'm for the AI sovereignty, meaning you
- 11:02need to own the hardware and
- 11:04infrastructure so you can trust and
- 11:07control everything. This is the journey,
- 11:09but it's not first of all it's not for
- 11:11everybody, but also is is a journey,
- 11:13it's not an on off.
- 11:15What I can see with this deep seek
- 11:16harness that the journey is getting
- 11:18better and better because now we have
- 11:20this harness that has MIT license,
- 11:24meaning you can use it in your project.
- 11:27You can use it yourself, implement,
- 11:29change it, do whatever, and use for
- 11:31commercial. So, what I'm suggesting
- 11:33here,
- 11:34if you are watching this video, probably
- 11:36is because you have access to this
- 11:38model.
- 11:40But, even if you don't, try this uh
- 11:42DeepSqueak Harness.
- 11:44If you're one of those lucky people that
- 11:46own a
- 11:48two DJX Spark
- 11:50or equivalent,
- 11:51try DeepSqueak version full flash,
- 11:54because can run on two of them with this
- 11:57harness, that probably they could work
- 12:00amazingly together. I'm at the moment of
- 12:02after a few months of uh
- 12:05of struggling with this AI,
- 12:07I'm excited again. I think it's a
- 12:10similar moment of when Open Claw came
- 12:13out, which was called uh
- 12:15Clawbot, I think. Can't even remember
- 12:17anymore.
- 12:19So, when it came out, it just was a
- 12:23massive push forward, because everybody
- 12:25finally had this agent that can do a
- 12:27lot. They control your computer, work
- 12:29for you, and so on. So, now I think what
- 12:33DeepSqueak did with their harness is a
- 12:36similar jump.
- 12:38Allows this model to do so much more,
- 12:42which is incredible.
- 12:43I will need to test this for longer, of
- 12:45course. I can't say,
- 12:47you know, that is the final solution,
- 12:50but at moment, this is what I feel. I
- 12:52feel like I never needed a more powerful
- 12:56model than this when I use this harness.
- 12:59Something that I I can't say the same
- 13:01with uh Hermes or Open Code.
- 13:05I think it's a killer combo. I need to
- 13:07understand this DeepSqueak harness even
- 13:09more. At moment, it is not fully
- 13:11finished. The way it works it's uh
- 13:14surprisingly well. So, if I were you,
- 13:17first of all, go to this DeepSqueak.com
- 13:21harness and uh and read it through. Like
- 13:24I said, I will do a new video, but they
- 13:26are definitely managing the things
- 13:28differently. And now, what is possible
- 13:30to achieve with this harness and the
- 13:32Qwen 3.8 27 billion is impressive.
- 13:37This is going to be my new way of
- 13:40working from now on. I will need to
- 13:43migrate and learn how to do it from
- 13:45Hermes to this model.
- 13:48Again, the good thing that is happening
- 13:50that is I basically already built this
- 13:53agent. This agent they are going to
- 13:55mentor.
- 13:56This can be a an add-on for Resonant OS
- 13:59and again shows the importance of
- 14:01working in add-ons or plugins like they
- 14:05call it because now it's easy to
- 14:07implement new things.
- 14:10And what I can even believe is that I
- 14:14built this
- 14:15100% with my local AI.
- 14:19Not needing a more powerful system.
- 14:23Everything that I asked to do has been
- 14:25done. Not getting lost and so on. One
- 14:29thing that I want to say about the
- 14:30psychic harness connected together with
- 14:33this Qwen 3.8 27 billion
- 14:36is they use a lot of token.
- 14:39They use a lot of tokens.
- 14:41But,
- 14:44the result is there. We know that the
- 14:46direction was increase the thinking time
- 14:48and you you will perform better. Here is
- 14:51happening and is incredible. Let's look
- 14:54at the what kind of hardware you need to
- 14:56run this model. In theory, according to
- 14:58Qwen, you can run it on 12 GB of RAM.
- 15:03Ideally, a 48 GB of RAM. And I would
- 15:06agree on the 48 GB of RAM because what
- 15:09you gain is actually a larger context
- 15:11window.
- 15:12I have the 32 GB of RAM. Okay, a Q6
- 15:16and 25 GB and look, I only lose a 0.2%
- 15:23from the Q8. So, the compromise is
- 15:26pretty acceptable. With the Q5, you lose
- 15:29a little bit. It's you lose 1.4.
- 15:32At 1.4, you say, "Okay, that is not
- 15:34much." And I mean, from this one, but
- 15:37you lose a 1.6
- 15:40from the top. It's still not that much,
- 15:42but from my benchmark that I run
- 15:45earlier,
- 15:47the bug hunting, the Q6 could find one
- 15:51more bug than the Q5.
- 15:55So, it is a little bit, you know,
- 15:57annoying, this thing. So, here's my
- 16:00suggestion.
- 16:01If you're looking for a bigger contest
- 16:03window, Q5 is fine. If you need to do
- 16:06some specific work that is more
- 16:08advanced,
- 16:10of course, you need to move into the Q6,
- 16:13ideally.
- 16:14If you have the 32 GB of RAM. I'm
- 16:17explaining this because it's important
- 16:20to understand what are the limits. When
- 16:22you go into the 16 GB of RAM, you start
- 16:25to lose quite a bit. At 12 GB of RAM,
- 16:29you lose a lot.
- 16:32So, while I never tried these two
- 16:35models, okay, I never tried them, I
- 16:38can't say how much you're losing from
- 16:40the model. This is something that you
- 16:41need to test yourself. But knowing that
- 16:43the minimum minimum requirement is 12,
- 16:46ideally, you want to be in the 24, 32,
- 16:5048 if you can do it. So, now, another
- 16:53element you need to consider is the size
- 16:55of the memory that is required to run
- 16:57this model.
- 16:58But the speed at which the model runs is
- 17:02connected to the bandwidth.
- 17:04So, I just listed here, I used these
- 17:07perplexities to do this research on the
- 17:1032 GB of RAM and 24 GB of RAM. So, here,
- 17:14for the 32 GB of RAM, the best one is
- 17:17the 1590, okay? This has a bandwidth
- 17:20that is insane, 1,792
- 17:23GB per second. That means the this model
- 17:27is working at around 100 and 10 token
- 17:30per second on my hardware. It's doing
- 17:32insanely well.
- 17:34So,
- 17:35when you go on this one,
- 17:38you can see that the AMD is doing 640.
- 17:43This is almost a third of the speed. So,
- 17:45you have an idea of what you can expect.
- 17:49Now, on the 24 GB of VRAM, we have a
- 17:53different kind of opportunity.
- 17:551,000 GB per second, which is pretty
- 17:58good. 3090 is the same. The good thing
- 18:01with 3090 that you can find them second
- 18:03hand much cheaper. This is the 3090
- 18:05idle. So, the 3090 normal 3090 is here,
- 18:09which is
- 18:11basically
- 18:12not not a big of a difference, okay? It
- 18:13doesn't really matter.
- 18:15But, from AMD, we have this card,
- 18:19which is still in commerce, that you can
- 18:21buy and it's not as expensive as those
- 18:254090. And the speed is almost there.
- 18:28Now, those are 24 GB of RAM. So, if I
- 18:31were in you and if you can afford, I
- 18:34mean, two of these together, they give
- 18:36you 48 GB of RAM, which again give us
- 18:41this 48 GB of RAM. I would not run the
- 18:44Q8. I would go for the Q6, but with a
- 18:46much bigger context window. Keep in mind
- 18:49that on my 32 GB of RAM, I can have
- 18:52131,000
- 18:54token context window. With 48, you
- 18:57probably, I guess, you could have the
- 18:59the total 260 something context window,
- 19:03which is a nice plus.
- 19:05Remember, Q8 is going to run slower than
- 19:08the the Q6, the Q6 slower than the Q5,
- 19:11and so on.
- 19:12Bandwidth, it is something to consider.
- 19:14At the moment, if you want to buy new,
- 19:16this is really good. If you're looking
- 19:18for second hand, this is pretty good.
- 19:21One more thing, this model can run on a
- 19:24DGX Spark and it runs, depending on the
- 19:26quantization, at around 20 30 token per
- 19:29second. At 30 token per second, in
- 19:32reality, you have a system that works,
- 19:34is usable, and I will probably run it
- 19:36and test it on my GX 10.
- 19:39This is the ASUS version of the DGX
- 19:41Spark. It's 128 GB of RAM, so I can run
- 19:44it with really large context window with
- 19:47multiple instances. So, that would be
- 19:49the advantage.
- 19:51Anyway, that's is what I wanted to share
- 19:53for today.
- 19:54To recap,
- 19:56a model that is not as good as I wanted.
- 19:58It didn't do the jump that I was
- 20:00expecting and wanting. Probably we're
- 20:02going to see it when they release the
- 20:03version four of it. But, combining this
- 20:07with Deep Seek harness, I saw something
- 20:11that I never saw before. It's early days
- 20:13to say, but so far, in the last 24
- 20:15hours, I use it a lot.
- 20:17And I never felt I needed to use a
- 20:22bigger model. Never once.
- 20:24[clears throat]
- 20:25That make me think that things are
- 20:27changing and locally I is gaining more
- 20:31and more value
- 20:33every single day.
- 20:34If you're interested in this kind of
- 20:35content and you want to join our
- 20:37community, first like and subscribe. In
- 20:39the description, you'll find the link to
- 20:41our Discord channel and there is a deep
- 20:44dive article. I will go deeper in
- 20:47everything that I'm talking inside this
- 20:49video. I will add all the links that I
- 20:52discussed over here and if possible, I
- 20:54will see how it goes, but I might also
- 20:57share
- 20:58this augmentor power by this DSH. This
- 21:02is an extension for the browser and by
- 21:06adding it, you can have this sidecar.
- 21:10Which, honestly, is amazing. It is
- 21:13actually amazing.
- 21:15I say if I can, because what is missing
- 21:18before I can share it is a way for you
- 21:21to add your own model. At the moment,
- 21:24there's no
- 21:25configuration system where you can add
- 21:28your API or local model or something.
- 21:31Even though through the code, you could
- 21:33be able to ask to your agents to add the
- 21:36model, I want to make it through the
- 21:38interface so everybody can just add
- 21:40their, you know, details and run it
- 21:43locally or for any other cloud provider.
- 21:47So, if I can manage to do this in time,
- 21:50I will link inside the Substack article
- 21:53the code for this
- 21:55agent. If not, will be shared in the in
- 22:00the near future. Thanks for staying
- 22:01until the end. Again, like and
- 22:03subscribe. Watch the next video. Join
- 22:05the Discord. And ciao.
About this transcript
This page contains the full transcript of I Don't Need Frontier Models Anymore (Qwen 3.8 27B + DeepSeek Harness) by Manolo Remiddi, generated from the public captions YouTube serves with the video. The transcript has 3,334 words across 513 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.