YouTube2Text

I Don't Need Frontier Models Anymore (Qwen 3.8 27B + DeepSeek Harness) — Transcript

by Manolo Remiddi · 3,334 words · 513 segments · language en · Watch on YouTube

Full transcript

  1. 0:01When 3.8 27 billion is out, okay, this
  2. 0:04is the model that I was waiting for.
  3. 0:07I've been using the previous version,
  4. 0:09the 3.6 27 billion, for quite a while
  5. 0:12and it's become my daily drive. I do 80%
  6. 0:15of the work with that model.
  7. 0:17I like it a lot. I bought a new computer
  8. 0:20just to run that model.
  9. 0:22I now run it on the 5090 and it's
  10. 0:24performing amazing. So, you can imagine
  11. 0:27how much I was waiting for this new
  12. 0:29version.
  13. 0:31And unfortunately, I have mixed feeling
  14. 0:34about it.
  15. 0:35Now,
  16. 0:36I want to explain because this is really
  17. 0:38important. There is something extremely
  18. 0:40important to understand, but also I have
  19. 0:42an insane good news about this model.
  20. 0:45This is what when released and this is
  21. 0:48the reason why a lot of people are
  22. 0:49comparing it to Opus 4.6. I personally
  23. 0:53I'm not much interested in this model.
  24. 0:55I'm more interested to understand what
  25. 0:57I'm gaining from what I had before to
  26. 0:59what I can run today.
  27. 1:01And I can see that in some scenario, the
  28. 1:04gap is not that big,
  29. 1:06but other scenarios like here, agent
  30. 1:08decoding, the gap is huge from 13.3 to
  31. 1:1242.2, it's massive. And software
  32. 1:15engineer for 49.3 to 79, this is another
  33. 1:20massive step forward. On frontier agent
  34. 1:23task,
  35. 1:25from 10.6 to 20, this is 100%
  36. 1:28improvement. The gap sometimes are
  37. 1:32pretty noticeable. So, why I felt this
  38. 1:35mixed feeling when I I tested myself?
  39. 1:38Because
  40. 1:39I tested the Qwen 3.6 27 billion
  41. 1:44Q6
  42. 1:45with the Qwen 3.8 27 billion Q5. In my
  43. 1:50test, bug hunting, the Q6 found one more
  44. 1:54bug than and Q5, even though the Q5 was
  45. 1:57the newer version. So, I stopped the
  46. 2:00test and I said, "Okay, this is
  47. 2:02This is disappointing for
  48. 2:04us, but let's compare them with the same
  49. 2:06quantization." So, I also installed the
  50. 2:08Qwen 3.8 27 billion Q6, then I could see
  51. 2:12the difference. But, the difference was
  52. 2:14still not that big as much as I was
  53. 2:17hoping for.
  54. 2:19So, there is probably it's hard to tell,
  55. 2:21but let's say a 30% improvement, which
  56. 2:24is noticeable, but not another level.
  57. 2:27So, I was expecting much more. This is
  58. 2:29because Luna itself could do such a
  59. 2:33great job when improving reasoning. So,
  60. 2:36here is GPT 5.6 Luna non-reasoning 27.
  61. 2:41When we start increase reasoning,
  62. 2:44go to 34, then
  63. 2:4639,
  64. 2:48then
  65. 2:4947, and then GPT 5.6 Luna max 52. The
  66. 2:55jump is insane. It's more than double
  67. 2:58the capability.
  68. 3:00So, let's look at
  69. 3:02Qwen 3.6 27 billion non-reasoning 31.
  70. 3:06When does reasoning go to 38? I was
  71. 3:09expecting, okay, the new model is going
  72. 3:11to be better as a model, but plus then
  73. 3:13going to increase the reasoning, so it
  74. 3:15should push us towards this kind of
  75. 3:18level. We are probably in the 40, but I
  76. 3:21was hoping to get actually up to in the
  77. 3:2350s. That is reality. Considering that
  78. 3:27when non-reasoning 31 and Luna
  79. 3:30non-reasoning is 27. So, potentially, we
  80. 3:34could have done much more than 50 and
  81. 3:37reach, you know, this kind of level.
  82. 3:40The DeepSeek version 4 Pro.
  83. 3:43That would be amazing. The GLM 5.2 max.
  84. 3:46Okay, that would be an amazing
  85. 3:48achievement, but we didn't. So, probably
  86. 3:51I was expecting too much.
  87. 3:54But let me tell you the good news here.
  88. 3:58The good news is this model is capable.
  89. 4:02But
  90. 4:04intelligence is not just about the
  91. 4:06model.
  92. 4:07Is the harness around it.
  93. 4:10Before I explain why
  94. 4:12I'm talking about harness now, it
  95. 4:13because I found that the killer combo
  96. 4:16between this model and a harness. Until
  97. 4:20now I was using Hermes and Open Code.
  98. 4:24And while Hermes I find it that is a
  99. 4:27fantastic system.
  100. 4:30Recently, when they start pushing their
  101. 4:33own subscription, I felt like
  102. 4:36yeah, that is not cool anymore. I felt
  103. 4:39like a client.
  104. 4:40I'm a guest within their system, and
  105. 4:43this is the reason why I'm building with
  106. 4:45my community Resonate OS, so we can
  107. 4:48co-own it and it's ours.
  108. 4:50The moment they started pushing their
  109. 4:53subscription, the moment that people are
  110. 4:54starting asking me is do I need to
  111. 4:57subscribe to use Hermes?
  112. 4:59I could see that
  113. 5:01Hermes just took the wrong path.
  114. 5:05Okay, and therefore I was already open
  115. 5:07with the idea I need to look to
  116. 5:09something different.
  117. 5:11With Resonate OS, I was thinking, okay,
  118. 5:13for us everything needs to be an add-on,
  119. 5:16okay? All the elements need to be
  120. 5:19possible for the community to add to the
  121. 5:21system, so they can customize, they can
  122. 5:22change it, they can replace part, they
  123. 5:24can replace the memory, they can replace
  124. 5:26the main agent, they can replace all
  125. 5:28those parts, but at least we have a
  126. 5:30system that we can work together and
  127. 5:32build it and customize together.
  128. 5:35So, at one point I came across randomly
  129. 5:38to this new harness that came out
  130. 5:41recently. I think less than a a week ago
  131. 5:44Deep Seek released Deep Seek harness.
  132. 5:48Now, I believe that this is the best
  133. 5:51combo for when 3.8 27 billion.
  134. 5:56I started using this harness. I'm going
  135. 5:58to explain a little bit, even though I
  136. 5:59will probably need to do a dedicated
  137. 6:02video to this harness because it does
  138. 6:04something that no other harnesses do.
  139. 6:07What happened? I started using it
  140. 6:08together with this new version of Quen
  141. 6:11and everything I wanted to build so far,
  142. 6:14it built it. So, in the past, I used to
  143. 6:17use Quen 3.6 27 billion, let's say 80%
  144. 6:21of the time. Sometimes, it would find a
  145. 6:24problem that was not solvable. I had to
  146. 6:26change model and use one of those
  147. 6:29frontier model. Now, with this Deep Seek
  148. 6:32harness and the new 3.8 27 billion, I
  149. 6:35never
  150. 6:37had a problem to swap to a different
  151. 6:39model.
  152. 6:40So, what is happening here? When I went
  153. 6:43to this website, I read everything is a
  154. 6:45plugin. So, in my world, everything is
  155. 6:48an add-on, for them, everything is a
  156. 6:50plugin. It resonated and said, "Hmm, I
  157. 6:53need to look into this thing."
  158. 6:55I was coming from
  159. 6:57not totally sure about Hermes anymore.
  160. 7:00We are still building with our community
  161. 7:02resident of us.
  162. 7:03I see this that says everything is a
  163. 7:06plugin.
  164. 7:07Total alignment. I installed it.
  165. 7:09Actually, I didn't even install, I asked
  166. 7:11Hermes to install it for me and
  167. 7:13configure it with the new model
  168. 7:16and it's amazing. So, now, this is
  169. 7:20how it looks like, okay?
  170. 7:22And I can't show you while it's running
  171. 7:25because while it's running, I can't
  172. 7:27record the screen because memory
  173. 7:29management of the file format that I'm
  174. 7:32recording. What I can tell you here is
  175. 7:36that is just amazing. It's not a
  176. 7:39finished product, okay? Okay, so they
  177. 7:41call it
  178. 7:42developer preview. So it's not finished.
  179. 7:45It's there is a lot of things maybe I'm
  180. 7:47missing things that are not perfect. So
  181. 7:49sometimes it doesn't do compaction in
  182. 7:52the right moment.
  183. 7:54And therefore it's a kind of a crash the
  184. 7:56system. Nothing serious because you just
  185. 7:58say resume working and it start working
  186. 8:01without any problems. But nevertheless
  187. 8:03it's still a preview. But what it did
  188. 8:06they change how the context window is
  189. 8:09managed.
  190. 8:10What is happening here? It's
  191. 8:12unbelievable.
  192. 8:13There is a lot of new elements inside.
  193. 8:17So it takes track of everything that has
  194. 8:20been done.
  195. 8:21And require for whatever system that
  196. 8:24they use less context window. I can't
  197. 8:27explain because it's confusing to me so
  198. 8:29I need to read and understand it better.
  199. 8:31But what I know it's this 131,000 tokens
  200. 8:35that I have available for this model on
  201. 8:37my 5090.
  202. 8:39They just last longer. The compaction is
  203. 8:43so much more solid. And this system is
  204. 8:46being chatting for so long building for
  205. 8:49so long that is incredible. So if you
  206. 8:52see I don't know if you can see but here
  207. 8:54it had an input of 38 million tokens.
  208. 8:59Still going strong. What did I build
  209. 9:02with this? I build this the augmentor.
  210. 9:06This is my agent. Okay, this is how I
  211. 9:07call my agent. This is part of resonant
  212. 9:10OS is also included in resonant OS. The
  213. 9:12ability augmentor that is powered by
  214. 9:15this deep seek harness.
  215. 9:17What it does it's it takes total control
  216. 9:20of the browser and and just go around
  217. 9:24and click and do things for me. So what
  218. 9:27I ask here was to check out of my ladies
  219. 9:31videos to see the comments and suggest a
  220. 9:35few videos idea.
  221. 9:37So that was a way to send
  222. 9:39these agents into the browser, click
  223. 9:42around, and grab information.
  224. 9:46So, here is what came out. It checked
  225. 9:49four five different videos.
  226. 9:52It saw some of the comments, and then
  227. 9:54come out with some ideas.
  228. 9:56Now, it doesn't matter the ideas because
  229. 9:58it's honestly I don't from what I see
  230. 10:01maybe are valuable, maybe not, but
  231. 10:03doesn't matter because it didn't have
  232. 10:05the instruction of what is this channel
  233. 10:07is about and so on. So, at the moment it
  234. 10:08doesn't have much context. It just went
  235. 10:11there and grabbed information. But, what
  236. 10:13is important is what was capable of
  237. 10:16doing. I built the all thing with this
  238. 10:19Qwen 3.8.
  239. 10:21It works completely fine. Yeah, there is
  240. 10:24a a little things that need to be
  241. 10:25improved.
  242. 10:26And the most important part, I never had
  243. 10:29the need to use a bigger model.
  244. 10:32Something that always happened before
  245. 10:34when was doing something as complex as
  246. 10:37this kind of tasks. So, we have another
  247. 10:40deep seek moment, but this time the deep
  248. 10:42seek moment is not on the model, is on
  249. 10:44the harness, and is done by deep seek.
  250. 10:47Now, I'm extremely curious to see what
  251. 10:49this harness can do even with a more
  252. 10:52powerful model. If you've been following
  253. 10:54this channel, you know that I'm really
  254. 10:56pro local AI and so on.
  255. 10:59I'm for the AI sovereignty, meaning you
  256. 11:02need to own the hardware and
  257. 11:04infrastructure so you can trust and
  258. 11:07control everything. This is the journey,
  259. 11:09but it's not first of all it's not for
  260. 11:11everybody, but also is is a journey,
  261. 11:13it's not an on off.
  262. 11:15What I can see with this deep seek
  263. 11:16harness that the journey is getting
  264. 11:18better and better because now we have
  265. 11:20this harness that has MIT license,
  266. 11:24meaning you can use it in your project.
  267. 11:27You can use it yourself, implement,
  268. 11:29change it, do whatever, and use for
  269. 11:31commercial. So, what I'm suggesting
  270. 11:33here,
  271. 11:34if you are watching this video, probably
  272. 11:36is because you have access to this
  273. 11:38model.
  274. 11:40But, even if you don't, try this uh
  275. 11:42DeepSqueak Harness.
  276. 11:44If you're one of those lucky people that
  277. 11:46own a
  278. 11:48two DJX Spark
  279. 11:50or equivalent,
  280. 11:51try DeepSqueak version full flash,
  281. 11:54because can run on two of them with this
  282. 11:57harness, that probably they could work
  283. 12:00amazingly together. I'm at the moment of
  284. 12:02after a few months of uh
  285. 12:05of struggling with this AI,
  286. 12:07I'm excited again. I think it's a
  287. 12:10similar moment of when Open Claw came
  288. 12:13out, which was called uh
  289. 12:15Clawbot, I think. Can't even remember
  290. 12:17anymore.
  291. 12:19So, when it came out, it just was a
  292. 12:23massive push forward, because everybody
  293. 12:25finally had this agent that can do a
  294. 12:27lot. They control your computer, work
  295. 12:29for you, and so on. So, now I think what
  296. 12:33DeepSqueak did with their harness is a
  297. 12:36similar jump.
  298. 12:38Allows this model to do so much more,
  299. 12:42which is incredible.
  300. 12:43I will need to test this for longer, of
  301. 12:45course. I can't say,
  302. 12:47you know, that is the final solution,
  303. 12:50but at moment, this is what I feel. I
  304. 12:52feel like I never needed a more powerful
  305. 12:56model than this when I use this harness.
  306. 12:59Something that I I can't say the same
  307. 13:01with uh Hermes or Open Code.
  308. 13:05I think it's a killer combo. I need to
  309. 13:07understand this DeepSqueak harness even
  310. 13:09more. At moment, it is not fully
  311. 13:11finished. The way it works it's uh
  312. 13:14surprisingly well. So, if I were you,
  313. 13:17first of all, go to this DeepSqueak.com
  314. 13:21harness and uh and read it through. Like
  315. 13:24I said, I will do a new video, but they
  316. 13:26are definitely managing the things
  317. 13:28differently. And now, what is possible
  318. 13:30to achieve with this harness and the
  319. 13:32Qwen 3.8 27 billion is impressive.
  320. 13:37This is going to be my new way of
  321. 13:40working from now on. I will need to
  322. 13:43migrate and learn how to do it from
  323. 13:45Hermes to this model.
  324. 13:48Again, the good thing that is happening
  325. 13:50that is I basically already built this
  326. 13:53agent. This agent they are going to
  327. 13:55mentor.
  328. 13:56This can be a an add-on for Resonant OS
  329. 13:59and again shows the importance of
  330. 14:01working in add-ons or plugins like they
  331. 14:05call it because now it's easy to
  332. 14:07implement new things.
  333. 14:10And what I can even believe is that I
  334. 14:14built this
  335. 14:15100% with my local AI.
  336. 14:19Not needing a more powerful system.
  337. 14:23Everything that I asked to do has been
  338. 14:25done. Not getting lost and so on. One
  339. 14:29thing that I want to say about the
  340. 14:30psychic harness connected together with
  341. 14:33this Qwen 3.8 27 billion
  342. 14:36is they use a lot of token.
  343. 14:39They use a lot of tokens.
  344. 14:41But,
  345. 14:44the result is there. We know that the
  346. 14:46direction was increase the thinking time
  347. 14:48and you you will perform better. Here is
  348. 14:51happening and is incredible. Let's look
  349. 14:54at the what kind of hardware you need to
  350. 14:56run this model. In theory, according to
  351. 14:58Qwen, you can run it on 12 GB of RAM.
  352. 15:03Ideally, a 48 GB of RAM. And I would
  353. 15:06agree on the 48 GB of RAM because what
  354. 15:09you gain is actually a larger context
  355. 15:11window.
  356. 15:12I have the 32 GB of RAM. Okay, a Q6
  357. 15:16and 25 GB and look, I only lose a 0.2%
  358. 15:23from the Q8. So, the compromise is
  359. 15:26pretty acceptable. With the Q5, you lose
  360. 15:29a little bit. It's you lose 1.4.
  361. 15:32At 1.4, you say, "Okay, that is not
  362. 15:34much." And I mean, from this one, but
  363. 15:37you lose a 1.6
  364. 15:40from the top. It's still not that much,
  365. 15:42but from my benchmark that I run
  366. 15:45earlier,
  367. 15:47the bug hunting, the Q6 could find one
  368. 15:51more bug than the Q5.
  369. 15:55So, it is a little bit, you know,
  370. 15:57annoying, this thing. So, here's my
  371. 16:00suggestion.
  372. 16:01If you're looking for a bigger contest
  373. 16:03window, Q5 is fine. If you need to do
  374. 16:06some specific work that is more
  375. 16:08advanced,
  376. 16:10of course, you need to move into the Q6,
  377. 16:13ideally.
  378. 16:14If you have the 32 GB of RAM. I'm
  379. 16:17explaining this because it's important
  380. 16:20to understand what are the limits. When
  381. 16:22you go into the 16 GB of RAM, you start
  382. 16:25to lose quite a bit. At 12 GB of RAM,
  383. 16:29you lose a lot.
  384. 16:32So, while I never tried these two
  385. 16:35models, okay, I never tried them, I
  386. 16:38can't say how much you're losing from
  387. 16:40the model. This is something that you
  388. 16:41need to test yourself. But knowing that
  389. 16:43the minimum minimum requirement is 12,
  390. 16:46ideally, you want to be in the 24, 32,
  391. 16:5048 if you can do it. So, now, another
  392. 16:53element you need to consider is the size
  393. 16:55of the memory that is required to run
  394. 16:57this model.
  395. 16:58But the speed at which the model runs is
  396. 17:02connected to the bandwidth.
  397. 17:04So, I just listed here, I used these
  398. 17:07perplexities to do this research on the
  399. 17:1032 GB of RAM and 24 GB of RAM. So, here,
  400. 17:14for the 32 GB of RAM, the best one is
  401. 17:17the 1590, okay? This has a bandwidth
  402. 17:20that is insane, 1,792
  403. 17:23GB per second. That means the this model
  404. 17:27is working at around 100 and 10 token
  405. 17:30per second on my hardware. It's doing
  406. 17:32insanely well.
  407. 17:34So,
  408. 17:35when you go on this one,
  409. 17:38you can see that the AMD is doing 640.
  410. 17:43This is almost a third of the speed. So,
  411. 17:45you have an idea of what you can expect.
  412. 17:49Now, on the 24 GB of VRAM, we have a
  413. 17:53different kind of opportunity.
  414. 17:551,000 GB per second, which is pretty
  415. 17:58good. 3090 is the same. The good thing
  416. 18:01with 3090 that you can find them second
  417. 18:03hand much cheaper. This is the 3090
  418. 18:05idle. So, the 3090 normal 3090 is here,
  419. 18:09which is
  420. 18:11basically
  421. 18:12not not a big of a difference, okay? It
  422. 18:13doesn't really matter.
  423. 18:15But, from AMD, we have this card,
  424. 18:19which is still in commerce, that you can
  425. 18:21buy and it's not as expensive as those
  426. 18:254090. And the speed is almost there.
  427. 18:28Now, those are 24 GB of RAM. So, if I
  428. 18:31were in you and if you can afford, I
  429. 18:34mean, two of these together, they give
  430. 18:36you 48 GB of RAM, which again give us
  431. 18:41this 48 GB of RAM. I would not run the
  432. 18:44Q8. I would go for the Q6, but with a
  433. 18:46much bigger context window. Keep in mind
  434. 18:49that on my 32 GB of RAM, I can have
  435. 18:52131,000
  436. 18:54token context window. With 48, you
  437. 18:57probably, I guess, you could have the
  438. 18:59the total 260 something context window,
  439. 19:03which is a nice plus.
  440. 19:05Remember, Q8 is going to run slower than
  441. 19:08the the Q6, the Q6 slower than the Q5,
  442. 19:11and so on.
  443. 19:12Bandwidth, it is something to consider.
  444. 19:14At the moment, if you want to buy new,
  445. 19:16this is really good. If you're looking
  446. 19:18for second hand, this is pretty good.
  447. 19:21One more thing, this model can run on a
  448. 19:24DGX Spark and it runs, depending on the
  449. 19:26quantization, at around 20 30 token per
  450. 19:29second. At 30 token per second, in
  451. 19:32reality, you have a system that works,
  452. 19:34is usable, and I will probably run it
  453. 19:36and test it on my GX 10.
  454. 19:39This is the ASUS version of the DGX
  455. 19:41Spark. It's 128 GB of RAM, so I can run
  456. 19:44it with really large context window with
  457. 19:47multiple instances. So, that would be
  458. 19:49the advantage.
  459. 19:51Anyway, that's is what I wanted to share
  460. 19:53for today.
  461. 19:54To recap,
  462. 19:56a model that is not as good as I wanted.
  463. 19:58It didn't do the jump that I was
  464. 20:00expecting and wanting. Probably we're
  465. 20:02going to see it when they release the
  466. 20:03version four of it. But, combining this
  467. 20:07with Deep Seek harness, I saw something
  468. 20:11that I never saw before. It's early days
  469. 20:13to say, but so far, in the last 24
  470. 20:15hours, I use it a lot.
  471. 20:17And I never felt I needed to use a
  472. 20:22bigger model. Never once.
  473. 20:24[clears throat]
  474. 20:25That make me think that things are
  475. 20:27changing and locally I is gaining more
  476. 20:31and more value
  477. 20:33every single day.
  478. 20:34If you're interested in this kind of
  479. 20:35content and you want to join our
  480. 20:37community, first like and subscribe. In
  481. 20:39the description, you'll find the link to
  482. 20:41our Discord channel and there is a deep
  483. 20:44dive article. I will go deeper in
  484. 20:47everything that I'm talking inside this
  485. 20:49video. I will add all the links that I
  486. 20:52discussed over here and if possible, I
  487. 20:54will see how it goes, but I might also
  488. 20:57share
  489. 20:58this augmentor power by this DSH. This
  490. 21:02is an extension for the browser and by
  491. 21:06adding it, you can have this sidecar.
  492. 21:10Which, honestly, is amazing. It is
  493. 21:13actually amazing.
  494. 21:15I say if I can, because what is missing
  495. 21:18before I can share it is a way for you
  496. 21:21to add your own model. At the moment,
  497. 21:24there's no
  498. 21:25configuration system where you can add
  499. 21:28your API or local model or something.
  500. 21:31Even though through the code, you could
  501. 21:33be able to ask to your agents to add the
  502. 21:36model, I want to make it through the
  503. 21:38interface so everybody can just add
  504. 21:40their, you know, details and run it
  505. 21:43locally or for any other cloud provider.
  506. 21:47So, if I can manage to do this in time,
  507. 21:50I will link inside the Substack article
  508. 21:53the code for this
  509. 21:55agent. If not, will be shared in the in
  510. 22:00the near future. Thanks for staying
  511. 22:01until the end. Again, like and
  512. 22:03subscribe. Watch the next video. Join
  513. 22:05the Discord. And ciao.

About this transcript

This page contains the full transcript of I Don't Need Frontier Models Anymore (Qwen 3.8 27B + DeepSeek Harness) by Manolo Remiddi, generated from the public captions YouTube serves with the video. The transcript has 3,334 words across 513 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.