YouTube2Text

Is STRIX Better than SPARK? Now Launching w/new Software: AMD's Ryzen AI Halo Developer Workstation — Transcript

by Level1Techs · 5,913 words · 840 segments · language en · Watch on YouTube

Full transcript

  1. 0:00There it is. Ryzen AI Halo. AMD built
  2. 0:04Stricks Halo, the chipset on which this
  3. 0:05is based, as a workstation APU. We first
  4. 0:08saw it something like 18 months ago. We
  5. 0:12crazy internet randos looked at it then
  6. 0:14and we immediately saw the 128 gigs of
  7. 0:16memory and we said, "Cool. What if we
  8. 0:19use that to do crimes against model
  9. 0:21size?" [music]
  10. 0:28AMD was really early here. That was even
  11. 0:30before Nvidia launched their 128 gig DJX
  12. 0:33Spark. And weirdly, Stricks Halo worked.
  13. 0:35It definitely worked for AMD. I mean, 48
  14. 0:38gig coding models at roughly 60 tokens
  15. 0:41per second. A 92 gig mixture of experts
  16. 0:43model roughly 20 tokens per second.
  17. 0:45Local, quiet, on a desk, less than 200
  18. 0:49watts. So, why is this product coming
  19. 0:50out now? Well, there there were some
  20. 0:53problems with the software side of
  21. 0:55things. It was never the silicon. The
  22. 0:57problem was always everything around it.
  23. 0:59Rockom versions and Rockom modifications
  24. 1:02for Stricks Halo, memory allocation
  25. 1:04myths, BIOS settings, container bugs,
  26. 1:06Vulcan escape hatches, and enough GitHub
  27. 1:09issues to make you reconsider your life
  28. 1:11choices. At least the life choices that
  29. 1:14led to working on AI stuff in the first
  30. 1:16place. But this this is the story. This
  31. 1:20is the the the pin in the corkboard, if
  32. 1:22you will, for AMD having hit both
  33. 1:25software and hardware milestones. It's
  34. 1:28not just a known good stricks Halo
  35. 1:30system with Linux in my case, also
  36. 1:32available with Windows. Playbooks and
  37. 1:33firmware updates piped right into the
  38. 1:35Linux way. LM Studio support, that's
  39. 1:37what I'm running here. Comfy UI, which
  40. 1:39is in the background, that's not a
  41. 1:41problem. This is the standard bearer for
  42. 1:43what comes next. It points toward Gorgon
  43. 1:46Halo, a larger unified memory machine,
  44. 1:48192 gigs coming with that update in Q3,
  45. 1:51but also the rest of the Stricks Halo
  46. 1:54ecosystem is going to be pulled up
  47. 1:55behind this, which is great. Fixed
  48. 1:58playbooks, firmware updates, clarified
  49. 2:01memory behavior, known good rock stacks,
  50. 2:04and known good Vulcan stacks, known good
  51. 2:07playbooks. And that's going to elevate
  52. 2:08everything from Minis Forum and
  53. 2:11Framework Desktop and GMK Techch and all
  54. 2:13the other versions of Stricks Halo out
  55. 2:15there. And that is what we can rally
  56. 2:18around. This is the real milestone. It's
  57. 2:21not just a flagship box, but a reference
  58. 2:23point the whole platform can finally
  59. 2:26rally around. A Halo's Halo product, if
  60. 2:31you will. This is AMD's Ryzen AI Halo
  61. 2:33developer platform. And yes, the name is
  62. 2:35funny because it's Stricks Halo. It's
  63. 2:37already a Halo and this is already a
  64. 2:38Halo product. This is a Halo Halo
  65. 2:41product. Somewhere there's a product
  66. 2:43manager that's glowing. The hardware is
  67. 2:45familiar. Ryzen AI Max 395 plus 16 Zen 5
  68. 2:49cores. Radeon 860S graphics. 128 gigs of
  69. 2:51unified LPDDR5X.
  70. 2:53A compact attractive chassis available
  71. 2:56either with Linux or Windows. The price
  72. 2:58$4,000. Uh that price is maybe the first
  73. 3:01problem, but we need to chat about it.
  74. 3:04But more to the point, Framework, Minis
  75. 3:06Forum, GMK Techch, Beink, Corsair, HP,
  76. 3:08and others have already shipped Stricks
  77. 3:10Halo based systems, and some are much
  78. 3:12cheaper, some are much more expandable.
  79. 3:15I've got a 25 gig Ethernet card in here.
  80. 3:17The Framework Desktop also has a PCIe
  81. 3:19slot, but in their case, it's not
  82. 3:20accessible. But if you put the
  83. 3:22motherboard in a different case, it's
  84. 3:23totally okay. GMK Tech gives you an
  85. 3:26Oculink slot. You've got options. Almost
  86. 3:28all of these have multiple M.2 slots,
  87. 3:30except for the one from AMD. It's got a
  88. 3:322280 at least. So, I mean, I'll take it.
  89. 3:35But, yeah, it's a lot of fun. Oh, and
  90. 3:38the fact that I've got a 25 gig dual
  91. 3:40port Intel E810 Nick in the minis forum,
  92. 3:42I think that does change the
  93. 3:43conversation a little bit around
  94. 3:45clustering. I mean, if you look at the
  95. 3:48networking here, you've just got the one
  96. 3:50realtech 10 gig nick. And that's all
  97. 3:53that is in our AI Halo system. But you
  98. 3:55can still cluster it. You can still
  99. 3:57build some clustering stuff. In fact, I
  100. 3:59think you should check out Donado's
  101. 4:00video on this. He has been building a
  102. 4:03Stricks Halo cluster. He's the Stricks
  103. 4:05Halo toolboxes guy and cluster like this
  104. 4:08out of that. Yeah. Yeah, it works. It
  105. 4:09was a latency problem the whole time.
  106. 4:10But back to this. Why does this exist?
  107. 4:13Why what is AMD doing here? Cuz it's not
  108. 4:15remotely about Stricks Halo, at least
  109. 4:17not this late in the product cycle. I
  110. 4:19think this is AMD planting a flag and
  111. 4:22paving the way for Gorgon Halo, the next
  112. 4:25generation product that rhymes with this
  113. 4:26one. Probably just a chip swap and
  114. 4:29memory swap as well. That's right around
  115. 4:30the corner. Q3. So AMD says. So we have
  116. 4:34at the back, if we take a look at the
  117. 4:35rear IO, USBC power from the 240 watt
  118. 4:38power brick, 10 gigabit Ethernet, thanks
  119. 4:40to our realtech chipset, three other
  120. 4:42USBC ports. One is recommended for
  121. 4:44display port alt mode, and two others
  122. 4:46that support alt mode and USBC
  123. 4:48connectivity. There is no other
  124. 4:49high-speed networking, Oculink or even a
  125. 4:53second M.2 internally. It's just the
  126. 4:56one. At least I don't think I'm pretty
  127. 4:57sure there's not cuz I took it apart. we
  128. 4:59can all look together. Taking apart is
  129. 5:02easy and and and repeatable though.
  130. 5:03There are screws underneath the magnetic
  131. 5:05feet. Uh there's also a Kensington lock
  132. 5:07port, so that's appropriate for
  133. 5:09university labs and that sort of thing
  134. 5:11where maybe you might have one of these
  135. 5:12and it would walk off. But physically,
  136. 5:15this is uh one of the most liipian
  137. 5:18options for Stricks Halo. Now, the
  138. 5:20obvious comparison that I've alluded to,
  139. 5:21Nvidia DJX Spark. Both are small local
  140. 5:23AI workstations. Both have 128 gigs of
  141. 5:25unified memory. Both target inference
  142. 5:28and developer workflows. Both claim
  143. 5:31support for very large models under the
  144. 5:34right quantization and optimization
  145. 5:36parameters. Right now on Newegg, you can
  146. 5:40still buy Spark for about 4,000 to 4500.
  147. 5:43DGX Spark has the stronger platform
  148. 5:45story where Nvidia tends to always come
  149. 5:48out ahead. I mean CUDA, DGXOS, container
  150. 5:50workflows, and the scale out networking.
  151. 5:53The scaleout networking really is
  152. 5:54something special. Uh the networking
  153. 5:57difference matters. I think Spark's
  154. 5:58ConnectX path is a real pro feature and
  155. 6:00not a spec sheet decoration. If you are
  156. 6:02trying to stitch two boxes together,
  157. 6:05Nvidia starts ahead. And the skills that
  158. 6:08you gain from dealing with nickel NCCL,
  159. 6:10that's Nvidia's communication primitives
  160. 6:12library. The skills from that will
  161. 6:14transfer to multi-million dollar AI
  162. 6:16clusters. So you you learn and you do
  163. 6:18stuff. AMD's counterpunch is flexibility
  164. 6:21and at least historically price. This is
  165. 6:24x86. It runs Linux. It has also always
  166. 6:28run Windows, whereas Nvidia is planning
  167. 6:31Windows support for their stuff, but
  168. 6:33it's not here yet. And AMD uses open-ish
  169. 6:36tooling. It does not require you to buy
  170. 6:38into the CUDA world view. And you know,
  171. 6:41for many inference workloads, especially
  172. 6:43token generation, stricks Halo is
  173. 6:46basically the same performance as DJX
  174. 6:47Spark. And that's because both systems
  175. 6:50are ultimately gated by the unified
  176. 6:53memory bandwidth and it's basically the
  177. 6:55same. It's very similar between the two
  178. 6:56platforms. The short version for view
  179. 6:58like anybody that doesn't have the
  180. 7:00attention span, Spark is cleaner for
  181. 7:03Nvidia's approach, but the Ryzen AI Halo
  182. 7:05is more flexible workstation. Spark has
  183. 7:07the better networking story. Halo has
  184. 7:09the better I could just use this like a
  185. 7:11PC story. Oh, and I'm not sleeping on
  186. 7:14networking over USB 4. That is an option
  187. 7:16here. here. You can use those USBC ports
  188. 7:18to connect to other machines, kind of
  189. 7:19like a Mac, but I want RDMA to be a part
  190. 7:22of that story when we talk about it
  191. 7:23because RDMMA is not here yet for that.
  192. 7:25Um, you can do it over USB 4,
  193. 7:28Thunderbolt, and it's on the order of
  194. 7:3020ish gigabit, give or take, but uh,
  195. 7:33yeah, I mean, I don't know. Nvidia made
  196. 7:36some weird choices with Spark 2, like
  197. 7:37they didn't use a 2280 M.2. 2280,
  198. 7:40everything on this table 2280, which is
  199. 7:42the only sane option. Why would they do
  200. 7:45that? unless they just wanted to put a
  201. 7:47crappy, you know, it's like, oh, we
  202. 7:48can't just buy a 2280. So, yeah, because
  203. 7:51[snorts] this is a pre-existing
  204. 7:52platform, Stricks Halo, uh, and because
  205. 7:56it's from AMD, I got to go off on a
  206. 7:58tangent for a second. There's a lot of
  207. 8:01bad information online about Stricks
  208. 8:02Halo, especially around memory and how
  209. 8:05the memory works in a unified platform.
  210. 8:06People say there's a difference between
  211. 8:08what Nvidia is doing with ARM and
  212. 8:10Stricks Halo. That is not true. it it's
  213. 8:12it's they they treat it like as if the
  214. 8:13machine has two physical memory pools.
  215. 8:15It does not. On Linux, the important
  216. 8:18mechanism that the GPU uses to map
  217. 8:20through the shared memory pool is
  218. 8:23through GTT and TTM. For most AI
  219. 8:25workloads, you do not need to reserve 96
  220. 8:28or 112 GB of memory permanently for the
  221. 8:30GPU. When I was doing gaming reviews on
  222. 8:33laptops that were based on stricks Halo,
  223. 8:36really killer machines like the ROG Flow
  224. 8:38Z13 and the HP G1A, I showed games like
  225. 8:42Clar Obscure running at a high frame
  226. 8:44rate but with only 512 megabytes of
  227. 8:47reserved VRAM. Except in some really
  228. 8:50legacy software scenarios, it's often
  229. 8:52the case that for AI workloads, this is
  230. 8:54the better option as well. keep the BIOS
  231. 8:57reserved GPU memory very small and then
  232. 8:59let the Linux kernel driver expose the
  233. 9:01large shared pool dynamically. The the
  234. 9:04old mental model for this and and why it
  235. 9:07shows up wrong a lot of the time I think
  236. 9:08comes from Windows. On Windows, the
  237. 9:11split is more visible and more annoying.
  238. 9:13AMD's variable graphics memory exists
  239. 9:16because software still asks things like
  240. 9:18how much VRAM do we have and it'll freak
  241. 9:20out if it encounters 512 megabytes. But
  242. 9:22most software has been updated to be
  243. 9:24aware of the thing that exists with
  244. 9:26unified memory platforms. This is a
  245. 9:28major problem for Nvidia too with their
  246. 9:30planned RTX Spark, the stuff that we saw
  247. 9:31at Compyex just just a few weeks ago
  248. 9:34because you know Windows and so of
  249. 9:36course I checked how they were handling
  250. 9:39Windows on the machines that I saw there
  251. 9:41and guess what Windows got fixed and the
  252. 9:43unified memory fix there benefits AMD
  253. 9:46too. That's sort of the funny part here.
  254. 9:48The work that Microsoft has done
  255. 9:50apparently fasttracked because now
  256. 9:53Nvidia needs it as well ends up greatly
  257. 9:55helping AMD's Windows Stricks Halo
  258. 9:58experience. I'm not got anything Windows
  259. 10:00related in this video. Probably a
  260. 10:02follow-up video if I can get my hands on
  261. 10:03the software image. But Stricks Halo in
  262. 10:05a nutshell, it is not a 96 gig GPU. It's
  263. 10:08128 gig unified machine. It can address
  264. 10:10a large pool of memory. And when the OS
  265. 10:12and the driver stack are configured
  266. 10:13correctly, I've got up to 112 GB of
  267. 10:16memory on this platform for the GPU
  268. 10:18without any issues or headaches or any
  269. 10:19really like special hackery that I had
  270. 10:21to do. It just works. So the people that
  271. 10:23complain about memory allocation or like
  272. 10:26freezing or hanging or something like
  273. 10:28that, they're they're not facing
  274. 10:30problems with the memory allocation.
  275. 10:32It's probably something else because
  276. 10:33it's it is a a wellexecuted unified
  277. 10:36platform. Sometimes an LM Studio just
  278. 10:39toggling off try MAPAP is all you need
  279. 10:42to do. That's the thing that hang people
  280. 10:43get hung up on sometimes. But I digress.
  281. 10:46So now AMD has this playbooks. This is
  282. 10:50the thing that I literally told Anush at
  283. 10:53Advancing AI. AMD needs a website of
  284. 10:55known good recipes that are part of your
  285. 10:58CI/CD process. Every time AMD pushes an
  286. 11:01update, every run through, every little
  287. 11:03thing, your CI/CD system kicks off and
  288. 11:06runs through all of these playbooks to
  289. 11:09make sure that it all works. Andy should
  290. 11:11have had this on launch day. The reason
  291. 11:13they did not is probably simple. Stricks
  292. 11:16Halo was not originally positioned as an
  293. 11:18AI developer appliance. These are useful
  294. 11:21playbooks for the individual as well
  295. 11:23because they run like these are common
  296. 11:25things that you would want to do. The
  297. 11:27community realized early on that strikes
  298. 11:29Halo was was going to be good, too good
  299. 11:31at inference to ignore. The playbook's
  300. 11:32turn, I spent all weekend trying to weed
  301. 11:35out if it runs better with Rockim or
  302. 11:36Vulcan into open the developer center,
  303. 11:40find the thing that's close to what you
  304. 11:41want to do, and run that. And out of the
  305. 11:44box, shipped with this thing, over 380
  306. 11:46gigs of stuff. It's got lemonade and web
  307. 11:48UI and VS Code. Well, you got to
  308. 11:50download VS Code, AMD Sync Workflows.
  309. 11:52There's a whole bunch of stuff on here
  310. 11:54that is ready to go out of the box in
  311. 11:56both in terms of both models and
  312. 11:58configuration. Lemonade server from AMD
  313. 12:00as I mentioned and comfy UI. This is
  314. 12:02this is amazing. This is AMD putting the
  315. 12:04most effort I have ever seen from AMD
  316. 12:06into making the client side experience
  317. 12:09smooth. And that is the real review.
  318. 12:11It's not really this box. It's the
  319. 12:12software. It's the mindset because all
  320. 12:14of this executed properly benefits all
  321. 12:18of this. That's what I meant when I said
  322. 12:19it's going to bring up everything. Yes,
  323. 12:21you could run 122 billion parameter uh
  324. 12:23class model on this stricks halo. That
  325. 12:26is an amazing sentence to be able to to
  326. 12:28utter in 2026, but I do not think the
  327. 12:31best daily driver layout is necessarily
  328. 12:34one enormous model eating the whole
  329. 12:37machine. The better version of this
  330. 12:38platform is probably a competent 27 to
  331. 12:4230 billion parameter coding or reasoning
  332. 12:44model as kind of the main brain with a
  333. 12:47constellation of smaller models around
  334. 12:49it. Uh maybe one for routing, maybe one
  335. 12:52for summarization, maybe one for
  336. 12:53retrieval cleanup. Retrieval augmented
  337. 12:55generation is a fun thing to explore,
  338. 12:56maybe one for vision or OCR, fast
  339. 12:59drafting, maybe one for tool use
  340. 13:01verification. Smaller models do prompt
  341. 13:04processing faster, respond sooner, and
  342. 13:06leave room for deep context and parallel
  343. 13:08work and make some of the slower running
  344. 13:10models a little bit more tolerable. How
  345. 13:12big a model can you cram into memory is
  346. 13:14not really that awesome. It's how many
  347. 13:17useful local intelligences can you keep
  348. 13:18alive at once I think okay so real world
  349. 13:22performance yes you can run 120 billion
  350. 13:24parameter model on this but that's not
  351. 13:26really the point and actually a dense
  352. 13:28120 billion parameter model uh would be
  353. 13:30difficult understand if you're just
  354. 13:32getting into this the difference between
  355. 13:33mixture of experts and a dense model and
  356. 13:36I'll use Quinn as an example there's a
  357. 13:3735 billion parameter Quinn model and
  358. 13:39there's a 27 billion parameter Quinn
  359. 13:41model the 27 billion parameter Quinn
  360. 13:43model on this is going to run 12 to 16
  361. 13:47to 20 tokens per second depending on
  362. 13:49some parameters and that's at a Q8 it's
  363. 13:528 bits per weight quantization you can
  364. 13:55run four and six bit but I think with
  365. 13:57Quinn you start to lose some of the
  366. 13:59fidelity of the model when you're much
  367. 14:01below Q8 it's actually uh natively 16
  368. 14:05bit so a 27 billion parameter dense
  369. 14:08model would be double that you need like
  370. 14:1048 gigs and yes you can run that on this
  371. 14:12but you're going to be in a relatively
  372. 14:15low uh token throughput situation.
  373. 14:18Mixture of experts, it's like Quinn 35B.
  374. 14:22It's a 35 billion parameter model, but
  375. 14:24that A3B means that only three billion
  376. 14:26parameters are active at once. Those
  377. 14:28kinds of models, mixture of experts
  378. 14:29models are great for systems like this
  379. 14:32that are memory bandwidth constrained
  380. 14:34because only three billion parameters
  381. 14:36are active at any one time. And the
  382. 14:38quality output between the 35 billion
  383. 14:41and 27 billion parameter models like how
  384. 14:43how good they are like how good they are
  385. 14:44at doing tasks is pretty similar. The
  386. 14:47setup that I have here is with Turnstone
  387. 14:50and I'm I'm doing something really
  388. 14:51interesting. Remember the 4GPU R9700
  389. 14:54system that I'm running that also has
  390. 14:56128 gigs of memory across four GPUs
  391. 14:58that's in the basement. This system is
  392. 15:01controlling it. This is an orchestrator
  393. 15:04and this is toward that whole agentic
  394. 15:07thing that everybody talks about. So
  395. 15:09Turnstone is a piece of software that's
  396. 15:11available on GitHub and it's open source
  397. 15:13and it is amazing and I need to do
  398. 15:15another video on it because it's kind of
  399. 15:17like Hermes or Open Claw but it's you
  400. 15:19know enterprise mindset. It does some
  401. 15:20things a little differently and again
  402. 15:23separate conversation but turnstone
  403. 15:26running on this with a you know Quinn 35
  404. 15:30billion parameter model is smart enough
  405. 15:32to be a supervisor. Now that is nowhere
  406. 15:34near as smart as the best models that
  407. 15:37you can get on you know Open Router or
  408. 15:40OpenAI or Anthropic. Those are frontier
  409. 15:43models. uh but you can connect those
  410. 15:45models to Turnstone through an API but
  411. 15:48you can also use this to minimize the
  412. 15:50amount of tokens that you spend in the
  413. 15:52cloud. You see this as an orchestrator
  414. 15:54doesn't have to be the smartest model to
  415. 15:57for you to get utility out of it. Um and
  416. 16:00with a 35 an A35B type uh model with
  417. 16:05multi-token prediction that's another
  418. 16:07thing that you probably want to try to
  419. 16:08do. It means that you try to get more
  420. 16:10than one token per uh prediction and the
  421. 16:12acceptance rate on that is typically 60
  422. 16:1470 80%. That'll improve your throughput
  423. 16:17in [snorts] addition to the whole only
  424. 16:19three billion parameters active. A lot
  425. 16:21of words to say that you can get 30 40
  426. 16:2350 60 tokens per second per session and
  427. 16:27three four sessions running in parallel.
  428. 16:29And so like with Turnstone, when you run
  429. 16:31an orchestrator like this and you give
  430. 16:33it a task, it can split up the task into
  431. 16:35four or five pieces and work on them
  432. 16:37simultaneously. And each piece that it's
  433. 16:39working on doesn't have to run with the
  434. 16:41same model. So like you don't have to be
  435. 16:42limited to Quinn. It's like, oh, the 35
  436. 16:45billion parameter that's going to fit
  437. 16:46in, you know, 48 gigs of VRAM. What do I
  438. 16:48do with the rest of the VRAM? You can
  439. 16:49run more models. You can run a model
  440. 16:51that's really good at OCR, really small
  441. 16:53model. Um, GMA from Google is also
  442. 16:56really good. They have some small models
  443. 16:58that are 3 4 billion parameters that are
  444. 17:01fantastic. Um, one of the reasons that
  445. 17:03Quinn is so amazing is because they
  446. 17:05publish papers and make it really easy
  447. 17:07for you to do your own fine-tuning. And
  448. 17:09so even though those those models are
  449. 17:11smaller, they're also larger. And there
  450. 17:13are smaller Quinn models that you can
  451. 17:15fine-tune on a platform like this. Now,
  452. 17:17it'll take a couple of weeks. like it'll
  453. 17:19run for a long time but you can set that
  454. 17:21up on a platform like this and then you
  455. 17:24have customized the model and trained it
  456. 17:27on your data. The difference between
  457. 17:29retrieval augmented generation and uh
  458. 17:32you know fine-tuning is that you have a
  459. 17:34lot of data that doesn't really change
  460. 17:35and you want the model to be read in on
  461. 17:38the data that you wanted to have it
  462. 17:39access to fine-tuning is probably worth
  463. 17:41it. But retrieval augmented generation
  464. 17:44is you sort of frontload a lot of data
  465. 17:47into the model with the requests. And so
  466. 17:50retrieval augmented generation you have
  467. 17:51like a vector database or something like
  468. 17:53that. You take information and you do
  469. 17:55some processing to it and at runtime it
  470. 17:59is injected into the model in a
  471. 18:01different way than fine-tuning does sort
  472. 18:03of kind of and so with every request the
  473. 18:06model has the context of the the
  474. 18:08documents that you've added and so
  475. 18:10that's why I say retrieval augmented
  476. 18:11generation it's generating tokens
  477. 18:13augmented by the fact that it retrieved
  478. 18:16the contents of the documents in a
  479. 18:18vector format or something like that not
  480. 18:20the raw format you got to process the
  481. 18:22documents that you're feeding for rag.
  482. 18:23Whereas fine-tuning, you run this
  483. 18:25process on the model, the model is
  484. 18:28modified, and then when you run the
  485. 18:30model, it has information about the
  486. 18:33documents that you gave it for the
  487. 18:35fine-tuning process. Both of those
  488. 18:37things are achievable on this platform
  489. 18:39because of good documentation that goes
  490. 18:41along with the model. And Quinn has been
  491. 18:44a standout among all of the options
  492. 18:46available in terms of how well it is
  493. 18:48documented. And it is, this platform is
  494. 18:50fast enough for you to to do those kinds
  495. 18:52of things both academically and in an
  496. 18:54actually useful kind of way. Now, for
  497. 18:56like software development and things
  498. 18:57like that, there's a lot of really
  499. 18:59low-level bog standard like janitorial
  500. 19:01type things like I need help organizing
  501. 19:03this git repo. I need help writing a
  502. 19:05unit test. I need help, you know, doing
  503. 19:06this evaluation. I need you to run
  504. 19:08through the unit tests or help me figure
  505. 19:10out why this one is failing. You don't
  506. 19:11need a frontier model for a lot of those
  507. 19:13kinds of really low-level tasks. And you
  508. 19:16can run that on here with Turnstone or
  509. 19:18any other harness that you want and that
  510. 19:20works fine and it runs acceptably fast
  511. 19:22in my opinion. The utility of these
  512. 19:25kinds of models and the state-of-the-art
  513. 19:27like how the software has moved the
  514. 19:29hardware over the years is a big part of
  515. 19:31the story here. The the software that's
  516. 19:33available now in mid to late 2026 is
  517. 19:37light years ahead of where we were a
  518. 19:38year ago, even though the hardware
  519. 19:40hasn't really changed a year ago. So it
  520. 19:41does feel like a new platform and it is
  521. 19:43really nice to see that level of
  522. 19:45development and speed. So the big thing
  523. 19:47with Agentic and benchmarking in a
  524. 19:49platform like this is don't think of AI
  525. 19:50as monolithic. Don't sit down. It's like
  526. 19:52we're going to oneshot a Tetris clone.
  527. 19:54If I ask this thing to do a Pac-Man
  528. 19:57clone or something like that, it's going
  529. 19:59to spawn a bunch of tasks. It's going to
  530. 20:01put a plan together and then there's
  531. 20:03going to be a bunch of agents that are
  532. 20:04running maybe on the same model, maybe
  533. 20:06on different models in order to complete
  534. 20:07the tasks and that kind of thing can
  535. 20:09function as a benchmark. But know that
  536. 20:11because of the memory bandwidth here,
  537. 20:12the performance for most models is on
  538. 20:14par with Nvidia's DGX Spark now in 2026.
  539. 20:18Now, I mentioned some off label use
  540. 20:20cases. All about the off label use
  541. 20:23cases.
  542. 20:24Thunderbolt, it's not really
  543. 20:26Thunderbolt, it's USB 4. Two of the
  544. 20:28three ports are capable of USB4 20 GB.
  545. 20:33So, here you can see we've got our
  546. 20:34external Razer uh X2 Thunderbolt 5 dock
  547. 20:39connected and it's got a big honken GPU
  548. 20:42in it and we've got 20 Gbit both ways.
  549. 20:44So, 40 Gbit and their 40 GB interface
  550. 20:47and we can see that our GPU came up uh
  551. 20:51properly. And this is a big deal because
  552. 20:53when Stricks Halo first launched, they
  553. 20:55didn't really plan on having GPUs. Like
  554. 20:58even the folks like me that took our
  555. 21:00framework desktop. This is one of the
  556. 21:02first stricts halo machines that I got
  557. 21:04and extended out the X4 slot. This X4
  558. 21:07slot can only provide 25 watts. You need
  559. 21:0975 watts for a GPU. So you got to do
  560. 21:11some herculean things to even get a GPU
  561. 21:14plugged into that PCIe slot which just
  562. 21:16four lanes. Wouldn't post because the
  563. 21:19BIOS wasn't plumbed to do that. It's a
  564. 21:21firmware. It's a firm AMD firmware team.
  565. 21:23Like it's like they're so isolated. They
  566. 21:26got to they got to come into the fold.
  567. 21:28We got to all work with the AMD firmware
  568. 21:30team. Um, and so this is a fix. This is
  569. 21:32this is a fix since launch, which is
  570. 21:34great. It's great to see this. And this
  571. 21:35worked well with uh some other different
  572. 21:37configurations, even running two of
  573. 21:38these things also for Thunderbolt
  574. 21:40networking. Now, you can do Thunderbolt
  575. 21:42networking and it is relatively low
  576. 21:44latency, but you really need um RDMA,
  577. 21:48not RDMA, RDMA, a remote direct memory
  578. 21:51access. Remote director memory access
  579. 21:53gives us would theoretically give us 20
  580. 21:55gigabit very low latency connection not
  581. 21:58through a real tech interface. It's not
  582. 22:00as good as you know something that's 100
  583. 22:02200 gigabit but it'll get the job done.
  584. 22:05I really think that AMD could have put
  585. 22:09something based on Polar or maybe a
  586. 22:12little FPGA in there. I think one of the
  587. 22:14like a competitive differentiation thing
  588. 22:16that AMD could have done is give you
  589. 22:17another M.2 too. And maybe with like a
  590. 22:20an FPGA that would give you maybe
  591. 22:23something you could experiment with in
  592. 22:24terms of oh, we can just do our own
  593. 22:26custom fabric. Maybe even direct PCIe
  594. 22:28connection because the PCIe it's not
  595. 22:30really that much of a lift to provide a
  596. 22:33software layer to do a direct PCIe to
  597. 22:36PCIe connection through, you know,
  598. 22:38Oculink or something like that with a
  599. 22:40little bit of custom glue. But that's
  600. 22:43not really the thing for this video. But
  601. 22:46this is really exciting that you have an
  602. 22:47interface. You can also kind of build an
  603. 22:48unholy machine here because okay, I've
  604. 22:50got 48 gigs of really high-speed GPU
  605. 22:53VRAM plus 128 gigs for mixture of
  606. 22:57experts type models. This is unholy. And
  607. 23:01we've done videos on that in the past
  608. 23:03where you're running uh a high
  609. 23:05performance GPU that has 16 32 48 96
  610. 23:10gigs of memory plus 128 gigs of much
  611. 23:13slower LP uh you know our our DDR5 here.
  612. 23:17Um, but that lets you run much larger
  613. 23:20mixture of experts models and also let
  614. 23:22you run smaller mixture of experts
  615. 23:24models much faster and also a bunch of,
  616. 23:27you know, agents or a bunch of models
  617. 23:30running simultaneously on a platform
  618. 23:32like this imminently doable. Oh, and AMD
  619. 23:34has found a way to make the NPU useful
  620. 23:37in this cuz remember it's designed as a
  621. 23:38laptop. There's an NPU. You can run
  622. 23:41multimodel scenarios and use the NPU.
  623. 23:43things like speech recognition and some
  624. 23:45of the larger models will also work on
  625. 23:47the NPU at shallow context with
  626. 23:49reasonable speed. Uh there's some
  627. 23:51amazing benchmarks of fast flow LM in
  628. 23:53exactly this scenario. You can run Llama
  629. 23:563.2 1 billion at over 50 tokens a second
  630. 23:58below an 8K context and the prefill
  631. 24:01speed with that is over 2,000 tokens per
  632. 24:03second. So if you've got a a pretty
  633. 24:06awesome carefully planned setup, you can
  634. 24:08do a lot with that. Now before AMD had a
  635. 24:11polished story, Stricks Halo made their
  636. 24:13own. Uh Benado that I mentioned, he made
  637. 24:15the Stricks Halo toolboxes. Those are
  638. 24:17the obvious example containerized
  639. 24:19environments for uh LLMs, image
  640. 24:22generation, fine-tuning, rock, vulcan,
  641. 24:24llama.cpp. Uh he's top of mind right now
  642. 24:28because we're working on something
  643. 24:29together and I think it'll be a lot of
  644. 24:30fun. But there are many other unsung
  645. 24:33heroes in the community for AMD building
  646. 24:36amazing things with Stricks Halo. Um
  647. 24:39Reddit sentiment tracks this perfectly.
  648. 24:42The happy users are running lemonade and
  649. 24:44llama.cpp LM Studio and models like
  650. 24:48Quinn or Gamma Sepron light llm also uh
  651. 24:51open web UI coding agents. I'm running
  652. 24:54quus which is a lot of fun. Uh multiple
  653. 24:56models at once there. They're they're
  654. 24:58they're creating a local box that is
  655. 25:00quiet, low power, private, always
  656. 25:02available and works well. The frustrated
  657. 25:04users are usually stuck on one of three
  658. 25:07things. Rockom edge cases image
  659. 25:10generation performance maybe memory
  660. 25:12configuration confusion things like the
  661. 25:14the try map setting in LM Studio has
  662. 25:17tripped up a lot of folks in our
  663. 25:19community especially if they you know
  664. 25:20read an older how-to it's like oh let's
  665. 25:22let's create 64 gigs for the GPU and 64
  666. 25:25gigs for the host and then LM Studio
  667. 25:27needs slightly more than 64 gigs and it
  668. 25:29trips over the try memory map because
  669. 25:31it's trying to load it twice it's trying
  670. 25:33to load it to system memory and then
  671. 25:34copy it over to GPU memory but that's
  672. 25:36not that's not how that should work at
  673. 25:38If you are going to go off script and
  674. 25:39use rockom, I would recommend the rock
  675. 25:43version of the rockom if you're not
  676. 25:44going to use one of Donado's toolboxes.
  677. 25:46Uh some people report excellent LLM
  678. 25:48performance but then miserable diffusion
  679. 25:50behavior and then they spend a lot of
  680. 25:51time troubleshooting that. Windows users
  681. 25:53in particular get stuck because of the
  682. 25:54whole shared memory thing and Windows
  683. 25:57reports things in a way that makes
  684. 25:59shared me. But that's going to be fixed
  685. 26:00with the updated version of Windows
  686. 26:02that's coming soon. Oh, I also did a
  687. 26:04clone. I tried to clone this image onto
  688. 26:06our other stricks halo machines like GMK
  689. 26:08tech ministorm also the minisformm NAS
  690. 26:11that I reviewed recently check that out
  691. 26:13and our framework desktop now framework
  692. 26:15desktop wouldn't boot because it ran
  693. 26:17into a systemd issue that is a system
  694. 26:19debug but generally most things work
  695. 26:21better than I expected the framework
  696. 26:23desktop tripping over the system debug
  697. 26:25you don't have to use systemd you can
  698. 26:27use something else mostly you don't have
  699. 26:28to though I'm sort of expecting that if
  700. 26:30AMD is going to say here is the formula
  701. 26:32you can do this that they might also
  702. 26:34will provide some hints around
  703. 26:35configuring DRO and kernel versions
  704. 26:36because that can also be a source of
  705. 26:39chaos and headache. Uh but no, it's not
  706. 26:42ready yet. When I move this over, it it
  707. 26:44depends on the UU IDs of the partition
  708. 26:47not changing. So when I got an update
  709. 26:48and it applied the update, it assumed
  710. 26:50that the UU ID partition the partition
  711. 26:52UU IDs had not changed, which was not
  712. 26:54true. Uh there is one concerning issue
  713. 26:55on GitHub. Well, there's a couple, but
  714. 26:576182 is a generalized stricks Halo
  715. 27:02thing. Users are hitting non-reoverable
  716. 27:04HSA memory faults when trying to load
  717. 27:06PyTorch or HIP models on some strict
  718. 27:08Halo machines. The issue report goes
  719. 27:10pretty deep. Multiple kernels, multiple
  720. 27:12rock versions, different BIOS CMA
  721. 27:14settings and AMD's uh official rock and
  722. 27:16pietorrch container and also the rock
  723. 27:18builds kernel parameters and and you
  724. 27:21know it's the same class of crash
  725. 27:23happening. This is also why the
  726. 27:24community deserves credit. things like
  727. 27:28Donado stricks, Halo Toolboxes, and the
  728. 27:30Reddit Discord forum crowd got a lot of
  729. 27:32this working before AMD's official story
  730. 27:34caught up. We have um we've all proved
  731. 27:37that the platform was worth caring
  732. 27:40about. Now, that rocking bug doesn't
  733. 27:42happen on this platform and it doesn't
  734. 27:45happen on the machines that I have as
  735. 27:47configured or configured really
  736. 27:49similarly to the way that the the AI
  737. 27:52Halo is. But there are obviously a lot
  738. 27:54of people that it is affecting even
  739. 27:55though the software is the same. It may
  740. 27:57be like our framework desktop where it
  741. 27:59is a systemd bug but the systemd bug is
  742. 28:02triggered by something in the BIOS and
  743. 28:05it's fixed in a newer version of systemd
  744. 28:07so it's just not not yet on the the
  745. 28:11platform here. So what's not to like?
  746. 28:13Nothing to do with stricks AI halo. Uh
  747. 28:15but the discrete playbooks I can't help
  748. 28:17but notice on the developer website it
  749. 28:19says coming soon. I mean, that's just
  750. 28:22too on brand for AMD. That that's kind
  751. 28:24of the microcosm. AMD gets the hardware
  752. 28:26right and it's amazing. The community is
  753. 28:28excited and then the software story
  754. 28:30arrives late. Uh, and also with
  755. 28:32footnotes. By the time AMD has polished
  756. 28:35the appliance, the ecosystem already
  757. 28:37has, you know, other options. You It's
  758. 28:40like, oh, we'll just get it working with
  759. 28:41Vulcan or, oh, something else will get
  760. 28:44put together. And then the community
  761. 28:45builds a tool chain and then before you
  762. 28:47know it, AMD has
  763. 28:49sort of working to do something that is
  764. 28:53parallel to but maybe not quite in line
  765. 28:55with what the community is already
  766. 28:56doing. But also Gorgon Halo is really
  767. 28:58close. So if you're a value buyer, maybe
  768. 29:02buying smart is waiting. I mean Gorgon
  769. 29:04Halo is going to be faster memory and up
  770. 29:06to 192 GB of unified memory for bigger
  771. 29:09bigger models in local AI. Now, that
  772. 29:11might matter, but I think the
  773. 29:12performance to VRAM ratio is already a
  774. 29:14bit out of whack at 128 GB with these
  775. 29:17crazy inflated prices. I mean, I'll take
  776. 29:19more memory if I can get it, but I mean,
  777. 29:21the bang for buck buyer may not want to
  778. 29:22spend $4,000 on a late cycle Stricks
  779. 29:24Halo box when cheaper Stricks Halo
  780. 29:27systems exist. And the Gorgon Halo
  781. 29:30refresh is visible on the horizon. Q3 is
  782. 29:32just right around the corner. But I
  783. 29:34think AMD sort of expects that. I don't
  784. 29:35think they really expect a lot of people
  785. 29:36to buy this particular aluminum box.
  786. 29:39Aluminium. AMD can ship a more cohesive
  787. 29:43local AI ecosystem experience. That's
  788. 29:45the story here. And you can be up and
  789. 29:48running on any of these platforms with
  790. 29:49LM Studio pretty much immediately and
  791. 29:52have a lot of fun. And it's amazing. And
  792. 29:54this website, the software model, and
  793. 29:56AMD making cohesive efforts to let you
  794. 29:59have this kind of fun and productivity
  795. 30:01with your six Halo box. That's what we
  796. 30:03need more of. This product is a Halo
  797. 30:05product in another sense. Not because
  798. 30:07AMD expects to sell millions, as I said,
  799. 30:09I really don't think they do, but
  800. 30:11because it lights the path. It is
  801. 30:13literally a pipe cleaner for the next
  802. 30:15generation of local AI machines on the
  803. 30:18software and hardware side, I guess, is
  804. 30:20what I'm saying. Developer workflows and
  805. 30:23working out the developer workflows are
  806. 30:25very, very important. And it's probably
  807. 30:27a good idea that that's not, you know,
  808. 30:30the un untapped masses of the internet.
  809. 30:33I think for where we are in 2026, it
  810. 30:35reminds me of what I've read about the
  811. 30:37early 1980s. Back then, everybody had
  812. 30:39heard of computers, but few people could
  813. 30:41imagine what having one at home would
  814. 30:44enable. And a lot of the same societal
  815. 30:46concerns now rhyme with what people were
  816. 30:49worried about then. Uh back then, I
  817. 30:50mean, a lot of people were concerned if
  818. 30:52the computer would replace their job in
  819. 30:54the 1980s.
  820. 30:56Sounds a lot like today. And these boxes
  821. 30:58are a glimpse of the alternative.
  822. 31:00powerful local models, private data,
  823. 31:03predictable costs, no token tax and no
  824. 31:05tools that you know live elsewhere. The
  825. 31:07tool lives on your desk. I think the
  826. 31:09future gets much more interesting when
  827. 31:11the machine that is thinking with you is
  828. 31:15yours and lives on your desk and I think
  829. 31:19it opens up a lot of possibilities which
  830. 31:21I'm very excited about for the future.
  831. 31:23I'm this level one. It is fantastic to
  832. 31:25see this kind of a launch. If I missed
  833. 31:26anything or you want to chat or run some
  834. 31:28tests or build some things or generate
  835. 31:30some images, uh, let me know. Hit me up
  836. 31:32at the forum. I'm signing out and I'll
  837. 31:33see you there. Eating cereal while
  838. 31:35driving a car.
  839. 31:38I'll take it. [music]
  840. 31:44[music]

About this transcript

This page contains the full transcript of Is STRIX Better than SPARK? Now Launching w/new Software: AMD's Ryzen AI Halo Developer Workstation by Level1Techs, generated from the public captions YouTube serves with the video. The transcript has 5,913 words across 840 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.