Is STRIX Better than SPARK? Now Launching w/new Software: AMD's Ryzen AI Halo Developer Workstation — Transcript
Full transcript
- 0:00There it is. Ryzen AI Halo. AMD built
- 0:04Stricks Halo, the chipset on which this
- 0:05is based, as a workstation APU. We first
- 0:08saw it something like 18 months ago. We
- 0:12crazy internet randos looked at it then
- 0:14and we immediately saw the 128 gigs of
- 0:16memory and we said, "Cool. What if we
- 0:19use that to do crimes against model
- 0:21size?" [music]
- 0:28AMD was really early here. That was even
- 0:30before Nvidia launched their 128 gig DJX
- 0:33Spark. And weirdly, Stricks Halo worked.
- 0:35It definitely worked for AMD. I mean, 48
- 0:38gig coding models at roughly 60 tokens
- 0:41per second. A 92 gig mixture of experts
- 0:43model roughly 20 tokens per second.
- 0:45Local, quiet, on a desk, less than 200
- 0:49watts. So, why is this product coming
- 0:50out now? Well, there there were some
- 0:53problems with the software side of
- 0:55things. It was never the silicon. The
- 0:57problem was always everything around it.
- 0:59Rockom versions and Rockom modifications
- 1:02for Stricks Halo, memory allocation
- 1:04myths, BIOS settings, container bugs,
- 1:06Vulcan escape hatches, and enough GitHub
- 1:09issues to make you reconsider your life
- 1:11choices. At least the life choices that
- 1:14led to working on AI stuff in the first
- 1:16place. But this this is the story. This
- 1:20is the the the pin in the corkboard, if
- 1:22you will, for AMD having hit both
- 1:25software and hardware milestones. It's
- 1:28not just a known good stricks Halo
- 1:30system with Linux in my case, also
- 1:32available with Windows. Playbooks and
- 1:33firmware updates piped right into the
- 1:35Linux way. LM Studio support, that's
- 1:37what I'm running here. Comfy UI, which
- 1:39is in the background, that's not a
- 1:41problem. This is the standard bearer for
- 1:43what comes next. It points toward Gorgon
- 1:46Halo, a larger unified memory machine,
- 1:48192 gigs coming with that update in Q3,
- 1:51but also the rest of the Stricks Halo
- 1:54ecosystem is going to be pulled up
- 1:55behind this, which is great. Fixed
- 1:58playbooks, firmware updates, clarified
- 2:01memory behavior, known good rock stacks,
- 2:04and known good Vulcan stacks, known good
- 2:07playbooks. And that's going to elevate
- 2:08everything from Minis Forum and
- 2:11Framework Desktop and GMK Techch and all
- 2:13the other versions of Stricks Halo out
- 2:15there. And that is what we can rally
- 2:18around. This is the real milestone. It's
- 2:21not just a flagship box, but a reference
- 2:23point the whole platform can finally
- 2:26rally around. A Halo's Halo product, if
- 2:31you will. This is AMD's Ryzen AI Halo
- 2:33developer platform. And yes, the name is
- 2:35funny because it's Stricks Halo. It's
- 2:37already a Halo and this is already a
- 2:38Halo product. This is a Halo Halo
- 2:41product. Somewhere there's a product
- 2:43manager that's glowing. The hardware is
- 2:45familiar. Ryzen AI Max 395 plus 16 Zen 5
- 2:49cores. Radeon 860S graphics. 128 gigs of
- 2:51unified LPDDR5X.
- 2:53A compact attractive chassis available
- 2:56either with Linux or Windows. The price
- 2:58$4,000. Uh that price is maybe the first
- 3:01problem, but we need to chat about it.
- 3:04But more to the point, Framework, Minis
- 3:06Forum, GMK Techch, Beink, Corsair, HP,
- 3:08and others have already shipped Stricks
- 3:10Halo based systems, and some are much
- 3:12cheaper, some are much more expandable.
- 3:15I've got a 25 gig Ethernet card in here.
- 3:17The Framework Desktop also has a PCIe
- 3:19slot, but in their case, it's not
- 3:20accessible. But if you put the
- 3:22motherboard in a different case, it's
- 3:23totally okay. GMK Tech gives you an
- 3:26Oculink slot. You've got options. Almost
- 3:28all of these have multiple M.2 slots,
- 3:30except for the one from AMD. It's got a
- 3:322280 at least. So, I mean, I'll take it.
- 3:35But, yeah, it's a lot of fun. Oh, and
- 3:38the fact that I've got a 25 gig dual
- 3:40port Intel E810 Nick in the minis forum,
- 3:42I think that does change the
- 3:43conversation a little bit around
- 3:45clustering. I mean, if you look at the
- 3:48networking here, you've just got the one
- 3:50realtech 10 gig nick. And that's all
- 3:53that is in our AI Halo system. But you
- 3:55can still cluster it. You can still
- 3:57build some clustering stuff. In fact, I
- 3:59think you should check out Donado's
- 4:00video on this. He has been building a
- 4:03Stricks Halo cluster. He's the Stricks
- 4:05Halo toolboxes guy and cluster like this
- 4:08out of that. Yeah. Yeah, it works. It
- 4:09was a latency problem the whole time.
- 4:10But back to this. Why does this exist?
- 4:13Why what is AMD doing here? Cuz it's not
- 4:15remotely about Stricks Halo, at least
- 4:17not this late in the product cycle. I
- 4:19think this is AMD planting a flag and
- 4:22paving the way for Gorgon Halo, the next
- 4:25generation product that rhymes with this
- 4:26one. Probably just a chip swap and
- 4:29memory swap as well. That's right around
- 4:30the corner. Q3. So AMD says. So we have
- 4:34at the back, if we take a look at the
- 4:35rear IO, USBC power from the 240 watt
- 4:38power brick, 10 gigabit Ethernet, thanks
- 4:40to our realtech chipset, three other
- 4:42USBC ports. One is recommended for
- 4:44display port alt mode, and two others
- 4:46that support alt mode and USBC
- 4:48connectivity. There is no other
- 4:49high-speed networking, Oculink or even a
- 4:53second M.2 internally. It's just the
- 4:56one. At least I don't think I'm pretty
- 4:57sure there's not cuz I took it apart. we
- 4:59can all look together. Taking apart is
- 5:02easy and and and repeatable though.
- 5:03There are screws underneath the magnetic
- 5:05feet. Uh there's also a Kensington lock
- 5:07port, so that's appropriate for
- 5:09university labs and that sort of thing
- 5:11where maybe you might have one of these
- 5:12and it would walk off. But physically,
- 5:15this is uh one of the most liipian
- 5:18options for Stricks Halo. Now, the
- 5:20obvious comparison that I've alluded to,
- 5:21Nvidia DJX Spark. Both are small local
- 5:23AI workstations. Both have 128 gigs of
- 5:25unified memory. Both target inference
- 5:28and developer workflows. Both claim
- 5:31support for very large models under the
- 5:34right quantization and optimization
- 5:36parameters. Right now on Newegg, you can
- 5:40still buy Spark for about 4,000 to 4500.
- 5:43DGX Spark has the stronger platform
- 5:45story where Nvidia tends to always come
- 5:48out ahead. I mean CUDA, DGXOS, container
- 5:50workflows, and the scale out networking.
- 5:53The scaleout networking really is
- 5:54something special. Uh the networking
- 5:57difference matters. I think Spark's
- 5:58ConnectX path is a real pro feature and
- 6:00not a spec sheet decoration. If you are
- 6:02trying to stitch two boxes together,
- 6:05Nvidia starts ahead. And the skills that
- 6:08you gain from dealing with nickel NCCL,
- 6:10that's Nvidia's communication primitives
- 6:12library. The skills from that will
- 6:14transfer to multi-million dollar AI
- 6:16clusters. So you you learn and you do
- 6:18stuff. AMD's counterpunch is flexibility
- 6:21and at least historically price. This is
- 6:24x86. It runs Linux. It has also always
- 6:28run Windows, whereas Nvidia is planning
- 6:31Windows support for their stuff, but
- 6:33it's not here yet. And AMD uses open-ish
- 6:36tooling. It does not require you to buy
- 6:38into the CUDA world view. And you know,
- 6:41for many inference workloads, especially
- 6:43token generation, stricks Halo is
- 6:46basically the same performance as DJX
- 6:47Spark. And that's because both systems
- 6:50are ultimately gated by the unified
- 6:53memory bandwidth and it's basically the
- 6:55same. It's very similar between the two
- 6:56platforms. The short version for view
- 6:58like anybody that doesn't have the
- 7:00attention span, Spark is cleaner for
- 7:03Nvidia's approach, but the Ryzen AI Halo
- 7:05is more flexible workstation. Spark has
- 7:07the better networking story. Halo has
- 7:09the better I could just use this like a
- 7:11PC story. Oh, and I'm not sleeping on
- 7:14networking over USB 4. That is an option
- 7:16here. here. You can use those USBC ports
- 7:18to connect to other machines, kind of
- 7:19like a Mac, but I want RDMA to be a part
- 7:22of that story when we talk about it
- 7:23because RDMMA is not here yet for that.
- 7:25Um, you can do it over USB 4,
- 7:28Thunderbolt, and it's on the order of
- 7:3020ish gigabit, give or take, but uh,
- 7:33yeah, I mean, I don't know. Nvidia made
- 7:36some weird choices with Spark 2, like
- 7:37they didn't use a 2280 M.2. 2280,
- 7:40everything on this table 2280, which is
- 7:42the only sane option. Why would they do
- 7:45that? unless they just wanted to put a
- 7:47crappy, you know, it's like, oh, we
- 7:48can't just buy a 2280. So, yeah, because
- 7:51[snorts] this is a pre-existing
- 7:52platform, Stricks Halo, uh, and because
- 7:56it's from AMD, I got to go off on a
- 7:58tangent for a second. There's a lot of
- 8:01bad information online about Stricks
- 8:02Halo, especially around memory and how
- 8:05the memory works in a unified platform.
- 8:06People say there's a difference between
- 8:08what Nvidia is doing with ARM and
- 8:10Stricks Halo. That is not true. it it's
- 8:12it's they they treat it like as if the
- 8:13machine has two physical memory pools.
- 8:15It does not. On Linux, the important
- 8:18mechanism that the GPU uses to map
- 8:20through the shared memory pool is
- 8:23through GTT and TTM. For most AI
- 8:25workloads, you do not need to reserve 96
- 8:28or 112 GB of memory permanently for the
- 8:30GPU. When I was doing gaming reviews on
- 8:33laptops that were based on stricks Halo,
- 8:36really killer machines like the ROG Flow
- 8:38Z13 and the HP G1A, I showed games like
- 8:42Clar Obscure running at a high frame
- 8:44rate but with only 512 megabytes of
- 8:47reserved VRAM. Except in some really
- 8:50legacy software scenarios, it's often
- 8:52the case that for AI workloads, this is
- 8:54the better option as well. keep the BIOS
- 8:57reserved GPU memory very small and then
- 8:59let the Linux kernel driver expose the
- 9:01large shared pool dynamically. The the
- 9:04old mental model for this and and why it
- 9:07shows up wrong a lot of the time I think
- 9:08comes from Windows. On Windows, the
- 9:11split is more visible and more annoying.
- 9:13AMD's variable graphics memory exists
- 9:16because software still asks things like
- 9:18how much VRAM do we have and it'll freak
- 9:20out if it encounters 512 megabytes. But
- 9:22most software has been updated to be
- 9:24aware of the thing that exists with
- 9:26unified memory platforms. This is a
- 9:28major problem for Nvidia too with their
- 9:30planned RTX Spark, the stuff that we saw
- 9:31at Compyex just just a few weeks ago
- 9:34because you know Windows and so of
- 9:36course I checked how they were handling
- 9:39Windows on the machines that I saw there
- 9:41and guess what Windows got fixed and the
- 9:43unified memory fix there benefits AMD
- 9:46too. That's sort of the funny part here.
- 9:48The work that Microsoft has done
- 9:50apparently fasttracked because now
- 9:53Nvidia needs it as well ends up greatly
- 9:55helping AMD's Windows Stricks Halo
- 9:58experience. I'm not got anything Windows
- 10:00related in this video. Probably a
- 10:02follow-up video if I can get my hands on
- 10:03the software image. But Stricks Halo in
- 10:05a nutshell, it is not a 96 gig GPU. It's
- 10:08128 gig unified machine. It can address
- 10:10a large pool of memory. And when the OS
- 10:12and the driver stack are configured
- 10:13correctly, I've got up to 112 GB of
- 10:16memory on this platform for the GPU
- 10:18without any issues or headaches or any
- 10:19really like special hackery that I had
- 10:21to do. It just works. So the people that
- 10:23complain about memory allocation or like
- 10:26freezing or hanging or something like
- 10:28that, they're they're not facing
- 10:30problems with the memory allocation.
- 10:32It's probably something else because
- 10:33it's it is a a wellexecuted unified
- 10:36platform. Sometimes an LM Studio just
- 10:39toggling off try MAPAP is all you need
- 10:42to do. That's the thing that hang people
- 10:43get hung up on sometimes. But I digress.
- 10:46So now AMD has this playbooks. This is
- 10:50the thing that I literally told Anush at
- 10:53Advancing AI. AMD needs a website of
- 10:55known good recipes that are part of your
- 10:58CI/CD process. Every time AMD pushes an
- 11:01update, every run through, every little
- 11:03thing, your CI/CD system kicks off and
- 11:06runs through all of these playbooks to
- 11:09make sure that it all works. Andy should
- 11:11have had this on launch day. The reason
- 11:13they did not is probably simple. Stricks
- 11:16Halo was not originally positioned as an
- 11:18AI developer appliance. These are useful
- 11:21playbooks for the individual as well
- 11:23because they run like these are common
- 11:25things that you would want to do. The
- 11:27community realized early on that strikes
- 11:29Halo was was going to be good, too good
- 11:31at inference to ignore. The playbook's
- 11:32turn, I spent all weekend trying to weed
- 11:35out if it runs better with Rockim or
- 11:36Vulcan into open the developer center,
- 11:40find the thing that's close to what you
- 11:41want to do, and run that. And out of the
- 11:44box, shipped with this thing, over 380
- 11:46gigs of stuff. It's got lemonade and web
- 11:48UI and VS Code. Well, you got to
- 11:50download VS Code, AMD Sync Workflows.
- 11:52There's a whole bunch of stuff on here
- 11:54that is ready to go out of the box in
- 11:56both in terms of both models and
- 11:58configuration. Lemonade server from AMD
- 12:00as I mentioned and comfy UI. This is
- 12:02this is amazing. This is AMD putting the
- 12:04most effort I have ever seen from AMD
- 12:06into making the client side experience
- 12:09smooth. And that is the real review.
- 12:11It's not really this box. It's the
- 12:12software. It's the mindset because all
- 12:14of this executed properly benefits all
- 12:18of this. That's what I meant when I said
- 12:19it's going to bring up everything. Yes,
- 12:21you could run 122 billion parameter uh
- 12:23class model on this stricks halo. That
- 12:26is an amazing sentence to be able to to
- 12:28utter in 2026, but I do not think the
- 12:31best daily driver layout is necessarily
- 12:34one enormous model eating the whole
- 12:37machine. The better version of this
- 12:38platform is probably a competent 27 to
- 12:4230 billion parameter coding or reasoning
- 12:44model as kind of the main brain with a
- 12:47constellation of smaller models around
- 12:49it. Uh maybe one for routing, maybe one
- 12:52for summarization, maybe one for
- 12:53retrieval cleanup. Retrieval augmented
- 12:55generation is a fun thing to explore,
- 12:56maybe one for vision or OCR, fast
- 12:59drafting, maybe one for tool use
- 13:01verification. Smaller models do prompt
- 13:04processing faster, respond sooner, and
- 13:06leave room for deep context and parallel
- 13:08work and make some of the slower running
- 13:10models a little bit more tolerable. How
- 13:12big a model can you cram into memory is
- 13:14not really that awesome. It's how many
- 13:17useful local intelligences can you keep
- 13:18alive at once I think okay so real world
- 13:22performance yes you can run 120 billion
- 13:24parameter model on this but that's not
- 13:26really the point and actually a dense
- 13:28120 billion parameter model uh would be
- 13:30difficult understand if you're just
- 13:32getting into this the difference between
- 13:33mixture of experts and a dense model and
- 13:36I'll use Quinn as an example there's a
- 13:3735 billion parameter Quinn model and
- 13:39there's a 27 billion parameter Quinn
- 13:41model the 27 billion parameter Quinn
- 13:43model on this is going to run 12 to 16
- 13:47to 20 tokens per second depending on
- 13:49some parameters and that's at a Q8 it's
- 13:528 bits per weight quantization you can
- 13:55run four and six bit but I think with
- 13:57Quinn you start to lose some of the
- 13:59fidelity of the model when you're much
- 14:01below Q8 it's actually uh natively 16
- 14:05bit so a 27 billion parameter dense
- 14:08model would be double that you need like
- 14:1048 gigs and yes you can run that on this
- 14:12but you're going to be in a relatively
- 14:15low uh token throughput situation.
- 14:18Mixture of experts, it's like Quinn 35B.
- 14:22It's a 35 billion parameter model, but
- 14:24that A3B means that only three billion
- 14:26parameters are active at once. Those
- 14:28kinds of models, mixture of experts
- 14:29models are great for systems like this
- 14:32that are memory bandwidth constrained
- 14:34because only three billion parameters
- 14:36are active at any one time. And the
- 14:38quality output between the 35 billion
- 14:41and 27 billion parameter models like how
- 14:43how good they are like how good they are
- 14:44at doing tasks is pretty similar. The
- 14:47setup that I have here is with Turnstone
- 14:50and I'm I'm doing something really
- 14:51interesting. Remember the 4GPU R9700
- 14:54system that I'm running that also has
- 14:56128 gigs of memory across four GPUs
- 14:58that's in the basement. This system is
- 15:01controlling it. This is an orchestrator
- 15:04and this is toward that whole agentic
- 15:07thing that everybody talks about. So
- 15:09Turnstone is a piece of software that's
- 15:11available on GitHub and it's open source
- 15:13and it is amazing and I need to do
- 15:15another video on it because it's kind of
- 15:17like Hermes or Open Claw but it's you
- 15:19know enterprise mindset. It does some
- 15:20things a little differently and again
- 15:23separate conversation but turnstone
- 15:26running on this with a you know Quinn 35
- 15:30billion parameter model is smart enough
- 15:32to be a supervisor. Now that is nowhere
- 15:34near as smart as the best models that
- 15:37you can get on you know Open Router or
- 15:40OpenAI or Anthropic. Those are frontier
- 15:43models. uh but you can connect those
- 15:45models to Turnstone through an API but
- 15:48you can also use this to minimize the
- 15:50amount of tokens that you spend in the
- 15:52cloud. You see this as an orchestrator
- 15:54doesn't have to be the smartest model to
- 15:57for you to get utility out of it. Um and
- 16:00with a 35 an A35B type uh model with
- 16:05multi-token prediction that's another
- 16:07thing that you probably want to try to
- 16:08do. It means that you try to get more
- 16:10than one token per uh prediction and the
- 16:12acceptance rate on that is typically 60
- 16:1470 80%. That'll improve your throughput
- 16:17in [snorts] addition to the whole only
- 16:19three billion parameters active. A lot
- 16:21of words to say that you can get 30 40
- 16:2350 60 tokens per second per session and
- 16:27three four sessions running in parallel.
- 16:29And so like with Turnstone, when you run
- 16:31an orchestrator like this and you give
- 16:33it a task, it can split up the task into
- 16:35four or five pieces and work on them
- 16:37simultaneously. And each piece that it's
- 16:39working on doesn't have to run with the
- 16:41same model. So like you don't have to be
- 16:42limited to Quinn. It's like, oh, the 35
- 16:45billion parameter that's going to fit
- 16:46in, you know, 48 gigs of VRAM. What do I
- 16:48do with the rest of the VRAM? You can
- 16:49run more models. You can run a model
- 16:51that's really good at OCR, really small
- 16:53model. Um, GMA from Google is also
- 16:56really good. They have some small models
- 16:58that are 3 4 billion parameters that are
- 17:01fantastic. Um, one of the reasons that
- 17:03Quinn is so amazing is because they
- 17:05publish papers and make it really easy
- 17:07for you to do your own fine-tuning. And
- 17:09so even though those those models are
- 17:11smaller, they're also larger. And there
- 17:13are smaller Quinn models that you can
- 17:15fine-tune on a platform like this. Now,
- 17:17it'll take a couple of weeks. like it'll
- 17:19run for a long time but you can set that
- 17:21up on a platform like this and then you
- 17:24have customized the model and trained it
- 17:27on your data. The difference between
- 17:29retrieval augmented generation and uh
- 17:32you know fine-tuning is that you have a
- 17:34lot of data that doesn't really change
- 17:35and you want the model to be read in on
- 17:38the data that you wanted to have it
- 17:39access to fine-tuning is probably worth
- 17:41it. But retrieval augmented generation
- 17:44is you sort of frontload a lot of data
- 17:47into the model with the requests. And so
- 17:50retrieval augmented generation you have
- 17:51like a vector database or something like
- 17:53that. You take information and you do
- 17:55some processing to it and at runtime it
- 17:59is injected into the model in a
- 18:01different way than fine-tuning does sort
- 18:03of kind of and so with every request the
- 18:06model has the context of the the
- 18:08documents that you've added and so
- 18:10that's why I say retrieval augmented
- 18:11generation it's generating tokens
- 18:13augmented by the fact that it retrieved
- 18:16the contents of the documents in a
- 18:18vector format or something like that not
- 18:20the raw format you got to process the
- 18:22documents that you're feeding for rag.
- 18:23Whereas fine-tuning, you run this
- 18:25process on the model, the model is
- 18:28modified, and then when you run the
- 18:30model, it has information about the
- 18:33documents that you gave it for the
- 18:35fine-tuning process. Both of those
- 18:37things are achievable on this platform
- 18:39because of good documentation that goes
- 18:41along with the model. And Quinn has been
- 18:44a standout among all of the options
- 18:46available in terms of how well it is
- 18:48documented. And it is, this platform is
- 18:50fast enough for you to to do those kinds
- 18:52of things both academically and in an
- 18:54actually useful kind of way. Now, for
- 18:56like software development and things
- 18:57like that, there's a lot of really
- 18:59low-level bog standard like janitorial
- 19:01type things like I need help organizing
- 19:03this git repo. I need help writing a
- 19:05unit test. I need help, you know, doing
- 19:06this evaluation. I need you to run
- 19:08through the unit tests or help me figure
- 19:10out why this one is failing. You don't
- 19:11need a frontier model for a lot of those
- 19:13kinds of really low-level tasks. And you
- 19:16can run that on here with Turnstone or
- 19:18any other harness that you want and that
- 19:20works fine and it runs acceptably fast
- 19:22in my opinion. The utility of these
- 19:25kinds of models and the state-of-the-art
- 19:27like how the software has moved the
- 19:29hardware over the years is a big part of
- 19:31the story here. The the software that's
- 19:33available now in mid to late 2026 is
- 19:37light years ahead of where we were a
- 19:38year ago, even though the hardware
- 19:40hasn't really changed a year ago. So it
- 19:41does feel like a new platform and it is
- 19:43really nice to see that level of
- 19:45development and speed. So the big thing
- 19:47with Agentic and benchmarking in a
- 19:49platform like this is don't think of AI
- 19:50as monolithic. Don't sit down. It's like
- 19:52we're going to oneshot a Tetris clone.
- 19:54If I ask this thing to do a Pac-Man
- 19:57clone or something like that, it's going
- 19:59to spawn a bunch of tasks. It's going to
- 20:01put a plan together and then there's
- 20:03going to be a bunch of agents that are
- 20:04running maybe on the same model, maybe
- 20:06on different models in order to complete
- 20:07the tasks and that kind of thing can
- 20:09function as a benchmark. But know that
- 20:11because of the memory bandwidth here,
- 20:12the performance for most models is on
- 20:14par with Nvidia's DGX Spark now in 2026.
- 20:18Now, I mentioned some off label use
- 20:20cases. All about the off label use
- 20:23cases.
- 20:24Thunderbolt, it's not really
- 20:26Thunderbolt, it's USB 4. Two of the
- 20:28three ports are capable of USB4 20 GB.
- 20:33So, here you can see we've got our
- 20:34external Razer uh X2 Thunderbolt 5 dock
- 20:39connected and it's got a big honken GPU
- 20:42in it and we've got 20 Gbit both ways.
- 20:44So, 40 Gbit and their 40 GB interface
- 20:47and we can see that our GPU came up uh
- 20:51properly. And this is a big deal because
- 20:53when Stricks Halo first launched, they
- 20:55didn't really plan on having GPUs. Like
- 20:58even the folks like me that took our
- 21:00framework desktop. This is one of the
- 21:02first stricts halo machines that I got
- 21:04and extended out the X4 slot. This X4
- 21:07slot can only provide 25 watts. You need
- 21:0975 watts for a GPU. So you got to do
- 21:11some herculean things to even get a GPU
- 21:14plugged into that PCIe slot which just
- 21:16four lanes. Wouldn't post because the
- 21:19BIOS wasn't plumbed to do that. It's a
- 21:21firmware. It's a firm AMD firmware team.
- 21:23Like it's like they're so isolated. They
- 21:26got to they got to come into the fold.
- 21:28We got to all work with the AMD firmware
- 21:30team. Um, and so this is a fix. This is
- 21:32this is a fix since launch, which is
- 21:34great. It's great to see this. And this
- 21:35worked well with uh some other different
- 21:37configurations, even running two of
- 21:38these things also for Thunderbolt
- 21:40networking. Now, you can do Thunderbolt
- 21:42networking and it is relatively low
- 21:44latency, but you really need um RDMA,
- 21:48not RDMA, RDMA, a remote direct memory
- 21:51access. Remote director memory access
- 21:53gives us would theoretically give us 20
- 21:55gigabit very low latency connection not
- 21:58through a real tech interface. It's not
- 22:00as good as you know something that's 100
- 22:02200 gigabit but it'll get the job done.
- 22:05I really think that AMD could have put
- 22:09something based on Polar or maybe a
- 22:12little FPGA in there. I think one of the
- 22:14like a competitive differentiation thing
- 22:16that AMD could have done is give you
- 22:17another M.2 too. And maybe with like a
- 22:20an FPGA that would give you maybe
- 22:23something you could experiment with in
- 22:24terms of oh, we can just do our own
- 22:26custom fabric. Maybe even direct PCIe
- 22:28connection because the PCIe it's not
- 22:30really that much of a lift to provide a
- 22:33software layer to do a direct PCIe to
- 22:36PCIe connection through, you know,
- 22:38Oculink or something like that with a
- 22:40little bit of custom glue. But that's
- 22:43not really the thing for this video. But
- 22:46this is really exciting that you have an
- 22:47interface. You can also kind of build an
- 22:48unholy machine here because okay, I've
- 22:50got 48 gigs of really high-speed GPU
- 22:53VRAM plus 128 gigs for mixture of
- 22:57experts type models. This is unholy. And
- 23:01we've done videos on that in the past
- 23:03where you're running uh a high
- 23:05performance GPU that has 16 32 48 96
- 23:10gigs of memory plus 128 gigs of much
- 23:13slower LP uh you know our our DDR5 here.
- 23:17Um, but that lets you run much larger
- 23:20mixture of experts models and also let
- 23:22you run smaller mixture of experts
- 23:24models much faster and also a bunch of,
- 23:27you know, agents or a bunch of models
- 23:30running simultaneously on a platform
- 23:32like this imminently doable. Oh, and AMD
- 23:34has found a way to make the NPU useful
- 23:37in this cuz remember it's designed as a
- 23:38laptop. There's an NPU. You can run
- 23:41multimodel scenarios and use the NPU.
- 23:43things like speech recognition and some
- 23:45of the larger models will also work on
- 23:47the NPU at shallow context with
- 23:49reasonable speed. Uh there's some
- 23:51amazing benchmarks of fast flow LM in
- 23:53exactly this scenario. You can run Llama
- 23:563.2 1 billion at over 50 tokens a second
- 23:58below an 8K context and the prefill
- 24:01speed with that is over 2,000 tokens per
- 24:03second. So if you've got a a pretty
- 24:06awesome carefully planned setup, you can
- 24:08do a lot with that. Now before AMD had a
- 24:11polished story, Stricks Halo made their
- 24:13own. Uh Benado that I mentioned, he made
- 24:15the Stricks Halo toolboxes. Those are
- 24:17the obvious example containerized
- 24:19environments for uh LLMs, image
- 24:22generation, fine-tuning, rock, vulcan,
- 24:24llama.cpp. Uh he's top of mind right now
- 24:28because we're working on something
- 24:29together and I think it'll be a lot of
- 24:30fun. But there are many other unsung
- 24:33heroes in the community for AMD building
- 24:36amazing things with Stricks Halo. Um
- 24:39Reddit sentiment tracks this perfectly.
- 24:42The happy users are running lemonade and
- 24:44llama.cpp LM Studio and models like
- 24:48Quinn or Gamma Sepron light llm also uh
- 24:51open web UI coding agents. I'm running
- 24:54quus which is a lot of fun. Uh multiple
- 24:56models at once there. They're they're
- 24:58they're creating a local box that is
- 25:00quiet, low power, private, always
- 25:02available and works well. The frustrated
- 25:04users are usually stuck on one of three
- 25:07things. Rockom edge cases image
- 25:10generation performance maybe memory
- 25:12configuration confusion things like the
- 25:14the try map setting in LM Studio has
- 25:17tripped up a lot of folks in our
- 25:19community especially if they you know
- 25:20read an older how-to it's like oh let's
- 25:22let's create 64 gigs for the GPU and 64
- 25:25gigs for the host and then LM Studio
- 25:27needs slightly more than 64 gigs and it
- 25:29trips over the try memory map because
- 25:31it's trying to load it twice it's trying
- 25:33to load it to system memory and then
- 25:34copy it over to GPU memory but that's
- 25:36not that's not how that should work at
- 25:38If you are going to go off script and
- 25:39use rockom, I would recommend the rock
- 25:43version of the rockom if you're not
- 25:44going to use one of Donado's toolboxes.
- 25:46Uh some people report excellent LLM
- 25:48performance but then miserable diffusion
- 25:50behavior and then they spend a lot of
- 25:51time troubleshooting that. Windows users
- 25:53in particular get stuck because of the
- 25:54whole shared memory thing and Windows
- 25:57reports things in a way that makes
- 25:59shared me. But that's going to be fixed
- 26:00with the updated version of Windows
- 26:02that's coming soon. Oh, I also did a
- 26:04clone. I tried to clone this image onto
- 26:06our other stricks halo machines like GMK
- 26:08tech ministorm also the minisformm NAS
- 26:11that I reviewed recently check that out
- 26:13and our framework desktop now framework
- 26:15desktop wouldn't boot because it ran
- 26:17into a systemd issue that is a system
- 26:19debug but generally most things work
- 26:21better than I expected the framework
- 26:23desktop tripping over the system debug
- 26:25you don't have to use systemd you can
- 26:27use something else mostly you don't have
- 26:28to though I'm sort of expecting that if
- 26:30AMD is going to say here is the formula
- 26:32you can do this that they might also
- 26:34will provide some hints around
- 26:35configuring DRO and kernel versions
- 26:36because that can also be a source of
- 26:39chaos and headache. Uh but no, it's not
- 26:42ready yet. When I move this over, it it
- 26:44depends on the UU IDs of the partition
- 26:47not changing. So when I got an update
- 26:48and it applied the update, it assumed
- 26:50that the UU ID partition the partition
- 26:52UU IDs had not changed, which was not
- 26:54true. Uh there is one concerning issue
- 26:55on GitHub. Well, there's a couple, but
- 26:576182 is a generalized stricks Halo
- 27:02thing. Users are hitting non-reoverable
- 27:04HSA memory faults when trying to load
- 27:06PyTorch or HIP models on some strict
- 27:08Halo machines. The issue report goes
- 27:10pretty deep. Multiple kernels, multiple
- 27:12rock versions, different BIOS CMA
- 27:14settings and AMD's uh official rock and
- 27:16pietorrch container and also the rock
- 27:18builds kernel parameters and and you
- 27:21know it's the same class of crash
- 27:23happening. This is also why the
- 27:24community deserves credit. things like
- 27:28Donado stricks, Halo Toolboxes, and the
- 27:30Reddit Discord forum crowd got a lot of
- 27:32this working before AMD's official story
- 27:34caught up. We have um we've all proved
- 27:37that the platform was worth caring
- 27:40about. Now, that rocking bug doesn't
- 27:42happen on this platform and it doesn't
- 27:45happen on the machines that I have as
- 27:47configured or configured really
- 27:49similarly to the way that the the AI
- 27:52Halo is. But there are obviously a lot
- 27:54of people that it is affecting even
- 27:55though the software is the same. It may
- 27:57be like our framework desktop where it
- 27:59is a systemd bug but the systemd bug is
- 28:02triggered by something in the BIOS and
- 28:05it's fixed in a newer version of systemd
- 28:07so it's just not not yet on the the
- 28:11platform here. So what's not to like?
- 28:13Nothing to do with stricks AI halo. Uh
- 28:15but the discrete playbooks I can't help
- 28:17but notice on the developer website it
- 28:19says coming soon. I mean, that's just
- 28:22too on brand for AMD. That that's kind
- 28:24of the microcosm. AMD gets the hardware
- 28:26right and it's amazing. The community is
- 28:28excited and then the software story
- 28:30arrives late. Uh, and also with
- 28:32footnotes. By the time AMD has polished
- 28:35the appliance, the ecosystem already
- 28:37has, you know, other options. You It's
- 28:40like, oh, we'll just get it working with
- 28:41Vulcan or, oh, something else will get
- 28:44put together. And then the community
- 28:45builds a tool chain and then before you
- 28:47know it, AMD has
- 28:49sort of working to do something that is
- 28:53parallel to but maybe not quite in line
- 28:55with what the community is already
- 28:56doing. But also Gorgon Halo is really
- 28:58close. So if you're a value buyer, maybe
- 29:02buying smart is waiting. I mean Gorgon
- 29:04Halo is going to be faster memory and up
- 29:06to 192 GB of unified memory for bigger
- 29:09bigger models in local AI. Now, that
- 29:11might matter, but I think the
- 29:12performance to VRAM ratio is already a
- 29:14bit out of whack at 128 GB with these
- 29:17crazy inflated prices. I mean, I'll take
- 29:19more memory if I can get it, but I mean,
- 29:21the bang for buck buyer may not want to
- 29:22spend $4,000 on a late cycle Stricks
- 29:24Halo box when cheaper Stricks Halo
- 29:27systems exist. And the Gorgon Halo
- 29:30refresh is visible on the horizon. Q3 is
- 29:32just right around the corner. But I
- 29:34think AMD sort of expects that. I don't
- 29:35think they really expect a lot of people
- 29:36to buy this particular aluminum box.
- 29:39Aluminium. AMD can ship a more cohesive
- 29:43local AI ecosystem experience. That's
- 29:45the story here. And you can be up and
- 29:48running on any of these platforms with
- 29:49LM Studio pretty much immediately and
- 29:52have a lot of fun. And it's amazing. And
- 29:54this website, the software model, and
- 29:56AMD making cohesive efforts to let you
- 29:59have this kind of fun and productivity
- 30:01with your six Halo box. That's what we
- 30:03need more of. This product is a Halo
- 30:05product in another sense. Not because
- 30:07AMD expects to sell millions, as I said,
- 30:09I really don't think they do, but
- 30:11because it lights the path. It is
- 30:13literally a pipe cleaner for the next
- 30:15generation of local AI machines on the
- 30:18software and hardware side, I guess, is
- 30:20what I'm saying. Developer workflows and
- 30:23working out the developer workflows are
- 30:25very, very important. And it's probably
- 30:27a good idea that that's not, you know,
- 30:30the un untapped masses of the internet.
- 30:33I think for where we are in 2026, it
- 30:35reminds me of what I've read about the
- 30:37early 1980s. Back then, everybody had
- 30:39heard of computers, but few people could
- 30:41imagine what having one at home would
- 30:44enable. And a lot of the same societal
- 30:46concerns now rhyme with what people were
- 30:49worried about then. Uh back then, I
- 30:50mean, a lot of people were concerned if
- 30:52the computer would replace their job in
- 30:54the 1980s.
- 30:56Sounds a lot like today. And these boxes
- 30:58are a glimpse of the alternative.
- 31:00powerful local models, private data,
- 31:03predictable costs, no token tax and no
- 31:05tools that you know live elsewhere. The
- 31:07tool lives on your desk. I think the
- 31:09future gets much more interesting when
- 31:11the machine that is thinking with you is
- 31:15yours and lives on your desk and I think
- 31:19it opens up a lot of possibilities which
- 31:21I'm very excited about for the future.
- 31:23I'm this level one. It is fantastic to
- 31:25see this kind of a launch. If I missed
- 31:26anything or you want to chat or run some
- 31:28tests or build some things or generate
- 31:30some images, uh, let me know. Hit me up
- 31:32at the forum. I'm signing out and I'll
- 31:33see you there. Eating cereal while
- 31:35driving a car.
- 31:38I'll take it. [music]
- 31:44[music]
About this transcript
This page contains the full transcript of Is STRIX Better than SPARK? Now Launching w/new Software: AMD's Ryzen AI Halo Developer Workstation by Level1Techs, generated from the public captions YouTube serves with the video. The transcript has 5,913 words across 840 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.