YouTube2Text

M5 Ultra vs 2 DGX Sparks… The Number You're Not Looking At — Transcript

by Alex Ziskind · 3,740 words · 533 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Are you thinking what I'm thinking?
  2. 0:01These edges are way too sharp. You
  3. 0:03should not let your kids play with them.
  4. 0:04It even says so on the box.
  5. 0:06>> I've got the M5 Ultra Mac Studio here,
  6. 0:09256 GB and two DGX Sparks. And go. And
  7. 0:15there they go. Okay, both wrote exactly
  8. 0:17800 tokens, 38.7 tokens a second on the
  9. 0:20M5 Ultra and 34.3 on the dual DGX Spark
  10. 0:24cluster. That's basically neck and neck,
  11. 0:26but
  12. 0:27we'll discover that there are
  13. 0:30quite a number of differences here. And
  14. 0:32this is the closest these two are going
  15. 0:34to get for the rest of the video. So,
  16. 0:35here's what we've got. On this side,
  17. 0:36this is the brand new M5 Ultra, top of
  18. 0:39the line right now with 256 gigs of
  19. 0:42memory, all in one box. And in this
  20. 0:44corner, we've got two DGX Sparks. Each
  21. 0:47one of them has 128 GB. That's also 256
  22. 0:51when you add them together. And they're
  23. 0:52connected by a one fat 200 gigabit
  24. 0:56cable. That cable runs something called
  25. 0:58Rocky.
  26. 0:59Not that Rocky. RDMA over converged
  27. 1:02Ethernet. It's like an acronym within an
  28. 1:04acronym. R is for RDMA, which is remote
  29. 1:07direct memory access. So, basically,
  30. 1:09they can write into each other's memory
  31. 1:11directly. The Mac as configured here is
  32. 1:13$14,000
  33. 1:16because it's got the 8 TB drive in
  34. 1:18there. The Sparks have 4 TB each, so
  35. 1:22together they're also 8 TB. And the
  36. 1:24Sparks, when they came out, they were
  37. 1:25four grand each. Now, they're almost
  38. 1:27five grand each. So, that's 10. Still
  39. 1:29cheaper than this.
  40. 1:31>> Plus the expensive cable. If you're a
  41. 1:32developer running big models locally or
  42. 1:35you want to service a small team, you
  43. 1:37should watch this video.
  44. 1:39And if you're not, you should also watch
  45. 1:40this video cuz it's pretty cool stuff.
  46. 1:42Merlin AI, it's an all-in-one AI tool
  47. 1:45and they gave my audience a big
  48. 1:46discount. I keep multiple AI tools
  49. 1:48around because each one is good at
  50. 1:50something, but it gets really expensive
  51. 1:53and bouncing between tabs breaks my
  52. 1:55focus. Merlin AI puts Chat GPT, Claude,
  53. 1:58Gemini, and more in one place so I can
  54. 2:01pick the best one for the moment,
  55. 2:02whether I'm coding, researching, or
  56. 2:05writing for a video. Watch this. I click
  57. 2:06the Merlin AI extension, chat with the
  58. 2:08web page to summarize what I'm reading,
  59. 2:11and pull out the important parts. And I
  60. 2:12even have my choice of models right at
  61. 2:14my fingertips. If I need something
  62. 2:16deeper, I turn on deep research, and it
  63. 2:18builds a clean structured report from
  64. 2:19multiple sources. And it also has quick
  65. 2:22modes like web, academic, and Reddit
  66. 2:24search. If you pay separately, Chat GPT
  67. 2:26is $20, Claude is $20, Gemini is $20,
  68. 2:30and that adds up fast. Merlin AI is
  69. 2:32cheaper because they buy AI API access
  70. 2:35in bulk. APIs cost less than the $20
  71. 2:37plans, and most people don't even use
  72. 2:40$20 worth of API in a month. And here's
  73. 2:42the discount. I click pricing, continue,
  74. 2:45it takes me to Stripe, I enter the promo
  75. 2:47code, and the total drops to $60 for the
  76. 2:49year. That's basically five bucks a
  77. 2:51month. I don't know how long this deal
  78. 2:52will be available, so grab it soon. The
  79. 2:54link is in the description.
  80. 2:58Now, two Sparks just don't turn into one
  81. 3:01big computer magically. vLLM, which is
  82. 3:05this high-throughput and
  83. 3:06memory-efficient inference and serving
  84. 3:08engine for LLMs. This is the software
  85. 3:11that most Nvidia setups run, and it
  86. 3:13splits the model across both of them. In
  87. 3:15this case, it's called tensor parallel.
  88. 3:17There's other kinds of parallelism, and
  89. 3:19I talk about that in other videos, but
  90. 3:21today we're doing tensor parallel. And
  91. 3:23what that means is that every layer of
  92. 3:25the model gets cut in half, one half on
  93. 3:27each Spark. For every single token, both
  94. 3:30boxes do their half. Then, they swap
  95. 3:33results over the cable before the next
  96. 3:35layer can even start. So, that's a lot
  97. 3:36of back and forth over a cable like
  98. 3:39this. This is a QSFP cable, it's called.
  99. 3:41That's just that standard right there.
  100. 3:44So, why bother with two? Well, modern
  101. 3:46models today are perfect fit for two of
  102. 3:49them. DeepSeek V4 flash, for example.
  103. 3:52That's about 150 160 GB. It doesn't fit
  104. 3:56on one 128 gig Spark. And yeah, I tried
  105. 3:59loading it once. It didn't work out so
  106. 4:01well. Plus, you need extra space for
  107. 4:03context. Once you load it across two of
  108. 4:05them, each node is holding just a little
  109. 4:07bit over 100 gigs. I got it loaded on
  110. 4:09both of them right now. We got 111 gigs.
  111. 4:11It's just showing one of them right now
  112. 4:13out of 128. And that's with extra system
  113. 4:16stuff going on, too. So, each one gets
  114. 4:17half the model plus a little room to
  115. 4:19work. And the Mac just loads the whole
  116. 4:21thing on one box. Just a quick note, I'm
  117. 4:24running the same model on both machines,
  118. 4:26but it's not the same file. Each file is
  119. 4:29rounded down to four bits or quantized,
  120. 4:31so it's smaller and faster. The Sparks
  121. 4:33are running vLLM, like I mentioned.
  122. 4:35Llama.cpp, a popular tool, works there,
  123. 4:38too, but vLLM is Nvidia's go-to for this
  124. 4:42kind of split. On the Mac, I tested both
  125. 4:44Llama.cpp and MLX. Sometimes one wins
  126. 4:47and sometimes the other one wins, and
  127. 4:48I'll point that out as we go.
  128. 4:52So, I'm going to be running Deep Seek V4
  129. 4:53Flash Qwen 3.8 Flash next. That's
  130. 4:57another pretty new one that's requires
  131. 4:59two Sparks cuz it's large enough. Now,
  132. 5:01every time you generate tokens, it
  133. 5:03actually happens in two steps. And I
  134. 5:05know some of you already know all this,
  135. 5:07but this is for the newcomers. First,
  136. 5:09the model reads your whole prompt all at
  137. 5:11once. That's pure math. It's matrix
  138. 5:14multiplication. So, whichever box has
  139. 5:16more compute wins. Usually, that stuff
  140. 5:19happens on the GPU. So, the more
  141. 5:20powerful GPUs
  142. 5:22get the job done faster. And that time
  143. 5:24is your time to first token, that
  144. 5:28waiting of the GPU number crunching.
  145. 5:29Then, it takes that answer and writes
  146. 5:31one token at a time. And every token is
  147. 5:34another trip through memory for the
  148. 5:36model's weights. Weights is just
  149. 5:38basically a collection of numbers, a
  150. 5:39huge collection of numbers, many, many
  151. 5:41gigabytes of collections of numbers.
  152. 5:42That's the files that you download. So,
  153. 5:44because for every token we use in the
  154. 5:46memory, the writing of the answer comes
  155. 5:48down to memory speed. In other words,
  156. 5:50the second part of inference is token
  157. 5:53generation, and it's reliant on memory
  158. 5:55bandwidth. So, let's take a look at
  159. 5:57writing. That's that second part. And
  160. 5:59this is the race from the start, the one
  161. 6:01I showed you in the beginning. A
  162. 6:02slightly different
  163. 6:03version of it, but very close numbers
  164. 6:05cuz I ran it multiple times. A short
  165. 6:07prompt with one user, I'm pointing over
  166. 6:09here cuz I have my charts over here. In
  167. 6:10DeepSeek, we're getting about 38 tokens
  168. 6:13per second on the M5 Ultra, and also
  169. 6:16about 38 tokens per second on the dual
  170. 6:18Sparks. That's basically a tie. Both are
  171. 6:21going a pretty decent speed. This is a
  172. 6:23big model, so that's not bad at all.
  173. 6:25With Qwen, we have a little bit of a
  174. 6:27difference there. 45 tokens per second
  175. 6:29on the Mac, and 38 tokens per second on
  176. 6:31the dual Sparks. That one I'll give to
  177. 6:32the Mac. Why did that happen? Well, my
  178. 6:35best guess is memory. Apple says the M5
  179. 6:38Ultra has 1.2 terabytes per second of
  180. 6:42memory bandwidth. That's a lot. Each
  181. 6:44Spark is rated at 273 GB a second,
  182. 6:47significantly less. And on top of that,
  183. 6:49they're syncing over a cable that I
  184. 6:51measured at about 111 GB, not the 200
  185. 6:54it's rated for, but still pretty good.
  186. 6:56But I didn't test how much that cable
  187. 6:58actually slows things down, so take this
  188. 7:00with a grain of salt. But the Mac has
  189. 7:02another trick. The engine you pick for
  190. 7:04the model matters a lot. For example,
  191. 7:06DeepSeek on llama.cpp writes at about 40
  192. 7:10tokens per second. But you take the same
  193. 7:12model and use MLX instead, and you're
  194. 7:14getting 53 tokens per second. That's 34%
  195. 7:18faster just from switching software. And
  196. 7:20on llama.cpp served exactly the same way
  197. 7:22the Mac and the Sparks were within 2%.
  198. 7:25So, yeah, that's faster than Sparks' 38
  199. 7:27tokens a second, but I only tested
  200. 7:29DeepSeek on MLX in process. That's the
  201. 7:32benchmark talking straight to the
  202. 7:33engine, not served over HTTP like a real
  203. 7:36chat server. Anyway, I'm collecting
  204. 7:38information, okay?
  205. 7:40I'm I'm trying my best to do all the
  206. 7:42tests that I can. There's a lot of tests
  207. 7:43that can be done. But so far the Mac is
  208. 7:45looking pretty good.
  209. 7:48Now just a quick little primer on
  210. 7:50tokens. A token is about 3/4 of a word,
  211. 7:53about there. So a common thing you'll
  212. 7:55hear is 32K or 32,000 tokens, which is
  213. 7:58kind of like a typical modern starting
  214. 8:01point. You start from there and you go
  215. 8:02up for context size. And that's about
  216. 8:0424,000 words, which is roughly a
  217. 8:06100-page document. How long is it going
  218. 8:08to take you to read a 100-page document,
  219. 8:10huh?
  220. 8:12I bet it's not going to take
  221. 8:1430 seconds or
  222. 8:16however many seconds. We'll find out.
  223. 8:18Now,
  224. 8:20reading or prompt processing, that's the
  225. 8:23GPU cranking away, remember? I started
  226. 8:25small at 2,000 tokens, both start fast.
  227. 8:29But the Sparks turn out to be twice as
  228. 8:32quick. Quinn, 1.5 seconds on the Mac,
  229. 8:350.85 seconds on the dual Sparks. Deep
  230. 8:39Seek, 2 and 1/2 seconds on the Mac, 1
  231. 8:40and 1/4 seconds on the dual Sparks. You
  232. 8:43won't care.
  233. 8:44It's only a 2,000 token prompt. You
  234. 8:46probably won't even notice. If you
  235. 8:48sneeze,
  236. 8:49by the time you wipe your nose it's
  237. 8:50done. But you might care
  238. 8:53later. We'll get to that. On Quinn, the
  239. 8:55Mac's faster writing even wins it back
  240. 8:57after a couple of hundred tokens. If we
  241. 8:59make that prompt a little longer, at
  242. 9:018,000 tokens Deep Seek takes about 4
  243. 9:04seconds on the Sparks and about 10
  244. 9:06seconds on the Mac. You double it to
  245. 9:0816,000 and the Mac's wait doubles also
  246. 9:12to 21 seconds. Now you're going to have
  247. 9:14to sneeze a few times and wipe your nose
  248. 9:17a few times, huh? Yeah, you're going to
  249. 9:18notice this one. So let's go big.
  250. 9:21Bigger, okay? I know you some of you
  251. 9:24going to say like, oh, 32,000 is not
  252. 9:26big. Okay, let's just take it one step
  253. 9:27at a time, okay? This time I'm using a
  254. 9:29real code base, 32,000, boom. There's 14
  255. 9:32Python files in here. Huh? The DGX
  256. 9:34Sparks started streaming already, which
  257. 9:36means they're generating tokens. Now
  258. 9:38we're past the calculation stage and the
  259. 9:40Mac is still
  260. 9:42reading the prompt. It's still
  261. 9:43processing. We're at 17 seconds for the
  262. 9:45two DJX Sparks. We're done. And
  263. 9:49yeah, the Mac is still thinking. Still
  264. 9:51reading the prompt. Yeah, 50 seconds. I
  265. 9:53can hear it generating.
  266. 9:55Yeah. But yeah, that's a big difference
  267. 9:58there. 300 tokens generated, 32,000
  268. 10:00prompt tokens. Now, if you look down
  269. 10:02here, Sparks 17 seconds to Mac Studio's
  270. 10:0650 seconds. That's about three times the
  271. 10:09wait. But hold on, don't run away yet
  272. 10:12buying the Sparks.
  273. 10:14Some of you might still want to pick up
  274. 10:16that Mac.
  275. 10:18I'll let you know why. Now here on this
  276. 10:19chart, this is where I use the slightly
  277. 10:21shorter prompt, but the ratio is about
  278. 10:23the same. So we got 44 seconds on the M5
  279. 10:25Ultra and 16.5 on the dual Sparks.
  280. 10:28Quinn, 26 seconds on the Ultra, 11.2 on
  281. 10:32the Sparks. So 2.4 times faster. Not
  282. 10:35even close. And after that, it kind of
  283. 10:37keeps going. Double the prompt, double
  284. 10:40the wait. Now with the Sparks, I did
  285. 10:42push them all the way up to 128,000
  286. 10:45tokens. I didn't run the Mac at 128,000
  287. 10:47tokens because I kind of saw a pattern
  288. 10:49here and if it kept its 32K pace, it
  289. 10:52would have been over 3 minutes of
  290. 10:53waiting. So it only slows down as the
  291. 10:55prompt gets longer. The Sparks did it in
  292. 10:5772 seconds. Still, way faster than the
  293. 11:01M3 Ultra.
  294. 11:03Okay?
  295. 11:04All right. I'm not trying to make
  296. 11:05excuses. I'm just saying we we still
  297. 11:07have improvements overall.
  298. 11:11Well, let's go back to that code base
  299. 11:13race. Deep seek with the 32K prompt. The
  300. 11:17Spark gets about a 30-second head start.
  301. 11:19After that, on Llama CPP, the writing
  302. 11:21speed is a tie. Quinn's different. The
  303. 11:23Sparks start about 15 seconds ahead, but
  304. 11:26the Mac gains a few thousandths of a
  305. 11:28second in every token. Divide one by the
  306. 11:30other and by my math, the Mac needs a
  307. 11:33few thousand tokens in order to catch
  308. 11:35up. I didn't time that one, it's an
  309. 11:36estimate. But that's quite a lot, it's a
  310. 11:38whole new file, basically. So, big
  311. 11:40inputs favor the Spark. Big output
  312. 11:44favors the Mac, at least on Quant. This
  313. 11:46is one of the reasons we have a lot of
  314. 11:47interest in doing this kind of
  315. 11:48disaggregated prefill decode using both
  316. 11:51the Sparks and a Mac. Sparks will do the
  317. 11:54prefill, Mac will do the decode. I did a
  318. 11:56video about this, an early early
  319. 11:58prototype a couple months ago. I'll link
  320. 12:00to it down below, you can check it out.
  321. 12:02It's interesting, but this is a new
  322. 12:03project by this guy Ash Hart. You can
  323. 12:06check it out, he's on Twitter. He's
  324. 12:07posting a lot about this stuff now. But,
  325. 12:10there's one thing that takes the edge
  326. 12:12off. You always see demos of oh, uh
  327. 12:16you know, these things when they're
  328. 12:18kicked off fresh, the prompt processing
  329. 12:21speed is so different. That's not how it
  330. 12:22happens in the real world. In a real
  331. 12:24scenario, in the same session, the
  332. 12:27servers keep what they've already read.
  333. 12:29There's a cache. With 16,000 tokens of
  334. 12:32context, the first ask took 21 seconds
  335. 12:36on the Mac and 8.3 seconds on the
  336. 12:38Sparks. All right, we've already seen
  337. 12:40this stuff. But, the follow-up question,
  338. 12:423 seconds on the Mac, 1.4 seconds on the
  339. 12:45Sparks. Yes, the Sparks are still a
  340. 12:48little bit faster, but you pay for that
  341. 12:50long read only once per session, not
  342. 12:53every time. And I recently made a video
  343. 12:54for members of the channel uh detailing
  344. 12:57different techniques for how to speed
  345. 12:58things up. It includes prefix caching.
  346. 13:01Thanks to the members of the channel, by
  347. 13:02the way. Really appreciate you.
  348. 13:03Sometimes they get extra videos uh when
  349. 13:06I get a chance to record them.
  350. 13:08Appreciate you all, anyway. And that's
  351. 13:09how coding tools actually work. Your
  352. 13:11agent sends the code base once, the
  353. 13:14server keeps it, every new question only
  354. 13:16adds a little bit on top. So, that long
  355. 13:19wait is mostly an initial first question
  356. 13:21cost, not an every question cost. Where
  357. 13:24it hurts the most is when the context
  358. 13:27keeps changing. Obviously, a new repo,
  359. 13:29big new files, or a long session that
  360. 13:32outgrows the cache. Those are all
  361. 13:34possibilities. Another thing I found
  362. 13:36when running Frontier models is when
  363. 13:38you're changing the model like from Opus
  364. 13:405 to Opus 5.5 or Fable 1.1, you have to
  365. 13:44recalculate all that. And the initial
  366. 13:46hit is longer usually. But we're not
  367. 13:48talking about Frontier models now, we're
  368. 13:49talking about local. All right?
  369. 13:52Same ideas though apply.
  370. 13:55Can MLX rescue the Mac on reading speed?
  371. 13:59Well, on Qwen it's kind of a split.
  372. 14:01Llama.cpp writes faster and MLX reads
  373. 14:04faster. MLX gets through 32K in about 19
  374. 14:08seconds instead of 26. That's still
  375. 14:10behind the Sparks' 11 seconds. And for
  376. 14:13Deep Seek, even at MLX's best reading
  377. 14:15speed, 32,000 tokens would take at least
  378. 14:1822 seconds, probably closer to 30. Yes,
  379. 14:21that's also slower than the Sparks.
  380. 14:23Whether MLX lets the Mac win that race
  381. 14:25back, I don't know yet. I haven't run
  382. 14:27Deep Seek yet on MLX with longer prompts
  383. 14:29or multiple users. Told you, there's a
  384. 14:31lot of tests to do and I'm crunching
  385. 14:33through them.
  386. 14:36More than one user, or the number of
  387. 14:38concurrencies sometimes it's called, uh
  388. 14:40you'll see a bigger number quoted
  389. 14:42usually. And it's the total speed of the
  390. 14:45throughput of all the tokens generated
  391. 14:47for all the users. But careful with that
  392. 14:49number though, because it shows
  393. 14:51everybody's number. It's everyone's
  394. 14:52tokens added up. Each person only gets a
  395. 14:55slice. The server writes for everyone
  396. 14:57together. And when somebody new shows
  397. 14:59up, well, the server has to stop and
  398. 15:02read their prompt too, which uh is going
  399. 15:04to add to that initial calculation. On
  400. 15:06the Mac that reading is slower, past
  401. 15:08about four users, it's spending more
  402. 15:10time reading than writing. So, at that
  403. 15:12point everyone's slice shrinks a little
  404. 15:15bit. And you can see it. On the Mac,
  405. 15:16Deep Seek peaks at four users, 66 tokens
  406. 15:20per second there. That's total. Then it
  407. 15:22drops to 46 tokens per second for eight
  408. 15:25users. But the Sparks, they keep
  409. 15:27climbing to 70 tokens per second. And I
  410. 15:29haven't done 16, so I don't know where
  411. 15:31it drops off. Uh
  412. 15:34TBD. Now, give everybody a thousand
  413. 15:37tokens of chat history and at eight
  414. 15:40users, the Mac does about 11 tokens per
  415. 15:43second. The Sparks, 25. And that's
  416. 15:45total. But the real pain is the wait.
  417. 15:47Each person on the Mac waits over a
  418. 15:49minute for the first word and then gets
  419. 15:50about five tokens a second. On the
  420. 15:52Sparks, it's about 24 second wait and
  421. 15:54then about seven tokens per second. So,
  422. 15:56yeah, sharing these machines,
  423. 15:59keep it to yourself, all right? You can
  424. 16:01do it, but use smaller models, maybe.
  425. 16:05Or, yeah, there's there's different ways
  426. 16:07of splitting these machines up, but just
  427. 16:09don't use big models for a lot of
  428. 16:11people. It's not going to turn out so
  429. 16:12well. Looking at Quinn here, at eight
  430. 16:15users, the Sparks put out 120 tokens a
  431. 16:18second total. That's pretty good. Yeah,
  432. 16:20that's way better than Deep Seek. The
  433. 16:22Mac does 66 in Llama CPP and 70 in MLX.
  434. 16:26And for that, I used OMLX, which is a
  435. 16:28new tool. Newish. There it is. No more
  436. 16:31waiting on your Mac. Well, there is some
  437. 16:33waiting, okay? Really, the Mac is a
  438. 16:35machine for one person, maybe two. The
  439. 16:37Sparks are a little bit better for a
  440. 16:39small team, maybe four people. Eight?
  441. 16:43You're pushing it. Now, I didn't start
  442. 16:44out with two Sparks. I went to four and
  443. 16:46then eight. That was a much bigger
  444. 16:47project. It was actually experimental,
  445. 16:49more like uh with breakout cables and
  446. 16:52extra switches that I had to buy. I made
  447. 16:53a whole video about it. Compared to
  448. 16:55that, two Sparks is kind of a perfect
  449. 16:57setup. It's just one cable and Nvidia's
  450. 17:00developer site, build.nvidia.com/spark,
  451. 17:04has some really interesting recipes that
  452. 17:06are super easy to follow and it just
  453. 17:08works most of the time. Shows you how to
  454. 17:10connect two of them, how to run multiple
  455. 17:12workloads via LLM, everything. Pretty
  456. 17:14easy. You still have to line things up.
  457. 17:15I matched the OS, the kernel, the
  458. 17:17driver, the firmware on both of the
  459. 17:19machines. Then you have to set up V L M
  460. 17:21in multi-node with a pile of nickel and
  461. 17:24rocky environment variables. After that,
  462. 17:27you have to check the traffic going over
  463. 17:29the RDMA connection. What you get for
  464. 17:32all that is CUDA and V L L M. So
  465. 17:34basically, it's the same kind of stack
  466. 17:35that you'd run in the cloud. But on the
  467. 17:38Mac side, this machine isn't just for
  468. 17:39AI. It's also good at AI, but for
  469. 17:43example, my daily driver is an M5 Max
  470. 17:46MacBook Pro. Everything that I do on
  471. 17:48that machine, I can do it on the M5
  472. 17:50Ultra, but faster. That includes running
  473. 17:53large models. And for everyday work,
  474. 17:56it's zero setup, basically. I can
  475. 17:58offload a lot of stuff to it, like
  476. 18:00rendering my videos, for example, for
  477. 18:02this channel. And um sometimes the Mac
  478. 18:05gets a little bogged down. I run a lot
  479. 18:07of stuff on it. And right now, I'm using
  480. 18:0988 GB of memory. Yeah, it gets bogged
  481. 18:12down a little bit, even with 128 gigs of
  482. 18:15memory. So it's nice to be able to
  483. 18:16offload stuff very easily. And here is
  484. 18:19another thing. There they go, full GPU
  485. 18:22utilization on both machines. Or yeah,
  486. 18:25you only see one here, but they're both
  487. 18:27working on the sparks. And the Mac is
  488. 18:29using 100% of GPU as well. Oh boy.
  489. 18:34Now, during my M5 Ultra first look, I
  490. 18:36noticed that the machine was pretty
  491. 18:38toasty. And yeah, it is. But sparks also
  492. 18:41have been known to be toasty. You know,
  493. 18:43we're not alone here. Right now, I'm
  494. 18:45doing a very heavy load on the cluster
  495. 18:48here and the Mac Studio. And this is
  496. 18:50what I'm seeing. This is nuts. All
  497. 18:53right. Uh
  498. 18:54I don't want to pop a breaker here, but
  499. 18:55I might. 434 W being used by the M5
  500. 18:59Ultra and 410 by the spark cluster. So
  501. 19:03we're very close. I should say at idle,
  502. 19:06it's a very different story. They are
  503. 19:08pretty warm now.
  504. 19:09And I'm hearing noise coming out of
  505. 19:12everywhere. Just like noise throughout.
  506. 19:16They feel about the same actually. Yeah,
  507. 19:18they're both pretty orange. About 49°
  508. 19:22to 50° on the hottest part of the Mac
  509. 19:25Studio, about 48°
  510. 19:28on the Sparks. Let's take a look at the
  511. 19:30back. Ooh.
  512. 19:3256° on the grill, the back of the Mac
  513. 19:35Studio. And wow, 58 60 63 I saw in there
  514. 19:41on the Sparks. So yeah, both get pretty
  515. 19:43toasty. For AI on the Mac, there is a
  516. 19:45little setup depending on which stack
  517. 19:47you want, Llama CPP or MLX. OMLX is
  518. 19:51pretty easy. There's also newer ones
  519. 19:52like native without the E. And there is
  520. 19:55Splash. Some of these are so new I
  521. 19:57haven't even tried them yet. These are
  522. 19:59basically home brew installs. So is 256
  523. 20:02gigs across two boxes the same as 256 in
  524. 20:06one? Well, it holds the same model. It
  525. 20:08just reads a lot faster. But before you
  526. 20:10go out and drop 10 grand or more on one
  527. 20:13of these setups, look at your own work.
  528. 20:15How big are your prompts? How long are
  529. 20:17the answers to your prompts? You want to
  530. 20:18learn more about clustering the Sparks,
  531. 20:20watch this video here. Clustering Mac
  532. 20:22Studios, watch this video here. Thanks
  533. 20:25for watching and I'll see you next time.

About this transcript

This page contains the full transcript of M5 Ultra vs 2 DGX Sparks… The Number You're Not Looking At by Alex Ziskind, generated from the public captions YouTube serves with the video. The transcript has 3,740 words across 533 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.