YouTube2Text

8 RTX Pro 6000’s Wasn’t What I Expected — Transcript

by Alex Ziskind · 4,450 words · 630 segments · language en · Watch on YouTube

Full transcript

  1. 0:00So, I recently built this machine with
  2. 0:02four RTX Pro 6000s, but who needs four
  3. 0:04when you can have eight? This is
  4. 0:07definitely sold last week. Right before
  5. 0:08I dropped the box on my face, I got
  6. 0:11something new. Oh, yeah.
  7. 0:14This is the Camino Grando, and it's 768
  8. 0:18GB of VRAM all in one box. Now, the box
  9. 0:22is not that much bigger than the other
  10. 0:24box. Actually, it's quite small for
  11. 0:26having eight of these things in here,
  12. 0:27but it's heavy. This thing weighs almost
  13. 0:29as much as me. Eight RTX Pro 6000s are
  14. 0:33stacked in here. And yeah, they're close
  15. 0:36together. That's because this whole
  16. 0:37thing is liquid cooled. A little secret,
  17. 0:40I've actually been testing this thing
  18. 0:41for a few months. I was so excited about
  19. 0:43this thing. And there are a few things
  20. 0:45that I'd like to point out that are not
  21. 0:47great besides the obvious amazing thing
  22. 0:50that it's going to be super fast. And of
  23. 0:52course, it's going to hold gigantic
  24. 0:54models and run them all quickly. But for
  25. 0:57whom? Who needs something like this?
  26. 0:59Sure, one developer with a dozen coding
  27. 1:01agents could use it all at once or a
  28. 1:03whole team of developers can use it in
  29. 1:05the office at the same time. For
  30. 1:06example, a 400 GBTE LLM with over
  31. 1:09400,000 token context window. Yeah, I
  32. 1:12ran that all that one box, no cloud. So,
  33. 1:15of course, I had to put it through its
  34. 1:17paces.
  35. 1:19[music]
  36. 1:20So, here's the thing. Everybody talks
  37. 1:22about running local LLMs and usually
  38. 1:24that means one person, one laptop, one
  39. 1:27model, you get 30, 40 tokens per second
  40. 1:30and you're happy, right? But that's not
  41. 1:32how developers work anymore. That was a
  42. 1:34year ago and now it's completely
  43. 1:35different. Even if you're solo, you're
  44. 1:37not running one chat window. You got a
  45. 1:39coding agent on this repo, you got
  46. 1:41another one writing tests over there and
  47. 1:42another one reviewing and every one of
  48. 1:45them sends 50,000 tokens. not your
  49. 1:47prompt, but the whole context because
  50. 1:50yeah, you need to include all that.
  51. 1:51That's a team's worth of load from one
  52. 1:54person. Uh, I probably should rephrase
  53. 1:56that next time. And if you're actually a
  54. 1:58team, you multiply that by 10.
  55. 2:01Today, I'm looking at the Camino Grando.
  56. 2:03Camino has been around for a while, and
  57. 2:05their specialty is to build these liquid
  58. 2:07cool GPU workstations and servers. I did
  59. 2:09not buy this one. They let me borrow it
  60. 2:12for a little bit. And at current prices,
  61. 2:14just the GPUs alone are about 15 grand a
  62. 2:17piece and there's eight of them. So
  63. 2:19yeah, you better put these to good use.
  64. 2:21This thing comes as a 4 unitit chassis.
  65. 2:23You can rack mount it or you can desktop
  66. 2:26mount it, which is what I did. And it's
  67. 2:28121 lbs just the machine, but it came in
  68. 2:30a big box. And for insurance purposes, I
  69. 2:32did not carry it alone. I had help.
  70. 2:36Right inside, we got a single AMD epic
  71. 2:399474F,
  72. 2:4148 cores, 96 threads. It's an epic CPU.
  73. 2:45What can I say? 512 gigs of DDR5 is in
  74. 2:48there, too, which is an insane amount of
  75. 2:52RAM right now. I understand. But then
  76. 2:54there's also the reason we're all here,
  77. 2:56eight Nvidia RTX Pro 6000 Blackwell
  78. 2:59Server Edition. Each one has 96 GB of
  79. 3:02VRAM. That's 768 GB of VRAM total. And
  80. 3:05did I try to run GLM 5.2 on it? Yes, I
  81. 3:09did. Did it succeed? Sort of. Now, for
  82. 3:11scale, the RTX 5090 has 32 gigs of VRAM.
  83. 3:15So, this is like 24 50s worth of memory
  84. 3:18in one box. That's a good title for this
  85. 3:20video. How do you manage to fit eight of
  86. 3:22these in a 4unit box when each one is a
  87. 3:25two slot card? You can take that cooler
  88. 3:27off. Every GPU has a custom copper water
  89. 3:30block on it. It covers the die, the
  90. 3:31memory, and the VRM. And that turns the
  91. 3:34two slot card into just a single slot.
  92. 3:37Every pair of cards has its own little
  93. 3:39manifold. And the fittings are dripless
  94. 3:42quick disconnects, color-coded red and
  95. 3:44blue, so you can pull one GPU out
  96. 3:45without draining the loop. And the whole
  97. 3:47thing is one shared loop for CPUs and
  98. 3:50GPUs with a 450ml reservoir with pumps
  99. 3:53built into it. Now, you may be wondering
  100. 3:55what it takes to power something like
  101. 3:57this.
  102. 3:58Yeah, 6 1/2 kW. That's like having four
  103. 4:02space heaters running all plugged into
  104. 4:04the same wall and it's not going to
  105. 4:06work. You need something special for
  106. 4:07that. Now, the Grondo comes with four
  107. 4:09PSUs and each one of them has a regular
  108. 4:12standard plug. So, you can run it into
  109. 4:15different outlets at the same time. They
  110. 4:17all combine the power, but make sure
  111. 4:19they're on different circuits. Now,
  112. 4:21here's one thing that people often
  113. 4:23overlook with GPUs, with multiple GPUs,
  114. 4:25is PCIe lanes. This is why we need the
  115. 4:28Thread Ripper or the Epic. The Epic has
  116. 4:30128 lanes. So each GPU gets a certain
  117. 4:34set of lanes, 16 to be precise, or at
  118. 4:36least most of them. Seven of them get 16
  119. 4:39lanes and one of them gets eight lanes.
  120. 4:41Now, because GPUs are using up most of
  121. 4:43the lanes, you only have two M.2 slots,
  122. 4:46so you have to use them wisely. And I I
  123. 4:48did run into issues where I had to swap
  124. 4:51out models because I ran out of space.
  125. 4:54There's also no Envy Link, so all these
  126. 4:56cards have to talk to each other via
  127. 4:58PCIe. There's four hot swappable 200
  128. 5:00watt power supplies. That's 8 kW of
  129. 5:03capacity total. You probably don't want
  130. 5:05to use all eight. Yeah, it's going to
  131. 5:07destroy your stuff. 8 RTX Pro 6000
  132. 5:10running at their full 600 watt
  133. 5:12potential. That's a total of 4,800 watts
  134. 5:15just in GPUs. Now, I don't have a 8 kW
  135. 5:18circuit in here. So, I had two PSUs
  136. 5:21plugged into a 240 volt 20 amp circuit,
  137. 5:24one into a regular 120 volt 15 amp
  138. 5:27outlet, and another one in a Jackaryi
  139. 5:30battery. Yeah, it's on a jackeri. Hey,
  140. 5:33I'm only human. Okay,
  141. 5:36now just a quick bit of housekeeping. I
  142. 5:38ran the GPUs capped at 300 watts each
  143. 5:41instead of the default 600. Don't leave
  144. 5:43yet. Don't leave. I'll tell you why it
  145. 5:45worked fine. Partly because of my wiring
  146. 5:47situation. partly because when I pushed
  147. 5:50it at 600, the machine actually dropped
  148. 5:52on me a couple times overnight while I
  149. 5:55was doing my long runs. You're probably
  150. 5:56thinking, half the power, half the
  151. 5:58speed, right? Well, that turned out to
  152. 6:01be one of the more interesting finds in
  153. 6:02this project. Whenever I'm working away
  154. 6:04from home, I end up connecting to
  155. 6:05networks I basically know nothing about.
  156. 6:08A network I completely trust. Not
  157. 6:10really. So, before I start working, I
  158. 6:13connect to Surf Shark. My terminal
  159. 6:15sessions, repo traffic, and container
  160. 6:17downloads all travel through an AES 256
  161. 6:20encrypted tunnel, and Clean Web blocks a
  162. 6:22lot of the tracking and advertising junk
  163. 6:24that I don't want. There's also an
  164. 6:25independently audited no logs policy.
  165. 6:28Plus, RAM only servers that get wiped
  166. 6:31every time they restart. And with the
  167. 6:32amount of hardware I have and I travel
  168. 6:34with, unlimited devices is a pretty big
  169. 6:36deal. It covers my MacBook, my phone,
  170. 6:38and pretty much any other device I have
  171. 6:40in my backpack that connects to the
  172. 6:42internet. And when I'm checking a server
  173. 6:44or sshing back to the office, that extra
  174. 6:47layer makes a lot of sense. So head over
  175. 6:49to surfshark.com/alexiscin
  176. 6:51for four extra months free. And if it's
  177. 6:54not for you, there's a 30-day money back
  178. 6:56guarantee. Protect your connection. Get
  179. 6:58your work done. Now, back to the video.
  180. 7:02Let's talk about noise. The Camino says
  181. 7:0539 to 70 dB depending on which fans you
  182. 7:09get because you can customize that.
  183. 7:10There's a 6200 RPM version and a 3,000
  184. 7:13RPM version. But let me tell you this,
  185. 7:15you're not going to have this on your
  186. 7:17desk next to you.
  187. 7:27Okay, it calms down after a little bit.
  188. 7:32>> Even at its quietest setting, under full
  189. 7:35AGPU load, it gets pretty loud. Yeah, it
  190. 7:38could get pretty loud if you're sitting
  191. 7:40next to it. I'd move away a little bit
  192. 7:43more. Maybe even more.
  193. 7:48Can you still hear it? Yeah. You need to
  194. 7:52be far away from it. Now, the cooling is
  195. 7:55a different story. The hottest GPU I saw
  196. 7:57was 62 Celsius and the CPU peaked at 65.
  197. 8:01So, the liquid cooling here is not a
  198. 8:03gimmick. That's the reason this box
  199. 8:05exists. It's really good and efficient.
  200. 8:07[music]
  201. 8:09All right, let's go over some model
  202. 8:10results. Ubuntu 24.04, Nvidia Open
  203. 8:14Driver 610, CUDA 13, and VLM Nightly. I
  204. 8:18kept it updated as I was doing my
  205. 8:20testing. And we're using Tensor
  206. 8:22Parallelism across all eight GPUs.
  207. 8:24There's GLM 5.2 right on top. 433 GB of
  208. 8:29weights. And then a few other new ones
  209. 8:31came out, so I had to do those. Here's a
  210. 8:33list of five models I ran mostly in
  211. 8:35NVFP4 format, which is the 4-bit format,
  212. 8:38floating point that Blackwell supports.
  213. 8:41It's what it was built for. The big boy
  214. 8:42GLM 5.2. This one was served with 49,600
  215. 8:48token context window. Quen 3 235B.
  216. 8:51Getting a little bit long in the tooth
  217. 8:52that model, but it's pretty big. It's a
  218. 8:54mixture of experts model. 22 billion
  219. 8:56parameters active. 144 GB on disk. This
  220. 9:00size right here, 144 to about 185 for
  221. 9:04GLM Flash 5.3. This is kind of like the
  222. 9:06sweet spot for this type of machine.
  223. 9:09Sure, it can run bigger models, but then
  224. 9:11you're going to do context. You're going
  225. 9:12to have multiple sessions, multiple
  226. 9:14agents hitting it at the same time, so
  227. 9:16you want to probably stick around here.
  228. 9:17GLM 5.2 loads in about 4 1/2 minutes and
  229. 9:21lands at 738 GB of the 784 available.
  230. 9:25That's about 95% of the VRAM on this
  231. 9:27machine just gone for one model. But it
  232. 9:29fits. That's the whole point. There's
  233. 9:31basically no other single box you can
  234. 9:34put on a desk that does this. There's a
  235. 9:36DJX station, which I recently did a
  236. 9:38video on. That one has 252 GB of high
  237. 9:41bandwidth memory, which is way faster
  238. 9:43than this memory, but it doesn't have a
  239. 9:45total capacity to hold this model all in
  240. 9:48VRAM. By the way, the DJX Station is
  241. 9:50also really good at the smaller models.
  242. 9:52And I want to compare how the smaller
  243. 9:54models act on this machine versus the
  244. 9:57DJX station. Stay tuned for that.
  245. 10:00[music] Let's start where most of
  246. 10:02developers live. One developer, one
  247. 10:04agent. I know, I know you're using more
  248. 10:07than that now, but let's start at the
  249. 10:08baseline. 2048 token prompt, 128 tokens
  250. 10:11out. On a single RTX Pro 6000, a 30
  251. 10:15billion parameter class model is around
  252. 10:17100 tokens per second. So, what happens
  253. 10:19when you spread big models across eight
  254. 10:22cards over PCIe? Ah, GLM 5.2 48 tokens
  255. 10:27per second generation. Okay, that's a
  256. 10:30433 gig model doing 48. That's pretty
  257. 10:33good. Not blazing, but it's fast enough.
  258. 10:36We'll come back to that. Quen 3 235B 85.
  259. 10:39Wow. Okay. Now, here are some of the
  260. 10:42more modern models. Deepseek v4 flash
  261. 10:46102 tokens per second. GLM 5.3 104 and
  262. 10:50Quen 3.8 wins this one. 126 tokens per
  263. 10:54second. Quen 3.8 flash next by the way.
  264. 10:56All right. Quen 4 architecture. Boom.
  265. 10:58Oh, if you don't know what I'm talking
  266. 10:59about, Quen 3.8 models that came out
  267. 11:02earlier are actually Quen 3 model
  268. 11:05architecture or 3.8 and Quen 3.8 Flash.
  269. 11:09Next has Quen 4 architecture. Don't ask
  270. 11:12me why. I don't know why they did that.
  271. 11:14Okay. The other number you should be
  272. 11:15interested in prompt processing because
  273. 11:18this number matters a lot especially
  274. 11:20when you're using agents in a code
  275. 11:22editor. For example, GLM 5.2 about 2200
  276. 11:25tokens per second. Quen 3235B we're up a
  277. 11:29lot to 4779 tokens per second. GLM 5.3
  278. 11:33Flash 8,300. Deepseek V4 flash 8700 and
  279. 11:37Quen 3.8 Flash. Next, holy cow 12,600.
  280. 11:42That's fast. Now, here's what it looks
  281. 11:44like in practice for a single user for
  282. 11:47time to first token on a 2,00 token
  283. 11:50prompt. A quarter of a second on flash
  284. 11:52next up to almost a second for GLM 5.2
  285. 11:55and everything else is in between there.
  286. 11:57When we look back at that 12,700 tokens
  287. 12:00per second of prompt processing, that
  288. 12:01means a 2,00 token prompt is basically
  289. 12:04instant.
  290. 12:06Now, I gave it a 128,000 token prompt.
  291. 12:10That's basically a whole codebase. a
  292. 12:13small code base on a Mac, on a mini PC,
  293. 12:16or pretty much anything else I've
  294. 12:17tested. A 128,000 token prompt is kind
  295. 12:21of a let's go make coffee situation.
  296. 12:24[snorts]
  297. 12:25>> Oh, okay. Coffee can wait. I'm still
  298. 12:28waiting for the M5 Ultras, which are
  299. 12:30about to come out. That might change
  300. 12:32things a little bit, but on everything
  301. 12:33older, you're going to be waiting
  302. 12:34minutes. [music] As the prompt length
  303. 12:36increases, your time to first token also
  304. 12:39increases, sometimes quite a bit. Gwen
  305. 12:413.8 8 flash. Next, we're at 14 1/2
  306. 12:44seconds to first token at 128,000
  307. 12:47[music]
  308. 12:48tokens. For GLM 5.3 flash, we're at
  309. 12:51about 18 and 12. Deepseek V4 flash,
  310. 12:54we're at 26. And in GLM 5.2, wo, 50
  311. 12:59seconds. That's the big boy, right? So,
  312. 13:02yeah, to be expected. Still though,
  313. 13:04under a minute for the biggest model and
  314. 13:07about 15 seconds for the fastest one.
  315. 13:09Quen 3 235B didn't make it that far. We
  316. 13:12only got 32 and uh then it kind of
  317. 13:15crashed on me for the rest of the times.
  318. 13:17Sorry. So this is what eight GPUs
  319. 13:19earning their keep looks like because
  320. 13:21prompt processing that's the first stage
  321. 13:23of inference before token generation
  322. 13:26happens. This is all computebound. So
  323. 13:28everything happens on the GPU chip. But
  324. 13:30then the second stage is token
  325. 13:32generation. So what happens to
  326. 13:34generation speed once all that context
  327. 13:36is sitting in memory? Huh? Every line is
  328. 13:39pretty much flat here, huh? Nothing
  329. 13:43happens. Flash Next is still generating
  330. 13:46at about 120 to 122 tokens per second
  331. 13:49even at 128,000 token context. GLM 5.2
  332. 13:54pretty steady around what is that 40 50
  333. 13:57none of them slow down. That's pretty
  334. 13:59incredible.
  335. 14:01Now, quick aside here because this
  336. 14:03matters to agents. I'm using VLM here to
  337. 14:05serve the models and VLM has prefix
  338. 14:08caching. If your conversation has about
  339. 14:1116,000 tokens in it and you send it
  340. 14:13another message without caching, GLM 5.2
  341. 14:17takes 5.9 seconds to first token. And
  342. 14:20with caching, 0.83
  343. 14:23seconds, 7 times faster. And you kind of
  344. 14:26see the same pattern all throughout.
  345. 14:27You'll say, "Alex, that's obvious. Turn
  346. 14:29on caching, right?" Well, especially
  347. 14:31with Agentic Flows, you want to have
  348. 14:33that option on.
  349. 14:36Okay, remember the power cap? I ran the
  350. 14:39same test at 300 watts and then 600
  351. 14:41watts. And you said 300 is going to be
  352. 14:43slower than 600. Well, actually, I'm the
  353. 14:45one that said it, but you were thinking
  354. 14:46it, weren't you? GLM 5.2 48 tokens per
  355. 14:49second at 300 watts. Quen 3235b 85. GLM
  356. 14:545.2, this is 32 users now, by the way.
  357. 14:57We got 111 tokens per second at 300
  358. 15:00watts. And Quen 3235B
  359. 15:0364 users 254 tokens per second. What
  360. 15:07does this look like at 600 watts? Boom.
  361. 15:10It's a wash. I wouldn't even call that
  362. 15:12close to being different. Really? Well,
  363. 15:14actually 254 and 252. Yeah, I know.
  364. 15:18Okay, calm down. It's close enough. And
  365. 15:20look at the actual GPU power draw. At
  366. 15:22300 watt caps, the eight GPU together
  367. 15:25pull about 1,550 watts during inference.
  368. 15:28This is for GLM 5.2.
  369. 15:30>> That's just one of the plugs and it's
  370. 15:33drawing 126 watts idle.
  371. 15:36>> And at 600, they pull about 1,700 W per
  372. 15:40card. The peak I ever saw was 268 watts,
  373. 15:44which means that these cards are memory
  374. 15:46bandwidth bound during inference. and
  375. 15:48they're talking to each other over PCIe,
  376. 15:50so they never get anywhere near the 600
  377. 15:52watt limit. Doubling the power limit
  378. 15:54bought me nothing more than heat and a
  379. 15:56less stable box. Which brings me to a
  380. 15:58question that I've had for a while. The
  381. 15:59workstation edition, which has 600 watt
  382. 16:01cap versus the Max Q edition, and that's
  383. 16:04going to be another video, I think. So,
  384. 16:06stay tuned for that. Make sure you
  385. 16:07subscribe. By the way, just sitting
  386. 16:09there with a model loaded and nothing
  387. 16:10happening, the GPUs are pulling about
  388. 16:13700 watts idle. just sitting there.
  389. 16:15[snorts] It's doing nothing expensively.
  390. 16:19So, make sure you pay the power bill.
  391. 16:23Don't get me wrong, this is not magic.
  392. 16:26Every time those AGPUs sync up, they go
  393. 16:28over PCIe and you feel it. It's not Envy
  394. 16:31Link. I would love to test some Envy
  395. 16:34Link on this channel. So, if you're
  396. 16:35listening and you have some of those
  397. 16:37laying around, let me know. Get in
  398. 16:39touch. A100s, anyone? H200s, Quen 3 235B
  399. 16:42on four GPUs with tensor parallel 4 did
  400. 16:4558 tokens per second in an earlier test.
  401. 16:47I did on all eight GPUs TP8 it did 37.
  402. 16:52So slower more GPUs slower because of
  403. 16:55the all reduce over PCIe costs more
  404. 16:58[music] than extra compute gives you.
  405. 17:00There's the balance that you need to
  406. 17:02maintain and you need to tweak it. So
  407. 17:03for a smaller model that would fit on
  408. 17:06four cards, not all eight, you would
  409. 17:08actually probably better off running two
  410. 17:10copies on four cards each and serve
  411. 17:12twice the people, not one copy on eight.
  412. 17:15GLM 5.2 at 48 tokens per second is with
  413. 17:18the plain official recipe. There's no
  414. 17:20spec decoding here. If you're curious
  415. 17:22about speculative decoding, I made a
  416. 17:23separate video on that. I'll link to it
  417. 17:24down below. And with MTP turned on, I
  418. 17:26got about 100 tokens per second in
  419. 17:29earlier testing. But yeah, that config
  420. 17:30was a little bit fussier to make. Now,
  421. 17:32the ecosystem is still catching up to
  422. 17:34these GPUs. These are still considered
  423. 17:36pretty new, even though they've been
  424. 17:37around for now almost 2 years in some
  425. 17:40form or another. For example, GLM 5.3
  426. 17:43Flash would not start on any stock VLM
  427. 17:46image. I tried six different
  428. 17:48configurations, got the same kernel
  429. 17:50error every time because the model uses
  430. 17:52an attention variant the Blackwell
  431. 17:54workstation kernel doesn't handle yet.
  432. 17:56It's too new. So, once VLM upstreams it
  433. 17:58and supports it, you might get slightly
  434. 18:00different numbers. It may be better
  435. 18:02even.
  436. 18:04All right, let's try this out.
  437. 18:07This is going to be nuts. All four of
  438. 18:09these machines are going to ping the
  439. 18:11Grandondo, which is obviously not in
  440. 18:13this room, but you might be able to hear
  441. 18:15it. Yeah, it's spinning up.
  442. 18:22This is Maya. She's doing a test suite
  443. 18:24for orders API. This is Ravi. He's
  444. 18:26working on a metrics dashboard. Lena is
  445. 18:29hardening the inventory report script.
  446. 18:32Yeah, I didn't write this. Okay. But
  447. 18:34it's still a test. And Tom over there.
  448. 18:37[laughter]
  449. 18:38Tom containerizing the Q worker. Yeah.
  450. 18:41So, they're all busy. Okay. I just
  451. 18:43That's the whole point here. I need to
  452. 18:45keep them busy. [laughter]
  453. 18:48All right. Now, we're cooking. Each one
  454. 18:50of these is doing about a bunch. A
  455. 18:54bunch. Whatever the agent needs. Each
  456. 18:56one of these launched multiple agents
  457. 18:58and sub agents all pointing at the grand
  458. 19:02all at the same time.
  459. 19:04Some of these require permission. Let's
  460. 19:06do always. Boom. And [music] confirm. He
  461. 19:09must be doing something dangerous there.
  462. 19:11Maya or Lena. This is showing the
  463. 19:13utilization and the VRAMm usage on each
  464. 19:16of the GPUs running on the Grando. So,
  465. 19:19we're 100% utilization, jumping up to
  466. 19:23about 185 to 190 watts per GPU and 90
  467. 19:28out of 96 GB used on each of those RTX
  468. 19:32Pro 6000s. And this is running DeepSync
  469. 19:34V4 Flash. It's going pretty fast. It's
  470. 19:38going decently fast. And all these
  471. 19:40agents and sub agents are all getting
  472. 19:43their share. It's plenty. We're gonna go
  473. 19:46with 32 agents on each machine to start
  474. 19:49with. And boom.
  475. 19:51Three, two, one, and go. [laughter]
  476. 19:56Woo. That's spinning up. You might be
  477. 20:00able to hear it from the other room.
  478. 20:01Nice. So, we got a total of 128
  479. 20:04concurrent agents. Each one has 32. And
  480. 20:07it's chugging away about 3,000 tokens
  481. 20:09per second. All right, let's stop that.
  482. 20:11Let's go to 128
  483. 20:13agents.
  484. 20:17per machine. Oh my gosh, this is going
  485. 20:19to be crazy. Launch all. Oh boy, it's
  486. 20:21doing it. 512 concurrent agents. Wow.
  487. 20:26Utilization is about the same. 100%. It
  488. 20:29better be. Finally. Let's do 256. And
  489. 20:33that's 256 agents for Maya, Ravi, Lena,
  490. 20:37and Tom.
  491. 20:39The four ninjas over here.
  492. 20:43Let's go. Launch all.
  493. 20:45Three, two, one. Boom. Well, [laughter]
  494. 20:50it's not liking it. It's definitely not
  495. 20:54liking it. Oh boy. It's trying. It's
  496. 20:57trying. 1,000 concurrent agents. Hey, it
  497. 21:02did it. We're at 6,000 to 7,000 tokens
  498. 21:05per second [laughter]
  499. 21:07and it's actually doing it. It had to
  500. 21:10spin up. Wow. KV cash 41% and zero
  501. 21:14errors. Time to first token is pretty
  502. 21:16high there. And it's not super happy,
  503. 21:20but it's completing the work. Since I've
  504. 21:23been talking here, we we've done over
  505. 21:26630,000
  506. 21:28tokens total. Not bad. Not bad. Job well
  507. 21:31done, humans. You can all go home.
  508. 21:35Bye, Maya. I'll miss you.
  509. 21:38>> [music]
  510. 21:39>> So I ran 1 2 4 8 16 32 64 [music]
  511. 21:44concurrent users against all five
  512. 21:47models. Same 448 token prompt. No queue
  513. 21:50and every user actually in flight at the
  514. 21:53same time. Quen 3.8 flash. Next one user
  515. 21:56gets 126 tokens per second. And as we
  516. 21:59add more users, the total throughput
  517. 22:01goes up because we get more tokens per
  518. 22:03second generated totally by the machine.
  519. 22:05But but the number of tokens per second
  520. 22:09per user goes down. So with four users
  521. 22:12we get 90. Eight users 82. And by the
  522. 22:15way, this is either eight users or it
  523. 22:17could be one developer running eight
  524. 22:18agents in parallel. So here we've got 82
  525. 22:21tokens per second for all eight of your
  526. 22:23agents, which is still pretty good.
  527. 22:25That's faster per agent than most people
  528. 22:28get from one agent in a cloud API. 16
  529. 22:30users, 61, 32 users, 37, and 64 users.
  530. 22:35We're down to 20.6. Not amazing down
  531. 22:38there. Let's take a look at these other
  532. 22:39models here. And you can see where we
  533. 22:41stand. Quen 3.8 being the fastest, of
  534. 22:43course. GLM 5.3 Flash and Deepc4 Flash
  535. 22:46are about the same. Here are the numbers
  536. 22:48for 1. Here are the numbers for 4
  537. 22:51[music] 8 and then ridiculous 64. Are
  538. 22:54you running 64 agents at the same time?
  539. 22:56Tell the truth. Come on. What's the most
  540. 22:58number of agents you ran at the same
  541. 22:59time? Put it down in the comments. And
  542. 23:01the machine is putting out 530 tokens
  543. 23:03per second aggregate at that point with
  544. 23:05a peak decode or token generation burst
  545. 23:08up to 2,880. Here are the numbers for
  546. 23:11one user, four users, eight users, and
  547. 23:14finally we got 64 users. There's that
  548. 23:16530 from Quen 3.8 flash. Next, wo time
  549. 23:21to first token looks uh pretty
  550. 23:22interesting, right? Especially for GLM
  551. 23:245.2 there. Time to first token is that
  552. 23:26wait before anything at all shows up. At
  553. 23:28eight users, we're at about a second and
  554. 23:31a half for all three Flash models. 5.5
  555. 23:34seconds for GLM 5.2. At 16, we're at 2
  556. 23:381/2 seconds and almost 10 seconds for
  557. 23:40GLM 5.2. And at 64,
  558. 23:44you know, we're getting to be in a place
  559. 23:47where you might not want to be,
  560. 23:48especially with GLM 5.2. Waiting 33
  561. 23:51seconds for time to first open. GH, I'm
  562. 23:54sorry you have to go through that. We're
  563. 23:55so impatient these days, aren't we? This
  564. 23:57thing is literally writing our stuff for
  565. 23:59us. And come on, hurry up. Task masters.
  566. 24:03Thanks for staying through the whole
  567. 24:05video, by the way. And thanks to the
  568. 24:06members of the channel. I appreciate you
  569. 24:07all. Short version, with the right
  570. 24:10models, 16 developers on coding agents
  571. 24:13or one developer with 16 agents. Every
  572. 24:16one of these models is getting 60 tokens
  573. 24:18per second with a 2 and 1/2 second wait
  574. 24:21time from one box just sitting hopefully
  575. 24:24in another room. And for reference,
  576. 24:25storage review, which had the exact same
  577. 24:28machine as me. It got shipped to me from
  578. 24:30them. They did a write up on this. They
  579. 24:33ran Claude Code sessions against Miniax
  580. 24:36on the same machine, and they got 39
  581. 24:37tokens per second per user at eight
  582. 24:40sessions, which is about what you get
  583. 24:42from Claude Opus through the API. So,
  584. 24:45would I buy one of these? Well, I mean,
  585. 24:48uh, I can't, but here's how I think of
  586. 24:50it. If you're a solo developer, this is
  587. 24:53the machine where you stop thinking
  588. 24:55about the model as a thing you wait for.
  589. 24:58Also, if you're a solo developer buying
  590. 25:00one of these, you better be making some
  591. 25:02decent money off of your gigs or renting
  592. 25:05it out or something. The main target
  593. 25:07audience here is obviously going to be
  594. 25:09businesses, small businesses, medium
  595. 25:11businesses, where there's teams of
  596. 25:12people working off of these things.
  597. 25:14Eight agents at 80 tokens per second
  598. 25:16each. A whole code base in the prompt in
  599. 25:1814 seconds. It's kind of overkill for
  600. 25:20one person if you ask me, but it's one
  601. 25:22of two boxes that I've recently tested
  602. 25:24where I don't feel like I'm limited by
  603. 25:26how many agents I could run. And if
  604. 25:27you're one person that builds like a
  605. 25:29team of agents, hey, it's not that
  606. 25:32crazy. The other thing you should think
  607. 25:33about is this is really kind of new
  608. 25:37hardware still. [music] So, you're going
  609. 25:38to need to get comfortable with nightly
  610. 25:40VLM builds. Maybe getting your hands a
  611. 25:42little bit dirty with the
  612. 25:43behindthe-scenes activities or if you
  613. 25:45just want to stand it up and run it. The
  614. 25:47recipes are out there. Nvidia has a
  615. 25:49bunch of recipes on their site. VLM has
  616. 25:52recipes which I've been using as well.
  617. 25:54Pretty [music] handy. They don't always
  618. 25:55list the exact thing that you have
  619. 25:57running. For example, uh here is a
  620. 25:59recipe for RTX Pro 6000, but four of
  621. 26:02them. So, you just need to modify it to
  622. 26:05your own size. For example, tensor
  623. 26:07parallel size, you'd set that to eight.
  624. 26:09Or if it's a small enough model, you can
  625. 26:11run a couple of instances of it. Like I
  626. 26:13mentioned earlier, I recently did that
  627. 26:15build with four RTX Pro [music] 6000s,
  628. 26:17and you can watch that right over here
  629. 26:18next. Thanks for watching, and I'll see
  630. 26:20you next time.

About this transcript

This page contains the full transcript of 8 RTX Pro 6000’s Wasn’t What I Expected by Alex Ziskind, generated from the public captions YouTube serves with the video. The transcript has 4,450 words across 630 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.