8 RTX Pro 6000’s Wasn’t What I Expected — Transcript
Full transcript
- 0:00So, I recently built this machine with
- 0:02four RTX Pro 6000s, but who needs four
- 0:04when you can have eight? This is
- 0:07definitely sold last week. Right before
- 0:08I dropped the box on my face, I got
- 0:11something new. Oh, yeah.
- 0:14This is the Camino Grando, and it's 768
- 0:18GB of VRAM all in one box. Now, the box
- 0:22is not that much bigger than the other
- 0:24box. Actually, it's quite small for
- 0:26having eight of these things in here,
- 0:27but it's heavy. This thing weighs almost
- 0:29as much as me. Eight RTX Pro 6000s are
- 0:33stacked in here. And yeah, they're close
- 0:36together. That's because this whole
- 0:37thing is liquid cooled. A little secret,
- 0:40I've actually been testing this thing
- 0:41for a few months. I was so excited about
- 0:43this thing. And there are a few things
- 0:45that I'd like to point out that are not
- 0:47great besides the obvious amazing thing
- 0:50that it's going to be super fast. And of
- 0:52course, it's going to hold gigantic
- 0:54models and run them all quickly. But for
- 0:57whom? Who needs something like this?
- 0:59Sure, one developer with a dozen coding
- 1:01agents could use it all at once or a
- 1:03whole team of developers can use it in
- 1:05the office at the same time. For
- 1:06example, a 400 GBTE LLM with over
- 1:09400,000 token context window. Yeah, I
- 1:12ran that all that one box, no cloud. So,
- 1:15of course, I had to put it through its
- 1:17paces.
- 1:19[music]
- 1:20So, here's the thing. Everybody talks
- 1:22about running local LLMs and usually
- 1:24that means one person, one laptop, one
- 1:27model, you get 30, 40 tokens per second
- 1:30and you're happy, right? But that's not
- 1:32how developers work anymore. That was a
- 1:34year ago and now it's completely
- 1:35different. Even if you're solo, you're
- 1:37not running one chat window. You got a
- 1:39coding agent on this repo, you got
- 1:41another one writing tests over there and
- 1:42another one reviewing and every one of
- 1:45them sends 50,000 tokens. not your
- 1:47prompt, but the whole context because
- 1:50yeah, you need to include all that.
- 1:51That's a team's worth of load from one
- 1:54person. Uh, I probably should rephrase
- 1:56that next time. And if you're actually a
- 1:58team, you multiply that by 10.
- 2:01Today, I'm looking at the Camino Grando.
- 2:03Camino has been around for a while, and
- 2:05their specialty is to build these liquid
- 2:07cool GPU workstations and servers. I did
- 2:09not buy this one. They let me borrow it
- 2:12for a little bit. And at current prices,
- 2:14just the GPUs alone are about 15 grand a
- 2:17piece and there's eight of them. So
- 2:19yeah, you better put these to good use.
- 2:21This thing comes as a 4 unitit chassis.
- 2:23You can rack mount it or you can desktop
- 2:26mount it, which is what I did. And it's
- 2:28121 lbs just the machine, but it came in
- 2:30a big box. And for insurance purposes, I
- 2:32did not carry it alone. I had help.
- 2:36Right inside, we got a single AMD epic
- 2:399474F,
- 2:4148 cores, 96 threads. It's an epic CPU.
- 2:45What can I say? 512 gigs of DDR5 is in
- 2:48there, too, which is an insane amount of
- 2:52RAM right now. I understand. But then
- 2:54there's also the reason we're all here,
- 2:56eight Nvidia RTX Pro 6000 Blackwell
- 2:59Server Edition. Each one has 96 GB of
- 3:02VRAM. That's 768 GB of VRAM total. And
- 3:05did I try to run GLM 5.2 on it? Yes, I
- 3:09did. Did it succeed? Sort of. Now, for
- 3:11scale, the RTX 5090 has 32 gigs of VRAM.
- 3:15So, this is like 24 50s worth of memory
- 3:18in one box. That's a good title for this
- 3:20video. How do you manage to fit eight of
- 3:22these in a 4unit box when each one is a
- 3:25two slot card? You can take that cooler
- 3:27off. Every GPU has a custom copper water
- 3:30block on it. It covers the die, the
- 3:31memory, and the VRM. And that turns the
- 3:34two slot card into just a single slot.
- 3:37Every pair of cards has its own little
- 3:39manifold. And the fittings are dripless
- 3:42quick disconnects, color-coded red and
- 3:44blue, so you can pull one GPU out
- 3:45without draining the loop. And the whole
- 3:47thing is one shared loop for CPUs and
- 3:50GPUs with a 450ml reservoir with pumps
- 3:53built into it. Now, you may be wondering
- 3:55what it takes to power something like
- 3:57this.
- 3:58Yeah, 6 1/2 kW. That's like having four
- 4:02space heaters running all plugged into
- 4:04the same wall and it's not going to
- 4:06work. You need something special for
- 4:07that. Now, the Grondo comes with four
- 4:09PSUs and each one of them has a regular
- 4:12standard plug. So, you can run it into
- 4:15different outlets at the same time. They
- 4:17all combine the power, but make sure
- 4:19they're on different circuits. Now,
- 4:21here's one thing that people often
- 4:23overlook with GPUs, with multiple GPUs,
- 4:25is PCIe lanes. This is why we need the
- 4:28Thread Ripper or the Epic. The Epic has
- 4:30128 lanes. So each GPU gets a certain
- 4:34set of lanes, 16 to be precise, or at
- 4:36least most of them. Seven of them get 16
- 4:39lanes and one of them gets eight lanes.
- 4:41Now, because GPUs are using up most of
- 4:43the lanes, you only have two M.2 slots,
- 4:46so you have to use them wisely. And I I
- 4:48did run into issues where I had to swap
- 4:51out models because I ran out of space.
- 4:54There's also no Envy Link, so all these
- 4:56cards have to talk to each other via
- 4:58PCIe. There's four hot swappable 200
- 5:00watt power supplies. That's 8 kW of
- 5:03capacity total. You probably don't want
- 5:05to use all eight. Yeah, it's going to
- 5:07destroy your stuff. 8 RTX Pro 6000
- 5:10running at their full 600 watt
- 5:12potential. That's a total of 4,800 watts
- 5:15just in GPUs. Now, I don't have a 8 kW
- 5:18circuit in here. So, I had two PSUs
- 5:21plugged into a 240 volt 20 amp circuit,
- 5:24one into a regular 120 volt 15 amp
- 5:27outlet, and another one in a Jackaryi
- 5:30battery. Yeah, it's on a jackeri. Hey,
- 5:33I'm only human. Okay,
- 5:36now just a quick bit of housekeeping. I
- 5:38ran the GPUs capped at 300 watts each
- 5:41instead of the default 600. Don't leave
- 5:43yet. Don't leave. I'll tell you why it
- 5:45worked fine. Partly because of my wiring
- 5:47situation. partly because when I pushed
- 5:50it at 600, the machine actually dropped
- 5:52on me a couple times overnight while I
- 5:55was doing my long runs. You're probably
- 5:56thinking, half the power, half the
- 5:58speed, right? Well, that turned out to
- 6:01be one of the more interesting finds in
- 6:02this project. Whenever I'm working away
- 6:04from home, I end up connecting to
- 6:05networks I basically know nothing about.
- 6:08A network I completely trust. Not
- 6:10really. So, before I start working, I
- 6:13connect to Surf Shark. My terminal
- 6:15sessions, repo traffic, and container
- 6:17downloads all travel through an AES 256
- 6:20encrypted tunnel, and Clean Web blocks a
- 6:22lot of the tracking and advertising junk
- 6:24that I don't want. There's also an
- 6:25independently audited no logs policy.
- 6:28Plus, RAM only servers that get wiped
- 6:31every time they restart. And with the
- 6:32amount of hardware I have and I travel
- 6:34with, unlimited devices is a pretty big
- 6:36deal. It covers my MacBook, my phone,
- 6:38and pretty much any other device I have
- 6:40in my backpack that connects to the
- 6:42internet. And when I'm checking a server
- 6:44or sshing back to the office, that extra
- 6:47layer makes a lot of sense. So head over
- 6:49to surfshark.com/alexiscin
- 6:51for four extra months free. And if it's
- 6:54not for you, there's a 30-day money back
- 6:56guarantee. Protect your connection. Get
- 6:58your work done. Now, back to the video.
- 7:02Let's talk about noise. The Camino says
- 7:0539 to 70 dB depending on which fans you
- 7:09get because you can customize that.
- 7:10There's a 6200 RPM version and a 3,000
- 7:13RPM version. But let me tell you this,
- 7:15you're not going to have this on your
- 7:17desk next to you.
- 7:27Okay, it calms down after a little bit.
- 7:32>> Even at its quietest setting, under full
- 7:35AGPU load, it gets pretty loud. Yeah, it
- 7:38could get pretty loud if you're sitting
- 7:40next to it. I'd move away a little bit
- 7:43more. Maybe even more.
- 7:48Can you still hear it? Yeah. You need to
- 7:52be far away from it. Now, the cooling is
- 7:55a different story. The hottest GPU I saw
- 7:57was 62 Celsius and the CPU peaked at 65.
- 8:01So, the liquid cooling here is not a
- 8:03gimmick. That's the reason this box
- 8:05exists. It's really good and efficient.
- 8:07[music]
- 8:09All right, let's go over some model
- 8:10results. Ubuntu 24.04, Nvidia Open
- 8:14Driver 610, CUDA 13, and VLM Nightly. I
- 8:18kept it updated as I was doing my
- 8:20testing. And we're using Tensor
- 8:22Parallelism across all eight GPUs.
- 8:24There's GLM 5.2 right on top. 433 GB of
- 8:29weights. And then a few other new ones
- 8:31came out, so I had to do those. Here's a
- 8:33list of five models I ran mostly in
- 8:35NVFP4 format, which is the 4-bit format,
- 8:38floating point that Blackwell supports.
- 8:41It's what it was built for. The big boy
- 8:42GLM 5.2. This one was served with 49,600
- 8:48token context window. Quen 3 235B.
- 8:51Getting a little bit long in the tooth
- 8:52that model, but it's pretty big. It's a
- 8:54mixture of experts model. 22 billion
- 8:56parameters active. 144 GB on disk. This
- 9:00size right here, 144 to about 185 for
- 9:04GLM Flash 5.3. This is kind of like the
- 9:06sweet spot for this type of machine.
- 9:09Sure, it can run bigger models, but then
- 9:11you're going to do context. You're going
- 9:12to have multiple sessions, multiple
- 9:14agents hitting it at the same time, so
- 9:16you want to probably stick around here.
- 9:17GLM 5.2 loads in about 4 1/2 minutes and
- 9:21lands at 738 GB of the 784 available.
- 9:25That's about 95% of the VRAM on this
- 9:27machine just gone for one model. But it
- 9:29fits. That's the whole point. There's
- 9:31basically no other single box you can
- 9:34put on a desk that does this. There's a
- 9:36DJX station, which I recently did a
- 9:38video on. That one has 252 GB of high
- 9:41bandwidth memory, which is way faster
- 9:43than this memory, but it doesn't have a
- 9:45total capacity to hold this model all in
- 9:48VRAM. By the way, the DJX Station is
- 9:50also really good at the smaller models.
- 9:52And I want to compare how the smaller
- 9:54models act on this machine versus the
- 9:57DJX station. Stay tuned for that.
- 10:00[music] Let's start where most of
- 10:02developers live. One developer, one
- 10:04agent. I know, I know you're using more
- 10:07than that now, but let's start at the
- 10:08baseline. 2048 token prompt, 128 tokens
- 10:11out. On a single RTX Pro 6000, a 30
- 10:15billion parameter class model is around
- 10:17100 tokens per second. So, what happens
- 10:19when you spread big models across eight
- 10:22cards over PCIe? Ah, GLM 5.2 48 tokens
- 10:27per second generation. Okay, that's a
- 10:30433 gig model doing 48. That's pretty
- 10:33good. Not blazing, but it's fast enough.
- 10:36We'll come back to that. Quen 3 235B 85.
- 10:39Wow. Okay. Now, here are some of the
- 10:42more modern models. Deepseek v4 flash
- 10:46102 tokens per second. GLM 5.3 104 and
- 10:50Quen 3.8 wins this one. 126 tokens per
- 10:54second. Quen 3.8 flash next by the way.
- 10:56All right. Quen 4 architecture. Boom.
- 10:58Oh, if you don't know what I'm talking
- 10:59about, Quen 3.8 models that came out
- 11:02earlier are actually Quen 3 model
- 11:05architecture or 3.8 and Quen 3.8 Flash.
- 11:09Next has Quen 4 architecture. Don't ask
- 11:12me why. I don't know why they did that.
- 11:14Okay. The other number you should be
- 11:15interested in prompt processing because
- 11:18this number matters a lot especially
- 11:20when you're using agents in a code
- 11:22editor. For example, GLM 5.2 about 2200
- 11:25tokens per second. Quen 3235B we're up a
- 11:29lot to 4779 tokens per second. GLM 5.3
- 11:33Flash 8,300. Deepseek V4 flash 8700 and
- 11:37Quen 3.8 Flash. Next, holy cow 12,600.
- 11:42That's fast. Now, here's what it looks
- 11:44like in practice for a single user for
- 11:47time to first token on a 2,00 token
- 11:50prompt. A quarter of a second on flash
- 11:52next up to almost a second for GLM 5.2
- 11:55and everything else is in between there.
- 11:57When we look back at that 12,700 tokens
- 12:00per second of prompt processing, that
- 12:01means a 2,00 token prompt is basically
- 12:04instant.
- 12:06Now, I gave it a 128,000 token prompt.
- 12:10That's basically a whole codebase. a
- 12:13small code base on a Mac, on a mini PC,
- 12:16or pretty much anything else I've
- 12:17tested. A 128,000 token prompt is kind
- 12:21of a let's go make coffee situation.
- 12:24[snorts]
- 12:25>> Oh, okay. Coffee can wait. I'm still
- 12:28waiting for the M5 Ultras, which are
- 12:30about to come out. That might change
- 12:32things a little bit, but on everything
- 12:33older, you're going to be waiting
- 12:34minutes. [music] As the prompt length
- 12:36increases, your time to first token also
- 12:39increases, sometimes quite a bit. Gwen
- 12:413.8 8 flash. Next, we're at 14 1/2
- 12:44seconds to first token at 128,000
- 12:47[music]
- 12:48tokens. For GLM 5.3 flash, we're at
- 12:51about 18 and 12. Deepseek V4 flash,
- 12:54we're at 26. And in GLM 5.2, wo, 50
- 12:59seconds. That's the big boy, right? So,
- 13:02yeah, to be expected. Still though,
- 13:04under a minute for the biggest model and
- 13:07about 15 seconds for the fastest one.
- 13:09Quen 3 235B didn't make it that far. We
- 13:12only got 32 and uh then it kind of
- 13:15crashed on me for the rest of the times.
- 13:17Sorry. So this is what eight GPUs
- 13:19earning their keep looks like because
- 13:21prompt processing that's the first stage
- 13:23of inference before token generation
- 13:26happens. This is all computebound. So
- 13:28everything happens on the GPU chip. But
- 13:30then the second stage is token
- 13:32generation. So what happens to
- 13:34generation speed once all that context
- 13:36is sitting in memory? Huh? Every line is
- 13:39pretty much flat here, huh? Nothing
- 13:43happens. Flash Next is still generating
- 13:46at about 120 to 122 tokens per second
- 13:49even at 128,000 token context. GLM 5.2
- 13:54pretty steady around what is that 40 50
- 13:57none of them slow down. That's pretty
- 13:59incredible.
- 14:01Now, quick aside here because this
- 14:03matters to agents. I'm using VLM here to
- 14:05serve the models and VLM has prefix
- 14:08caching. If your conversation has about
- 14:1116,000 tokens in it and you send it
- 14:13another message without caching, GLM 5.2
- 14:17takes 5.9 seconds to first token. And
- 14:20with caching, 0.83
- 14:23seconds, 7 times faster. And you kind of
- 14:26see the same pattern all throughout.
- 14:27You'll say, "Alex, that's obvious. Turn
- 14:29on caching, right?" Well, especially
- 14:31with Agentic Flows, you want to have
- 14:33that option on.
- 14:36Okay, remember the power cap? I ran the
- 14:39same test at 300 watts and then 600
- 14:41watts. And you said 300 is going to be
- 14:43slower than 600. Well, actually, I'm the
- 14:45one that said it, but you were thinking
- 14:46it, weren't you? GLM 5.2 48 tokens per
- 14:49second at 300 watts. Quen 3235b 85. GLM
- 14:545.2, this is 32 users now, by the way.
- 14:57We got 111 tokens per second at 300
- 15:00watts. And Quen 3235B
- 15:0364 users 254 tokens per second. What
- 15:07does this look like at 600 watts? Boom.
- 15:10It's a wash. I wouldn't even call that
- 15:12close to being different. Really? Well,
- 15:14actually 254 and 252. Yeah, I know.
- 15:18Okay, calm down. It's close enough. And
- 15:20look at the actual GPU power draw. At
- 15:22300 watt caps, the eight GPU together
- 15:25pull about 1,550 watts during inference.
- 15:28This is for GLM 5.2.
- 15:30>> That's just one of the plugs and it's
- 15:33drawing 126 watts idle.
- 15:36>> And at 600, they pull about 1,700 W per
- 15:40card. The peak I ever saw was 268 watts,
- 15:44which means that these cards are memory
- 15:46bandwidth bound during inference. and
- 15:48they're talking to each other over PCIe,
- 15:50so they never get anywhere near the 600
- 15:52watt limit. Doubling the power limit
- 15:54bought me nothing more than heat and a
- 15:56less stable box. Which brings me to a
- 15:58question that I've had for a while. The
- 15:59workstation edition, which has 600 watt
- 16:01cap versus the Max Q edition, and that's
- 16:04going to be another video, I think. So,
- 16:06stay tuned for that. Make sure you
- 16:07subscribe. By the way, just sitting
- 16:09there with a model loaded and nothing
- 16:10happening, the GPUs are pulling about
- 16:13700 watts idle. just sitting there.
- 16:15[snorts] It's doing nothing expensively.
- 16:19So, make sure you pay the power bill.
- 16:23Don't get me wrong, this is not magic.
- 16:26Every time those AGPUs sync up, they go
- 16:28over PCIe and you feel it. It's not Envy
- 16:31Link. I would love to test some Envy
- 16:34Link on this channel. So, if you're
- 16:35listening and you have some of those
- 16:37laying around, let me know. Get in
- 16:39touch. A100s, anyone? H200s, Quen 3 235B
- 16:42on four GPUs with tensor parallel 4 did
- 16:4558 tokens per second in an earlier test.
- 16:47I did on all eight GPUs TP8 it did 37.
- 16:52So slower more GPUs slower because of
- 16:55the all reduce over PCIe costs more
- 16:58[music] than extra compute gives you.
- 17:00There's the balance that you need to
- 17:02maintain and you need to tweak it. So
- 17:03for a smaller model that would fit on
- 17:06four cards, not all eight, you would
- 17:08actually probably better off running two
- 17:10copies on four cards each and serve
- 17:12twice the people, not one copy on eight.
- 17:15GLM 5.2 at 48 tokens per second is with
- 17:18the plain official recipe. There's no
- 17:20spec decoding here. If you're curious
- 17:22about speculative decoding, I made a
- 17:23separate video on that. I'll link to it
- 17:24down below. And with MTP turned on, I
- 17:26got about 100 tokens per second in
- 17:29earlier testing. But yeah, that config
- 17:30was a little bit fussier to make. Now,
- 17:32the ecosystem is still catching up to
- 17:34these GPUs. These are still considered
- 17:36pretty new, even though they've been
- 17:37around for now almost 2 years in some
- 17:40form or another. For example, GLM 5.3
- 17:43Flash would not start on any stock VLM
- 17:46image. I tried six different
- 17:48configurations, got the same kernel
- 17:50error every time because the model uses
- 17:52an attention variant the Blackwell
- 17:54workstation kernel doesn't handle yet.
- 17:56It's too new. So, once VLM upstreams it
- 17:58and supports it, you might get slightly
- 18:00different numbers. It may be better
- 18:02even.
- 18:04All right, let's try this out.
- 18:07This is going to be nuts. All four of
- 18:09these machines are going to ping the
- 18:11Grandondo, which is obviously not in
- 18:13this room, but you might be able to hear
- 18:15it. Yeah, it's spinning up.
- 18:22This is Maya. She's doing a test suite
- 18:24for orders API. This is Ravi. He's
- 18:26working on a metrics dashboard. Lena is
- 18:29hardening the inventory report script.
- 18:32Yeah, I didn't write this. Okay. But
- 18:34it's still a test. And Tom over there.
- 18:37[laughter]
- 18:38Tom containerizing the Q worker. Yeah.
- 18:41So, they're all busy. Okay. I just
- 18:43That's the whole point here. I need to
- 18:45keep them busy. [laughter]
- 18:48All right. Now, we're cooking. Each one
- 18:50of these is doing about a bunch. A
- 18:54bunch. Whatever the agent needs. Each
- 18:56one of these launched multiple agents
- 18:58and sub agents all pointing at the grand
- 19:02all at the same time.
- 19:04Some of these require permission. Let's
- 19:06do always. Boom. And [music] confirm. He
- 19:09must be doing something dangerous there.
- 19:11Maya or Lena. This is showing the
- 19:13utilization and the VRAMm usage on each
- 19:16of the GPUs running on the Grando. So,
- 19:19we're 100% utilization, jumping up to
- 19:23about 185 to 190 watts per GPU and 90
- 19:28out of 96 GB used on each of those RTX
- 19:32Pro 6000s. And this is running DeepSync
- 19:34V4 Flash. It's going pretty fast. It's
- 19:38going decently fast. And all these
- 19:40agents and sub agents are all getting
- 19:43their share. It's plenty. We're gonna go
- 19:46with 32 agents on each machine to start
- 19:49with. And boom.
- 19:51Three, two, one, and go. [laughter]
- 19:56Woo. That's spinning up. You might be
- 20:00able to hear it from the other room.
- 20:01Nice. So, we got a total of 128
- 20:04concurrent agents. Each one has 32. And
- 20:07it's chugging away about 3,000 tokens
- 20:09per second. All right, let's stop that.
- 20:11Let's go to 128
- 20:13agents.
- 20:17per machine. Oh my gosh, this is going
- 20:19to be crazy. Launch all. Oh boy, it's
- 20:21doing it. 512 concurrent agents. Wow.
- 20:26Utilization is about the same. 100%. It
- 20:29better be. Finally. Let's do 256. And
- 20:33that's 256 agents for Maya, Ravi, Lena,
- 20:37and Tom.
- 20:39The four ninjas over here.
- 20:43Let's go. Launch all.
- 20:45Three, two, one. Boom. Well, [laughter]
- 20:50it's not liking it. It's definitely not
- 20:54liking it. Oh boy. It's trying. It's
- 20:57trying. 1,000 concurrent agents. Hey, it
- 21:02did it. We're at 6,000 to 7,000 tokens
- 21:05per second [laughter]
- 21:07and it's actually doing it. It had to
- 21:10spin up. Wow. KV cash 41% and zero
- 21:14errors. Time to first token is pretty
- 21:16high there. And it's not super happy,
- 21:20but it's completing the work. Since I've
- 21:23been talking here, we we've done over
- 21:26630,000
- 21:28tokens total. Not bad. Not bad. Job well
- 21:31done, humans. You can all go home.
- 21:35Bye, Maya. I'll miss you.
- 21:38>> [music]
- 21:39>> So I ran 1 2 4 8 16 32 64 [music]
- 21:44concurrent users against all five
- 21:47models. Same 448 token prompt. No queue
- 21:50and every user actually in flight at the
- 21:53same time. Quen 3.8 flash. Next one user
- 21:56gets 126 tokens per second. And as we
- 21:59add more users, the total throughput
- 22:01goes up because we get more tokens per
- 22:03second generated totally by the machine.
- 22:05But but the number of tokens per second
- 22:09per user goes down. So with four users
- 22:12we get 90. Eight users 82. And by the
- 22:15way, this is either eight users or it
- 22:17could be one developer running eight
- 22:18agents in parallel. So here we've got 82
- 22:21tokens per second for all eight of your
- 22:23agents, which is still pretty good.
- 22:25That's faster per agent than most people
- 22:28get from one agent in a cloud API. 16
- 22:30users, 61, 32 users, 37, and 64 users.
- 22:35We're down to 20.6. Not amazing down
- 22:38there. Let's take a look at these other
- 22:39models here. And you can see where we
- 22:41stand. Quen 3.8 being the fastest, of
- 22:43course. GLM 5.3 Flash and Deepc4 Flash
- 22:46are about the same. Here are the numbers
- 22:48for 1. Here are the numbers for 4
- 22:51[music] 8 and then ridiculous 64. Are
- 22:54you running 64 agents at the same time?
- 22:56Tell the truth. Come on. What's the most
- 22:58number of agents you ran at the same
- 22:59time? Put it down in the comments. And
- 23:01the machine is putting out 530 tokens
- 23:03per second aggregate at that point with
- 23:05a peak decode or token generation burst
- 23:08up to 2,880. Here are the numbers for
- 23:11one user, four users, eight users, and
- 23:14finally we got 64 users. There's that
- 23:16530 from Quen 3.8 flash. Next, wo time
- 23:21to first token looks uh pretty
- 23:22interesting, right? Especially for GLM
- 23:245.2 there. Time to first token is that
- 23:26wait before anything at all shows up. At
- 23:28eight users, we're at about a second and
- 23:31a half for all three Flash models. 5.5
- 23:34seconds for GLM 5.2. At 16, we're at 2
- 23:381/2 seconds and almost 10 seconds for
- 23:40GLM 5.2. And at 64,
- 23:44you know, we're getting to be in a place
- 23:47where you might not want to be,
- 23:48especially with GLM 5.2. Waiting 33
- 23:51seconds for time to first open. GH, I'm
- 23:54sorry you have to go through that. We're
- 23:55so impatient these days, aren't we? This
- 23:57thing is literally writing our stuff for
- 23:59us. And come on, hurry up. Task masters.
- 24:03Thanks for staying through the whole
- 24:05video, by the way. And thanks to the
- 24:06members of the channel. I appreciate you
- 24:07all. Short version, with the right
- 24:10models, 16 developers on coding agents
- 24:13or one developer with 16 agents. Every
- 24:16one of these models is getting 60 tokens
- 24:18per second with a 2 and 1/2 second wait
- 24:21time from one box just sitting hopefully
- 24:24in another room. And for reference,
- 24:25storage review, which had the exact same
- 24:28machine as me. It got shipped to me from
- 24:30them. They did a write up on this. They
- 24:33ran Claude Code sessions against Miniax
- 24:36on the same machine, and they got 39
- 24:37tokens per second per user at eight
- 24:40sessions, which is about what you get
- 24:42from Claude Opus through the API. So,
- 24:45would I buy one of these? Well, I mean,
- 24:48uh, I can't, but here's how I think of
- 24:50it. If you're a solo developer, this is
- 24:53the machine where you stop thinking
- 24:55about the model as a thing you wait for.
- 24:58Also, if you're a solo developer buying
- 25:00one of these, you better be making some
- 25:02decent money off of your gigs or renting
- 25:05it out or something. The main target
- 25:07audience here is obviously going to be
- 25:09businesses, small businesses, medium
- 25:11businesses, where there's teams of
- 25:12people working off of these things.
- 25:14Eight agents at 80 tokens per second
- 25:16each. A whole code base in the prompt in
- 25:1814 seconds. It's kind of overkill for
- 25:20one person if you ask me, but it's one
- 25:22of two boxes that I've recently tested
- 25:24where I don't feel like I'm limited by
- 25:26how many agents I could run. And if
- 25:27you're one person that builds like a
- 25:29team of agents, hey, it's not that
- 25:32crazy. The other thing you should think
- 25:33about is this is really kind of new
- 25:37hardware still. [music] So, you're going
- 25:38to need to get comfortable with nightly
- 25:40VLM builds. Maybe getting your hands a
- 25:42little bit dirty with the
- 25:43behindthe-scenes activities or if you
- 25:45just want to stand it up and run it. The
- 25:47recipes are out there. Nvidia has a
- 25:49bunch of recipes on their site. VLM has
- 25:52recipes which I've been using as well.
- 25:54Pretty [music] handy. They don't always
- 25:55list the exact thing that you have
- 25:57running. For example, uh here is a
- 25:59recipe for RTX Pro 6000, but four of
- 26:02them. So, you just need to modify it to
- 26:05your own size. For example, tensor
- 26:07parallel size, you'd set that to eight.
- 26:09Or if it's a small enough model, you can
- 26:11run a couple of instances of it. Like I
- 26:13mentioned earlier, I recently did that
- 26:15build with four RTX Pro [music] 6000s,
- 26:17and you can watch that right over here
- 26:18next. Thanks for watching, and I'll see
- 26:20you next time.
About this transcript
This page contains the full transcript of 8 RTX Pro 6000’s Wasn’t What I Expected by Alex Ziskind, generated from the public captions YouTube serves with the video. The transcript has 4,450 words across 630 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.