M5 Ultra vs 2 DGX Sparks… The Number You're Not Looking At — Transcript
Full transcript
- 0:00Are you thinking what I'm thinking?
- 0:01These edges are way too sharp. You
- 0:03should not let your kids play with them.
- 0:04It even says so on the box.
- 0:06>> I've got the M5 Ultra Mac Studio here,
- 0:09256 GB and two DGX Sparks. And go. And
- 0:15there they go. Okay, both wrote exactly
- 0:17800 tokens, 38.7 tokens a second on the
- 0:20M5 Ultra and 34.3 on the dual DGX Spark
- 0:24cluster. That's basically neck and neck,
- 0:26but
- 0:27we'll discover that there are
- 0:30quite a number of differences here. And
- 0:32this is the closest these two are going
- 0:34to get for the rest of the video. So,
- 0:35here's what we've got. On this side,
- 0:36this is the brand new M5 Ultra, top of
- 0:39the line right now with 256 gigs of
- 0:42memory, all in one box. And in this
- 0:44corner, we've got two DGX Sparks. Each
- 0:47one of them has 128 GB. That's also 256
- 0:51when you add them together. And they're
- 0:52connected by a one fat 200 gigabit
- 0:56cable. That cable runs something called
- 0:58Rocky.
- 0:59Not that Rocky. RDMA over converged
- 1:02Ethernet. It's like an acronym within an
- 1:04acronym. R is for RDMA, which is remote
- 1:07direct memory access. So, basically,
- 1:09they can write into each other's memory
- 1:11directly. The Mac as configured here is
- 1:13$14,000
- 1:16because it's got the 8 TB drive in
- 1:18there. The Sparks have 4 TB each, so
- 1:22together they're also 8 TB. And the
- 1:24Sparks, when they came out, they were
- 1:25four grand each. Now, they're almost
- 1:27five grand each. So, that's 10. Still
- 1:29cheaper than this.
- 1:31>> Plus the expensive cable. If you're a
- 1:32developer running big models locally or
- 1:35you want to service a small team, you
- 1:37should watch this video.
- 1:39And if you're not, you should also watch
- 1:40this video cuz it's pretty cool stuff.
- 1:42Merlin AI, it's an all-in-one AI tool
- 1:45and they gave my audience a big
- 1:46discount. I keep multiple AI tools
- 1:48around because each one is good at
- 1:50something, but it gets really expensive
- 1:53and bouncing between tabs breaks my
- 1:55focus. Merlin AI puts Chat GPT, Claude,
- 1:58Gemini, and more in one place so I can
- 2:01pick the best one for the moment,
- 2:02whether I'm coding, researching, or
- 2:05writing for a video. Watch this. I click
- 2:06the Merlin AI extension, chat with the
- 2:08web page to summarize what I'm reading,
- 2:11and pull out the important parts. And I
- 2:12even have my choice of models right at
- 2:14my fingertips. If I need something
- 2:16deeper, I turn on deep research, and it
- 2:18builds a clean structured report from
- 2:19multiple sources. And it also has quick
- 2:22modes like web, academic, and Reddit
- 2:24search. If you pay separately, Chat GPT
- 2:26is $20, Claude is $20, Gemini is $20,
- 2:30and that adds up fast. Merlin AI is
- 2:32cheaper because they buy AI API access
- 2:35in bulk. APIs cost less than the $20
- 2:37plans, and most people don't even use
- 2:40$20 worth of API in a month. And here's
- 2:42the discount. I click pricing, continue,
- 2:45it takes me to Stripe, I enter the promo
- 2:47code, and the total drops to $60 for the
- 2:49year. That's basically five bucks a
- 2:51month. I don't know how long this deal
- 2:52will be available, so grab it soon. The
- 2:54link is in the description.
- 2:58Now, two Sparks just don't turn into one
- 3:01big computer magically. vLLM, which is
- 3:05this high-throughput and
- 3:06memory-efficient inference and serving
- 3:08engine for LLMs. This is the software
- 3:11that most Nvidia setups run, and it
- 3:13splits the model across both of them. In
- 3:15this case, it's called tensor parallel.
- 3:17There's other kinds of parallelism, and
- 3:19I talk about that in other videos, but
- 3:21today we're doing tensor parallel. And
- 3:23what that means is that every layer of
- 3:25the model gets cut in half, one half on
- 3:27each Spark. For every single token, both
- 3:30boxes do their half. Then, they swap
- 3:33results over the cable before the next
- 3:35layer can even start. So, that's a lot
- 3:36of back and forth over a cable like
- 3:39this. This is a QSFP cable, it's called.
- 3:41That's just that standard right there.
- 3:44So, why bother with two? Well, modern
- 3:46models today are perfect fit for two of
- 3:49them. DeepSeek V4 flash, for example.
- 3:52That's about 150 160 GB. It doesn't fit
- 3:56on one 128 gig Spark. And yeah, I tried
- 3:59loading it once. It didn't work out so
- 4:01well. Plus, you need extra space for
- 4:03context. Once you load it across two of
- 4:05them, each node is holding just a little
- 4:07bit over 100 gigs. I got it loaded on
- 4:09both of them right now. We got 111 gigs.
- 4:11It's just showing one of them right now
- 4:13out of 128. And that's with extra system
- 4:16stuff going on, too. So, each one gets
- 4:17half the model plus a little room to
- 4:19work. And the Mac just loads the whole
- 4:21thing on one box. Just a quick note, I'm
- 4:24running the same model on both machines,
- 4:26but it's not the same file. Each file is
- 4:29rounded down to four bits or quantized,
- 4:31so it's smaller and faster. The Sparks
- 4:33are running vLLM, like I mentioned.
- 4:35Llama.cpp, a popular tool, works there,
- 4:38too, but vLLM is Nvidia's go-to for this
- 4:42kind of split. On the Mac, I tested both
- 4:44Llama.cpp and MLX. Sometimes one wins
- 4:47and sometimes the other one wins, and
- 4:48I'll point that out as we go.
- 4:52So, I'm going to be running Deep Seek V4
- 4:53Flash Qwen 3.8 Flash next. That's
- 4:57another pretty new one that's requires
- 4:59two Sparks cuz it's large enough. Now,
- 5:01every time you generate tokens, it
- 5:03actually happens in two steps. And I
- 5:05know some of you already know all this,
- 5:07but this is for the newcomers. First,
- 5:09the model reads your whole prompt all at
- 5:11once. That's pure math. It's matrix
- 5:14multiplication. So, whichever box has
- 5:16more compute wins. Usually, that stuff
- 5:19happens on the GPU. So, the more
- 5:20powerful GPUs
- 5:22get the job done faster. And that time
- 5:24is your time to first token, that
- 5:28waiting of the GPU number crunching.
- 5:29Then, it takes that answer and writes
- 5:31one token at a time. And every token is
- 5:34another trip through memory for the
- 5:36model's weights. Weights is just
- 5:38basically a collection of numbers, a
- 5:39huge collection of numbers, many, many
- 5:41gigabytes of collections of numbers.
- 5:42That's the files that you download. So,
- 5:44because for every token we use in the
- 5:46memory, the writing of the answer comes
- 5:48down to memory speed. In other words,
- 5:50the second part of inference is token
- 5:53generation, and it's reliant on memory
- 5:55bandwidth. So, let's take a look at
- 5:57writing. That's that second part. And
- 5:59this is the race from the start, the one
- 6:01I showed you in the beginning. A
- 6:02slightly different
- 6:03version of it, but very close numbers
- 6:05cuz I ran it multiple times. A short
- 6:07prompt with one user, I'm pointing over
- 6:09here cuz I have my charts over here. In
- 6:10DeepSeek, we're getting about 38 tokens
- 6:13per second on the M5 Ultra, and also
- 6:16about 38 tokens per second on the dual
- 6:18Sparks. That's basically a tie. Both are
- 6:21going a pretty decent speed. This is a
- 6:23big model, so that's not bad at all.
- 6:25With Qwen, we have a little bit of a
- 6:27difference there. 45 tokens per second
- 6:29on the Mac, and 38 tokens per second on
- 6:31the dual Sparks. That one I'll give to
- 6:32the Mac. Why did that happen? Well, my
- 6:35best guess is memory. Apple says the M5
- 6:38Ultra has 1.2 terabytes per second of
- 6:42memory bandwidth. That's a lot. Each
- 6:44Spark is rated at 273 GB a second,
- 6:47significantly less. And on top of that,
- 6:49they're syncing over a cable that I
- 6:51measured at about 111 GB, not the 200
- 6:54it's rated for, but still pretty good.
- 6:56But I didn't test how much that cable
- 6:58actually slows things down, so take this
- 7:00with a grain of salt. But the Mac has
- 7:02another trick. The engine you pick for
- 7:04the model matters a lot. For example,
- 7:06DeepSeek on llama.cpp writes at about 40
- 7:10tokens per second. But you take the same
- 7:12model and use MLX instead, and you're
- 7:14getting 53 tokens per second. That's 34%
- 7:18faster just from switching software. And
- 7:20on llama.cpp served exactly the same way
- 7:22the Mac and the Sparks were within 2%.
- 7:25So, yeah, that's faster than Sparks' 38
- 7:27tokens a second, but I only tested
- 7:29DeepSeek on MLX in process. That's the
- 7:32benchmark talking straight to the
- 7:33engine, not served over HTTP like a real
- 7:36chat server. Anyway, I'm collecting
- 7:38information, okay?
- 7:40I'm I'm trying my best to do all the
- 7:42tests that I can. There's a lot of tests
- 7:43that can be done. But so far the Mac is
- 7:45looking pretty good.
- 7:48Now just a quick little primer on
- 7:50tokens. A token is about 3/4 of a word,
- 7:53about there. So a common thing you'll
- 7:55hear is 32K or 32,000 tokens, which is
- 7:58kind of like a typical modern starting
- 8:01point. You start from there and you go
- 8:02up for context size. And that's about
- 8:0424,000 words, which is roughly a
- 8:06100-page document. How long is it going
- 8:08to take you to read a 100-page document,
- 8:10huh?
- 8:12I bet it's not going to take
- 8:1430 seconds or
- 8:16however many seconds. We'll find out.
- 8:18Now,
- 8:20reading or prompt processing, that's the
- 8:23GPU cranking away, remember? I started
- 8:25small at 2,000 tokens, both start fast.
- 8:29But the Sparks turn out to be twice as
- 8:32quick. Quinn, 1.5 seconds on the Mac,
- 8:350.85 seconds on the dual Sparks. Deep
- 8:39Seek, 2 and 1/2 seconds on the Mac, 1
- 8:40and 1/4 seconds on the dual Sparks. You
- 8:43won't care.
- 8:44It's only a 2,000 token prompt. You
- 8:46probably won't even notice. If you
- 8:48sneeze,
- 8:49by the time you wipe your nose it's
- 8:50done. But you might care
- 8:53later. We'll get to that. On Quinn, the
- 8:55Mac's faster writing even wins it back
- 8:57after a couple of hundred tokens. If we
- 8:59make that prompt a little longer, at
- 9:018,000 tokens Deep Seek takes about 4
- 9:04seconds on the Sparks and about 10
- 9:06seconds on the Mac. You double it to
- 9:0816,000 and the Mac's wait doubles also
- 9:12to 21 seconds. Now you're going to have
- 9:14to sneeze a few times and wipe your nose
- 9:17a few times, huh? Yeah, you're going to
- 9:18notice this one. So let's go big.
- 9:21Bigger, okay? I know you some of you
- 9:24going to say like, oh, 32,000 is not
- 9:26big. Okay, let's just take it one step
- 9:27at a time, okay? This time I'm using a
- 9:29real code base, 32,000, boom. There's 14
- 9:32Python files in here. Huh? The DGX
- 9:34Sparks started streaming already, which
- 9:36means they're generating tokens. Now
- 9:38we're past the calculation stage and the
- 9:40Mac is still
- 9:42reading the prompt. It's still
- 9:43processing. We're at 17 seconds for the
- 9:45two DJX Sparks. We're done. And
- 9:49yeah, the Mac is still thinking. Still
- 9:51reading the prompt. Yeah, 50 seconds. I
- 9:53can hear it generating.
- 9:55Yeah. But yeah, that's a big difference
- 9:58there. 300 tokens generated, 32,000
- 10:00prompt tokens. Now, if you look down
- 10:02here, Sparks 17 seconds to Mac Studio's
- 10:0650 seconds. That's about three times the
- 10:09wait. But hold on, don't run away yet
- 10:12buying the Sparks.
- 10:14Some of you might still want to pick up
- 10:16that Mac.
- 10:18I'll let you know why. Now here on this
- 10:19chart, this is where I use the slightly
- 10:21shorter prompt, but the ratio is about
- 10:23the same. So we got 44 seconds on the M5
- 10:25Ultra and 16.5 on the dual Sparks.
- 10:28Quinn, 26 seconds on the Ultra, 11.2 on
- 10:32the Sparks. So 2.4 times faster. Not
- 10:35even close. And after that, it kind of
- 10:37keeps going. Double the prompt, double
- 10:40the wait. Now with the Sparks, I did
- 10:42push them all the way up to 128,000
- 10:45tokens. I didn't run the Mac at 128,000
- 10:47tokens because I kind of saw a pattern
- 10:49here and if it kept its 32K pace, it
- 10:52would have been over 3 minutes of
- 10:53waiting. So it only slows down as the
- 10:55prompt gets longer. The Sparks did it in
- 10:5772 seconds. Still, way faster than the
- 11:01M3 Ultra.
- 11:03Okay?
- 11:04All right. I'm not trying to make
- 11:05excuses. I'm just saying we we still
- 11:07have improvements overall.
- 11:11Well, let's go back to that code base
- 11:13race. Deep seek with the 32K prompt. The
- 11:17Spark gets about a 30-second head start.
- 11:19After that, on Llama CPP, the writing
- 11:21speed is a tie. Quinn's different. The
- 11:23Sparks start about 15 seconds ahead, but
- 11:26the Mac gains a few thousandths of a
- 11:28second in every token. Divide one by the
- 11:30other and by my math, the Mac needs a
- 11:33few thousand tokens in order to catch
- 11:35up. I didn't time that one, it's an
- 11:36estimate. But that's quite a lot, it's a
- 11:38whole new file, basically. So, big
- 11:40inputs favor the Spark. Big output
- 11:44favors the Mac, at least on Quant. This
- 11:46is one of the reasons we have a lot of
- 11:47interest in doing this kind of
- 11:48disaggregated prefill decode using both
- 11:51the Sparks and a Mac. Sparks will do the
- 11:54prefill, Mac will do the decode. I did a
- 11:56video about this, an early early
- 11:58prototype a couple months ago. I'll link
- 12:00to it down below, you can check it out.
- 12:02It's interesting, but this is a new
- 12:03project by this guy Ash Hart. You can
- 12:06check it out, he's on Twitter. He's
- 12:07posting a lot about this stuff now. But,
- 12:10there's one thing that takes the edge
- 12:12off. You always see demos of oh, uh
- 12:16you know, these things when they're
- 12:18kicked off fresh, the prompt processing
- 12:21speed is so different. That's not how it
- 12:22happens in the real world. In a real
- 12:24scenario, in the same session, the
- 12:27servers keep what they've already read.
- 12:29There's a cache. With 16,000 tokens of
- 12:32context, the first ask took 21 seconds
- 12:36on the Mac and 8.3 seconds on the
- 12:38Sparks. All right, we've already seen
- 12:40this stuff. But, the follow-up question,
- 12:423 seconds on the Mac, 1.4 seconds on the
- 12:45Sparks. Yes, the Sparks are still a
- 12:48little bit faster, but you pay for that
- 12:50long read only once per session, not
- 12:53every time. And I recently made a video
- 12:54for members of the channel uh detailing
- 12:57different techniques for how to speed
- 12:58things up. It includes prefix caching.
- 13:01Thanks to the members of the channel, by
- 13:02the way. Really appreciate you.
- 13:03Sometimes they get extra videos uh when
- 13:06I get a chance to record them.
- 13:08Appreciate you all, anyway. And that's
- 13:09how coding tools actually work. Your
- 13:11agent sends the code base once, the
- 13:14server keeps it, every new question only
- 13:16adds a little bit on top. So, that long
- 13:19wait is mostly an initial first question
- 13:21cost, not an every question cost. Where
- 13:24it hurts the most is when the context
- 13:27keeps changing. Obviously, a new repo,
- 13:29big new files, or a long session that
- 13:32outgrows the cache. Those are all
- 13:34possibilities. Another thing I found
- 13:36when running Frontier models is when
- 13:38you're changing the model like from Opus
- 13:405 to Opus 5.5 or Fable 1.1, you have to
- 13:44recalculate all that. And the initial
- 13:46hit is longer usually. But we're not
- 13:48talking about Frontier models now, we're
- 13:49talking about local. All right?
- 13:52Same ideas though apply.
- 13:55Can MLX rescue the Mac on reading speed?
- 13:59Well, on Qwen it's kind of a split.
- 14:01Llama.cpp writes faster and MLX reads
- 14:04faster. MLX gets through 32K in about 19
- 14:08seconds instead of 26. That's still
- 14:10behind the Sparks' 11 seconds. And for
- 14:13Deep Seek, even at MLX's best reading
- 14:15speed, 32,000 tokens would take at least
- 14:1822 seconds, probably closer to 30. Yes,
- 14:21that's also slower than the Sparks.
- 14:23Whether MLX lets the Mac win that race
- 14:25back, I don't know yet. I haven't run
- 14:27Deep Seek yet on MLX with longer prompts
- 14:29or multiple users. Told you, there's a
- 14:31lot of tests to do and I'm crunching
- 14:33through them.
- 14:36More than one user, or the number of
- 14:38concurrencies sometimes it's called, uh
- 14:40you'll see a bigger number quoted
- 14:42usually. And it's the total speed of the
- 14:45throughput of all the tokens generated
- 14:47for all the users. But careful with that
- 14:49number though, because it shows
- 14:51everybody's number. It's everyone's
- 14:52tokens added up. Each person only gets a
- 14:55slice. The server writes for everyone
- 14:57together. And when somebody new shows
- 14:59up, well, the server has to stop and
- 15:02read their prompt too, which uh is going
- 15:04to add to that initial calculation. On
- 15:06the Mac that reading is slower, past
- 15:08about four users, it's spending more
- 15:10time reading than writing. So, at that
- 15:12point everyone's slice shrinks a little
- 15:15bit. And you can see it. On the Mac,
- 15:16Deep Seek peaks at four users, 66 tokens
- 15:20per second there. That's total. Then it
- 15:22drops to 46 tokens per second for eight
- 15:25users. But the Sparks, they keep
- 15:27climbing to 70 tokens per second. And I
- 15:29haven't done 16, so I don't know where
- 15:31it drops off. Uh
- 15:34TBD. Now, give everybody a thousand
- 15:37tokens of chat history and at eight
- 15:40users, the Mac does about 11 tokens per
- 15:43second. The Sparks, 25. And that's
- 15:45total. But the real pain is the wait.
- 15:47Each person on the Mac waits over a
- 15:49minute for the first word and then gets
- 15:50about five tokens a second. On the
- 15:52Sparks, it's about 24 second wait and
- 15:54then about seven tokens per second. So,
- 15:56yeah, sharing these machines,
- 15:59keep it to yourself, all right? You can
- 16:01do it, but use smaller models, maybe.
- 16:05Or, yeah, there's there's different ways
- 16:07of splitting these machines up, but just
- 16:09don't use big models for a lot of
- 16:11people. It's not going to turn out so
- 16:12well. Looking at Quinn here, at eight
- 16:15users, the Sparks put out 120 tokens a
- 16:18second total. That's pretty good. Yeah,
- 16:20that's way better than Deep Seek. The
- 16:22Mac does 66 in Llama CPP and 70 in MLX.
- 16:26And for that, I used OMLX, which is a
- 16:28new tool. Newish. There it is. No more
- 16:31waiting on your Mac. Well, there is some
- 16:33waiting, okay? Really, the Mac is a
- 16:35machine for one person, maybe two. The
- 16:37Sparks are a little bit better for a
- 16:39small team, maybe four people. Eight?
- 16:43You're pushing it. Now, I didn't start
- 16:44out with two Sparks. I went to four and
- 16:46then eight. That was a much bigger
- 16:47project. It was actually experimental,
- 16:49more like uh with breakout cables and
- 16:52extra switches that I had to buy. I made
- 16:53a whole video about it. Compared to
- 16:55that, two Sparks is kind of a perfect
- 16:57setup. It's just one cable and Nvidia's
- 17:00developer site, build.nvidia.com/spark,
- 17:04has some really interesting recipes that
- 17:06are super easy to follow and it just
- 17:08works most of the time. Shows you how to
- 17:10connect two of them, how to run multiple
- 17:12workloads via LLM, everything. Pretty
- 17:14easy. You still have to line things up.
- 17:15I matched the OS, the kernel, the
- 17:17driver, the firmware on both of the
- 17:19machines. Then you have to set up V L M
- 17:21in multi-node with a pile of nickel and
- 17:24rocky environment variables. After that,
- 17:27you have to check the traffic going over
- 17:29the RDMA connection. What you get for
- 17:32all that is CUDA and V L L M. So
- 17:34basically, it's the same kind of stack
- 17:35that you'd run in the cloud. But on the
- 17:38Mac side, this machine isn't just for
- 17:39AI. It's also good at AI, but for
- 17:43example, my daily driver is an M5 Max
- 17:46MacBook Pro. Everything that I do on
- 17:48that machine, I can do it on the M5
- 17:50Ultra, but faster. That includes running
- 17:53large models. And for everyday work,
- 17:56it's zero setup, basically. I can
- 17:58offload a lot of stuff to it, like
- 18:00rendering my videos, for example, for
- 18:02this channel. And um sometimes the Mac
- 18:05gets a little bogged down. I run a lot
- 18:07of stuff on it. And right now, I'm using
- 18:0988 GB of memory. Yeah, it gets bogged
- 18:12down a little bit, even with 128 gigs of
- 18:15memory. So it's nice to be able to
- 18:16offload stuff very easily. And here is
- 18:19another thing. There they go, full GPU
- 18:22utilization on both machines. Or yeah,
- 18:25you only see one here, but they're both
- 18:27working on the sparks. And the Mac is
- 18:29using 100% of GPU as well. Oh boy.
- 18:34Now, during my M5 Ultra first look, I
- 18:36noticed that the machine was pretty
- 18:38toasty. And yeah, it is. But sparks also
- 18:41have been known to be toasty. You know,
- 18:43we're not alone here. Right now, I'm
- 18:45doing a very heavy load on the cluster
- 18:48here and the Mac Studio. And this is
- 18:50what I'm seeing. This is nuts. All
- 18:53right. Uh
- 18:54I don't want to pop a breaker here, but
- 18:55I might. 434 W being used by the M5
- 18:59Ultra and 410 by the spark cluster. So
- 19:03we're very close. I should say at idle,
- 19:06it's a very different story. They are
- 19:08pretty warm now.
- 19:09And I'm hearing noise coming out of
- 19:12everywhere. Just like noise throughout.
- 19:16They feel about the same actually. Yeah,
- 19:18they're both pretty orange. About 49°
- 19:22to 50° on the hottest part of the Mac
- 19:25Studio, about 48°
- 19:28on the Sparks. Let's take a look at the
- 19:30back. Ooh.
- 19:3256° on the grill, the back of the Mac
- 19:35Studio. And wow, 58 60 63 I saw in there
- 19:41on the Sparks. So yeah, both get pretty
- 19:43toasty. For AI on the Mac, there is a
- 19:45little setup depending on which stack
- 19:47you want, Llama CPP or MLX. OMLX is
- 19:51pretty easy. There's also newer ones
- 19:52like native without the E. And there is
- 19:55Splash. Some of these are so new I
- 19:57haven't even tried them yet. These are
- 19:59basically home brew installs. So is 256
- 20:02gigs across two boxes the same as 256 in
- 20:06one? Well, it holds the same model. It
- 20:08just reads a lot faster. But before you
- 20:10go out and drop 10 grand or more on one
- 20:13of these setups, look at your own work.
- 20:15How big are your prompts? How long are
- 20:17the answers to your prompts? You want to
- 20:18learn more about clustering the Sparks,
- 20:20watch this video here. Clustering Mac
- 20:22Studios, watch this video here. Thanks
- 20:25for watching and I'll see you next time.
About this transcript
This page contains the full transcript of M5 Ultra vs 2 DGX Sparks… The Number You're Not Looking At by Alex Ziskind, generated from the public captions YouTube serves with the video. The transcript has 3,740 words across 533 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.