YouTube2Text

This Is What Happens When You CRUSH An AI Video Model — Transcript

by Alex Ziskind · 3,328 words · 499 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Here's a thing that nobody tells you. If
  2. 0:01you're running an AI video model
  3. 0:03locally, you're running it quantized.
  4. 0:05You just are. ComfyUI, probably the most
  5. 0:08popular tool for generating videos, it
  6. 0:10has you a GGUE version. And most
  7. 0:12workflows assume you grab the Q4 on day
  8. 0:14one and never looked back. Basically
  9. 0:16means the model weights are quantized
  10. 0:17down to four bits. But, nobody actually
  11. 0:19tells you what you give up to get this.
  12. 0:21So, I wanted to check it out. I took two
  13. 0:24video models and I ran each one all the
  14. 0:27way down an eight-step ladder from full
  15. 0:29precision at FP16 down all the way to
  16. 0:32two bits. Some of these look okay at two
  17. 0:34bits, actually. But, we'll go through
  18. 0:35that. Wen 2.2 14 billion parameters
  19. 0:39text-to-video and LTX 2.3 22 billion
  20. 0:43parameter model. This one is a bit
  21. 0:44spicier. And that's because it generates
  22. 0:47the video and the audio together, which
  23. 0:49means now the picture and the sound can
  24. 0:51break separately. I ran the same prompt,
  25. 0:54the same seed, the same settings every
  26. 0:57single time. And as we walk down the
  27. 0:59ladder, only one thing actually changes,
  28. 1:02just the quantization. And all of this
  29. 1:04on a custom Linux box that I built
  30. 1:06running an RTX Pro 6000 Blackwell.
  31. 1:09That's 96 gigs of RAM, which comfortably
  32. 1:12fit all the precision models that I
  33. 1:14have. Now, here's what I didn't expect.
  34. 1:16Two completely different architectures
  35. 1:19built by two completely different teams
  36. 1:21and they eventually kind of fall apart
  37. 1:23in the same way. Let's go through it.
  38. 1:28So, this is the reference. For Wen, I
  39. 1:30got FP16.
  40. 1:33>> [music]
  41. 1:33>> Looking good, looking good. For LTX,
  42. 1:35it's basically the same thing. It's
  43. 1:36BF16, just slightly different
  44. 1:39quantization, but they're both full
  45. 1:41format or full precision. And this is
  46. 1:43what everything else we're going to
  47. 1:44compare gets [music] compared against.
  48. 1:46For Wen, I ran five prompts. Some are
  49. 1:49simple objects, some are humans, some
  50. 1:52busy scenes. For LTX, I measured a
  51. 1:55couple more things. Whether the words
  52. 1:56stayed correct
  53. 1:58>> Does this support the new Phantom 5090?
  54. 2:00>> and whether the audio still sounded
  55. 2:01intelligible.
  56. 2:02>> Did you try turning it off and on again?
  57. 2:04>> Then I scored the videos with some
  58. 2:05common benchmarks like SSIM, LPIPS, and
  59. 2:08prompt alignment tests. Plus the test of
  60. 2:10my actual eyeballs and ear holes. But,
  61. 2:13enough about my holes. Let's see the
  62. 2:15next level so we can actually get
  63. 2:16something to compare it against, shall
  64. 2:18we? So, here's when FP16, the whole
  65. 2:20model is 27 GB, quite chunky. And I do
  66. 2:23see pretty good character consistency.
  67. 2:26Looks pretty good. Except the very
  68. 2:27beginning, two frames or so, where it
  69. 2:29looks like an old, I don't know, TV CRT
  70. 2:32screen. But, after that it balances out
  71. 2:34and the character remains pretty
  72. 2:36consistent throughout. The leaves look
  73. 2:38normal. Let's see LTX.
  74. 2:40>> Does this support the new Phantom 5090?
  75. 2:42>> Oh, yeah. 24 lanes.
  76. 2:45Wait.
  77. 2:47THEY LIED TO US AGAIN!
  78. 2:50>> DOES THIS SUPPORT THE NEW PHANTOM 5090?
  79. 2:51>> That's crazy. Yeah, my prompt had him
  80. 2:54doing something funny, but uh that was
  81. 2:56just nuts. That's me and Dan at Micro
  82. 2:58Center, but our voices are just
  83. 3:01not at all the same. Although,
  84. 3:03the voices are very clear, very
  85. 3:05consistent. And after analyzing both of
  86. 3:07the image quality and the audio quality
  87. 3:10on this one, it's kind of hard to see,
  88. 3:12but these are our baselines and I will
  89. 3:13go through these very soon here.
  90. 3:15>> This tiny thing has quietly become one
  91. 3:18of the most useful tools I carry. As
  92. 3:19someone who constantly is bouncing
  93. 3:21between meetings, conferences, and
  94. 3:22content, I need a better way to keep
  95. 3:24track of the details. The hard part for
  96. 3:25me is I cannot stay fully present,
  97. 3:28listen, and get footage, and take clean
  98. 3:30notes all at the same time. So, now I
  99. 3:32use Plaud Note Pen S like a second brain
  100. 3:35for meetings, interviews, and event
  101. 3:37days. Most note-taking setups still
  102. 3:39leave me doing the worst part
  103. 3:40afterwards, sorting out who said what
  104. 3:42and what actually matters. This thing is
  105. 3:44actually really simple to use. One click
  106. 3:46starts the recording. The physical
  107. 3:48button helps mark key moments. And then
  108. 3:50it organizes everything for me after. It
  109. 3:53can capture up to 20 hours non-stop and
  110. 3:56then turns that into transcripts with
  111. 3:58speaker labels, clean summaries, and
  112. 4:00actionable to-do lists instead of one
  113. 4:02giant recording that I never revisit. I
  114. 4:04also really like ask Plod. It lets me
  115. 4:06pull up past information or think
  116. 4:07through next steps without digging
  117. 4:09through everything manually. The big win
  118. 4:11for me is less mental load and better
  119. 4:13follow through. And I'm saying that as
  120. 4:15someone who's been using Plod since
  121. 4:172024. I went through all the different
  122. 4:19versions. Obviously, use it where
  123. 4:21recording is appropriate and always get
  124. 4:23permission, but as a workflow tool, this
  125. 4:25thing has really been useful for me. Use
  126. 4:27my code Alex15off for 15% off. There's
  127. 4:30also a Prime Day deal plus 30-day free
  128. 4:32return policy. Check the link in the
  129. 4:34description.
  130. 4:37Now we cut to half the bits, and this is
  131. 4:39where it gets weird. Eight bits per
  132. 4:40weight, but specifically the FP8 format,
  133. 4:44so floating points. Half the bytes of
  134. 4:46FP16. FP8 is natively supported on this
  135. 4:49GPU, so that should be a slam dunk. FP8
  136. 4:52is 14 GB on disk. How's the quality? So
  137. 4:54here I started with the simplest thing
  138. 4:56that I rendered, and this is a just
  139. 4:57glass of water just sitting there.
  140. 4:59Camera's moving around just a little
  141. 5:01bit. Whatever changes you see here is
  142. 5:03just the model basically rendering
  143. 5:05things differently, even though the seed
  144. 5:06is exactly the same. The FP16 version
  145. 5:08looks just a little bit more realistic
  146. 5:10to me except for what are these lines
  147. 5:12walking across the table? I don't know,
  148. 5:14it might be raining or something or
  149. 5:15something else is moving in the room. On
  150. 5:17the right, it's a little choppier and a
  151. 5:19little bit more blurry, but if it wasn't
  152. 5:21next to the FP16 image, it would just
  153. 5:23pass. Here's how things change. FP8 is
  154. 5:25already off the baseline here. This is
  155. 5:27the static detail video. 0.19 on LPIPS,
  156. 5:30and LPIPS is basically like my eyeballs
  157. 5:34in software. If it was zero, that would
  158. 5:35mean it's identical, and the higher the
  159. 5:37numbers, it means it's more different
  160. 5:39from the original. Now here's the
  161. 5:40surprise. The very next level down, it's
  162. 5:43not really down, it's parallel, I'd say.
  163. 5:46It's Q80. So, it's still eight bits per
  164. 5:49weight, but in a different format. It's
  165. 5:52integer eight with K quant grouping. On
  166. 5:54the glass video, Q8 lands at 0.07. It's
  167. 5:58almost identical to full precision,
  168. 6:00while FP8 is about twice that much. And
  169. 6:02yeah, look at that. The glass actually
  170. 6:05looks almost identical to the FP16 in
  171. 6:08the video itself, according to my balls
  172. 6:10in the eye. Q8, pretty much the same
  173. 6:12disk size as FP8, but better fidelity.
  174. 6:15And guess what? The average across all
  175. 6:18the five prompts here, FP8 is still
  176. 6:21about twice as far from baseline as Q8.
  177. 6:24So, FP8, the format with hardware
  178. 6:26acceleration that everybody said was the
  179. 6:27future, it drifts further and further
  180. 6:30from full precision than the older int8
  181. 6:33approach. Same number of bits, the
  182. 6:34difference is just where it spends that
  183. 6:37precision. And here is the kicker, the
  184. 6:39red car, a simple motion prompt. There's
  185. 6:42FP16, at FP8, the car drives backwards.
  186. 6:45Not in any other quantization does this
  187. 6:48happen. Just FP8. And this time, it's
  188. 6:51not just slightly different rendering,
  189. 6:53it's just wrong. So, more bits is not
  190. 6:55the same as better. The format does more
  191. 6:57work than the bit count does. Remember
  192. 6:59that, because we're about to watch the
  193. 7:01same exact surprise happen in a
  194. 7:03completely different model.
  195. 7:07FP8 again, same hardware acceleration,
  196. 7:10same expected slam dunk.
  197. 7:11>> Does this support the new Phantom 5090?
  198. 7:13>> Looks about the same, right? Watch the
  199. 7:15face and listen at the end.
  200. 7:17>> Oh, yeah, 24 lanes.
  201. 7:20Wait.
  202. 7:21THEY LIED TO US AGAIN.
  203. 7:24>> [laughter]
  204. 7:24>> DID YOU CATCH THAT?
  205. 7:27WHISPER picks it up. Word error rate on
  206. 7:30FP8 is 0.18. And basically, word error
  207. 7:33rate is just how many words drifted from
  208. 7:35the baseline transcript. Zero is
  209. 7:38identical. This is the only quantization
  210. 7:41between BF16 and Q3KM with a non-zero
  211. 7:45WER, word error rate. The model added a
  212. 7:48laugh that does not exist in the
  213. 7:50baseline audio. SSIM point 87 LPIPS is
  214. 7:55at point 07, pretty good. By the way,
  215. 7:57the top three are for video quality
  216. 8:00comparisons, and the bottom two for LTX
  217. 8:02only are for audio. And here we have mel
  218. 8:04spectrogram MSE 26.7.
  219. 8:07And for video, by the way, the SSIM is
  220. 8:10the pixel match. So, one would be
  221. 8:13identical. Obviously, BF16 is identical
  222. 8:16to itself, so that's one. And here on
  223. 8:18FP8, we're sliding down to point 8. And
  224. 8:21mel MSC down here is the audio
  225. 8:23fingerprint. So, zero would be identical
  226. 8:25in this case. Anything higher than that
  227. 8:27is drift.
  228. 8:28>> Does this support the new Phantom 5090?
  229. 8:30>> Oh, yeah, 24 lanes.
  230. 8:33Wait.
  231. 8:36Stealing
  232. 8:37our words again.
  233. 8:38>> So, compare that to Q80,
  234. 8:41one level down on the ladder, sort of
  235. 8:44more like
  236. 8:45horizontal. Same eight bits per weight,
  237. 8:47of course. Different format, though.
  238. 8:49Look at the difference between BF16 and
  239. 8:51Q80. There's a huge difference in the
  240. 8:53amount of space it takes. 46 GB versus
  241. 8:5623, but they look identical. Even the
  242. 8:58motion blur in those exact moments, in
  243. 9:01those frames, is exactly the same. But
  244. 9:03FP8
  245. 9:05not. And Q8 is better on every single
  246. 9:08metric, both video and audio. So, FP8
  247. 9:11underperformed in one
  248. 9:14one one, I don't know how to say that
  249. 9:16properly, don't ask me. And it also
  250. 9:20underperformed in LTX. Two different
  251. 9:23model architectures, two different
  252. 9:24teams, two different training runs.
  253. 9:27Weird, right? Maybe not. So, if you've
  254. 9:29got a choice between FP8 and Q8 for a
  255. 9:32video model right now, take Q8. I'm
  256. 9:35going to skip over levels between Q6 and
  257. 9:39Q5 because they do get a little bit
  258. 9:42worse, but there's no real big jumps
  259. 9:45here until you get to Q4. In fact, with
  260. 9:47these glasses, Q4 looks pretty good to
  261. 9:49me, too. Same thing with the car. I see
  262. 9:52slight differences with the detail of
  263. 9:54the car itself, but overall, the motion,
  264. 9:58the clarity looks pretty good. Let's
  265. 10:00take a look at the lady here. Q6 and Q5
  266. 10:03both 12 and 11 GB respectively on disk.
  267. 10:06It definitely looks like the same exact
  268. 10:08lady. So, what happens at Q4? Can we
  269. 10:10save more space and actually get away
  270. 10:12with it?
  271. 10:13>> [music]
  272. 10:15>> Q4, 9 GB on disk. This is the kind of
  273. 10:19thing you'd run on a 24 GB consumer GPU.
  274. 10:22And this is where most of the community
  275. 10:24lives. And if you watch this in
  276. 10:26isolation, you'd probably think, "Huh,
  277. 10:28it's fine." Sure, the face is a little
  278. 10:31bit fuzzier and a little bit not as
  279. 10:33crisp, but it's passable. The forest
  280. 10:36still moves, but if you compare it, L
  281. 10:38pips says, "We are
  282. 10:41at 0.34." That's further away than FP8.
  283. 10:45For a complex scene, we're at 0.35. What
  284. 10:48about that glass of water? Not so bad
  285. 10:50here, 0.2. For L pips, that character
  286. 10:52consistency is getting there. Let me
  287. 10:54show you. Notice anything different
  288. 10:55about her hair? And also, it kind of
  289. 10:57doesn't look like the same woman
  290. 10:58anymore. Here's the complex scene. It's
  291. 11:00a little bit harder to tell here. Tokyo
  292. 11:02night market, quarter disk size of FP16,
  293. 11:06but it still holds together. This is
  294. 11:08kind of like a sweet spot over here, and
  295. 11:10you can stop here if you don't have any
  296. 11:13reason to get any smaller. Red car looks
  297. 11:15fine, and do comment down below if you
  298. 11:18notice any weirdnesses that I didn't
  299. 11:20notice. We save a ton of space in LTX
  300. 11:232.3. The Q4 is only 14 GB here. There's
  301. 11:26some extra weirdness going on here with
  302. 11:28the face there and my eyes and Dan's
  303. 11:31eyes, and the unexpected reaction makes
  304. 11:33its reappearance.
  305. 11:36But look at the mel spectrogram here. We
  306. 11:38just jumped to 46.9.
  307. 11:41One level up at Q5, it was just 9.9. So,
  308. 11:46the audio fidelity just dropped roughly
  309. 11:48five times in a single step. Now, check
  310. 11:50the top row, the video metrics. SSIM is
  311. 11:54.88, basically right where it was. LPIPS
  312. 11:56is .07, the picture is right about where
  313. 11:59it was for FP8. So, it's definitely
  314. 12:02gradually getting worse here. And Q4
  315. 12:04shows a big change, at least in the
  316. 12:06objective measurements, but not as much
  317. 12:08as audio. So, that's finding number two,
  318. 12:10audio degrades before video here, even
  319. 12:13though it might not be perceived as such
  320. 12:16or might not be as noticeable as video.
  321. 12:18If you're only watching the picture,
  322. 12:19you're going to miss that. And why does
  323. 12:21this matter? The picture degrading is
  324. 12:23something you might catch on a rewatch,
  325. 12:25but the audio degrading is something the
  326. 12:27viewer hears immediately. You ever watch
  327. 12:30a YouTube video with terrible audio? I
  328. 12:32hope I hope my audio is actually decent
  329. 12:34here. Can't be perfect, but I try. But,
  330. 12:37you can put up with pretty bad video.
  331. 12:40However, if you hear terrible audio,
  332. 12:42people will just click away right away.
  333. 12:44Very different tolerance levels. So,
  334. 12:45models like LTX, the ones that support
  335. 12:48audio, have to take extra special care.
  336. 12:50>> [music]
  337. 12:54>> All right, two bits per weight. This is
  338. 12:565 GB for when. This is the bottom of the
  339. 13:00ladder, and it shows. That glass is
  340. 13:04>> [laughter]
  341. 13:05>> uh pretty bad looking. There's even
  342. 13:07differences frame to frame. Look at the
  343. 13:08lady. Woo,
  344. 13:10that's terrible. And Q3 is very similar
  345. 13:12to Q4, where it does mess with the hair
  346. 13:16quite a bit from the original and
  347. 13:18doesn't look at all like the original.
  348. 13:20But, Q2 is a whole different level
  349. 13:22that's completely unusable here at this
  350. 13:23point. Look at this character
  351. 13:25consistency chart. LPIPS is 0.54.
  352. 13:29Tokyo market, the complex scene, we're
  353. 13:32at 0.57, even worse. Both up about 60%
  354. 13:36from Q4. Now, LTX at the same two bits
  355. 13:39does the exact same thing. Look at that
  356. 13:41first frame, looks pretty good, right?
  357. 13:42This is image to video, by the way, in
  358. 13:44case you didn't know. Small model, 8.3
  359. 13:46GB,
  360. 13:47but look what happens if I play this.
  361. 13:49>> Does this support the new Phantom 5090?
  362. 13:51Oh, yeah, 24 lanes.
  363. 13:54Wait.
  364. 13:57THEY LIED TO US [screaming] AGAIN!
  365. 13:59>> OH MY GOD, THAT WAS just like tugs at
  366. 14:01the heartstrings. The audio is very
  367. 14:03different. It sounds robotic until he
  368. 14:05screams, then it sounds like more of a
  369. 14:07somebody trying to get an Oscar. But
  370. 14:09look how a terribly video quality is at
  371. 14:12every frame. Anything that's moving is
  372. 14:14basically completely destroyed and
  373. 14:16smashed.
  374. 14:20Whoa. Look at the transcript here. Word
  375. 14:23error rate here for Q3 and Q2, it went
  376. 14:26from they lied to us again to they lied
  377. 14:29to us again. Caps in the middle of the
  378. 14:31sentence. These are not the prompts, by
  379. 14:32the way, these are the transcripts for
  380. 14:35that uh whisper test that are generated
  381. 14:37from the video. And for LTX, it's not
  382. 14:39just this one skit. LTX again, this is
  383. 14:42BF16, FP8 looking like a very different
  384. 14:45person. Q8, Q6, Q5
  385. 14:49look like BF16 pretty much, pretty much
  386. 14:51the same person. But the leaves are
  387. 14:53wrong, even in BF16. Those are not maple
  388. 14:56leaves, I don't know what that is, like
  389. 14:58a starfish in the shape of a leaf or a
  390. 15:00leaf in the shape of a starfish, I
  391. 15:02guess. The shirt is also very difficult
  392. 15:05here. Every single one of these has a
  393. 15:07different shirt. Q5 and Q4 are kind of
  394. 15:09the same. And look what happens with Q3.
  395. 15:12We have a different person here.
  396. 15:16>> Can you tell which one of me is the real
  397. 15:18one?
  398. 15:19>> She's not even looking at the camera.
  399. 15:20Forget about Q2.
  400. 15:24>> Can you tell which one of me is the real
  401. 15:26one?
  402. 15:27>> Yeah, it's not you. Actually, it's none
  403. 15:28of them, but some of these could pass
  404. 15:30for real one.
  405. 15:34>> Can you tell which one of me is the real
  406. 15:36one?
  407. 15:36>> Did you notice the audio difference
  408. 15:38between BF16 and Q2? Huge difference.
  409. 15:42This one sounds real. Here's Q2 again.
  410. 15:44>> Can you tell which one of me is the real
  411. 15:46one?
  412. 15:46>> Totally fake at this point, right? This
  413. 15:48is what her charts look like. Pretty
  414. 15:50gradual SSIM, LPIPS also pretty gradual,
  415. 15:53but you know, it gets up there. They all
  416. 15:55have a little bump on FP8, which makes
  417. 15:57it equivalent to Q4/Q3.
  418. 16:00Let's check out the tech support scene.
  419. 16:02BF16, listen to the audio, too.
  420. 16:05>> Did you try turning it off and on again?
  421. 16:09>> I AM THE IT DEPARTMENT.
  422. 16:15>> NOT SURE WHAT'S GOING ON over there.
  423. 16:17Something's going on. With FP8, the lady
  424. 16:19is totally different.
  425. 16:20>> Try turning it off and on again.
  426. 16:24>> I AM THE IT DEPARTMENT.
  427. 16:31>> [laughter]
  428. 16:32>> LOOK AT HER. SHE'S NOT IMPRESSED. And
  429. 16:34the lady looks different in pretty much
  430. 16:36every single one of these. Her shirt
  431. 16:38color is different, the wall decorations
  432. 16:41are different, what's on the computer
  433. 16:43screen is different. But once we get to
  434. 16:45the 11 GB Q3, things just start melting
  435. 16:48down. Check this out.
  436. 16:49>> Did you try turning it off and on again?
  437. 16:54>> I AM THE IT DEPARTMENT.
  438. 16:59>> KIND OF LOOKS LIKE MR. Smith from The
  439. 17:01Matrix. You know, she's trying to push
  440. 17:03his buttons, and she's winning. Q2.
  441. 17:06>> Did you try turning it off and on again?
  442. 17:10>> I am
  443. 17:12the
  444. 17:14IT DEPARTMENT.
  445. 17:16>> DID YOU TRY
  446. 17:17>> WOW, that one is like way off the rails.
  447. 17:19First of all, the lady sounds robotic.
  448. 17:21The video is just totally destroyed.
  449. 17:26So, faces are still the hardest part.
  450. 17:28Whisper for Q2 and Mel spectrogram for
  451. 17:32Q2 all just are off the chart terrible.
  452. 17:36So, by now it looks like two bits just
  453. 17:38wrecks everything. But, here's what I
  454. 17:40did not expect. The glass of water scene
  455. 17:43SSIM is at .74 at two bits per weight.
  456. 17:47It doesn't look great when I'm examining
  457. 17:48with my balls of eye, but the numbers
  458. 17:51are saying it's okay. So, take these
  459. 17:54tests also with a grain of salt. I
  460. 17:56didn't come up with these tests. These
  461. 17:57are pretty much standard tests. So,
  462. 18:00yeah. Haven't shown you this one yet,
  463. 18:02the marble rolling down the ramp. And
  464. 18:04even though each one of these frames by
  465. 18:06themselves might look okay, the physics
  466. 18:09of this thing is just terrible at every
  467. 18:12single weight level. Even at full
  468. 18:15precision here, the ball just doesn't
  469. 18:17roll like a real ball and then two balls
  470. 18:20get merged. That's just the model, not
  471. 18:23the quantization. It's just wrong, not
  472. 18:26necessarily worse. Although, at Q2, you
  473. 18:29could pretty much say it's worse. So,
  474. 18:30finding number three I'd say is that
  475. 18:33there is no universal cliff. And that's
  476. 18:35true in both models. If the thing you're
  477. 18:38making is a single hard object in
  478. 18:40motion, you can probably ride this all
  479. 18:43the way down. But, if you're making
  480. 18:44humans in it, the floor is a lot higher.
  481. 18:47Bottom line,
  482. 18:48run Q4 unless you've got a reason not
  483. 18:51to. And the sweet spot doesn't move when
  484. 18:53you switch model architectures. But,
  485. 18:54also if your output has audio, listen to
  486. 18:57it at Q4. The picture will trick you,
  487. 18:59but the voice will not. And if you
  488. 19:01remember one thing from this whole
  489. 19:02video, let it be this. Format matters
  490. 19:06more than bit count. The full ladders
  491. 19:08are linked down below, every clip and
  492. 19:09every chart. And thanks to this rig, I
  493. 19:11was able to generate these videos pretty
  494. 19:14fast. If you want to know whether this
  495. 19:1596 GB RTX Pro 6000 rig was actually
  496. 19:18worth it, that video is over here.
  497. 19:21Thanks for watching, [music] and I'll
  498. 19:22see you next time.
  499. 19:29>> [music]

About this transcript

This page contains the full transcript of This Is What Happens When You CRUSH An AI Video Model by Alex Ziskind, generated from the public captions YouTube serves with the video. The transcript has 3,328 words across 499 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.