This Is What Happens When You CRUSH An AI Video Model — Transcript
Full transcript
- 0:00Here's a thing that nobody tells you. If
- 0:01you're running an AI video model
- 0:03locally, you're running it quantized.
- 0:05You just are. ComfyUI, probably the most
- 0:08popular tool for generating videos, it
- 0:10has you a GGUE version. And most
- 0:12workflows assume you grab the Q4 on day
- 0:14one and never looked back. Basically
- 0:16means the model weights are quantized
- 0:17down to four bits. But, nobody actually
- 0:19tells you what you give up to get this.
- 0:21So, I wanted to check it out. I took two
- 0:24video models and I ran each one all the
- 0:27way down an eight-step ladder from full
- 0:29precision at FP16 down all the way to
- 0:32two bits. Some of these look okay at two
- 0:34bits, actually. But, we'll go through
- 0:35that. Wen 2.2 14 billion parameters
- 0:39text-to-video and LTX 2.3 22 billion
- 0:43parameter model. This one is a bit
- 0:44spicier. And that's because it generates
- 0:47the video and the audio together, which
- 0:49means now the picture and the sound can
- 0:51break separately. I ran the same prompt,
- 0:54the same seed, the same settings every
- 0:57single time. And as we walk down the
- 0:59ladder, only one thing actually changes,
- 1:02just the quantization. And all of this
- 1:04on a custom Linux box that I built
- 1:06running an RTX Pro 6000 Blackwell.
- 1:09That's 96 gigs of RAM, which comfortably
- 1:12fit all the precision models that I
- 1:14have. Now, here's what I didn't expect.
- 1:16Two completely different architectures
- 1:19built by two completely different teams
- 1:21and they eventually kind of fall apart
- 1:23in the same way. Let's go through it.
- 1:28So, this is the reference. For Wen, I
- 1:30got FP16.
- 1:33>> [music]
- 1:33>> Looking good, looking good. For LTX,
- 1:35it's basically the same thing. It's
- 1:36BF16, just slightly different
- 1:39quantization, but they're both full
- 1:41format or full precision. And this is
- 1:43what everything else we're going to
- 1:44compare gets [music] compared against.
- 1:46For Wen, I ran five prompts. Some are
- 1:49simple objects, some are humans, some
- 1:52busy scenes. For LTX, I measured a
- 1:55couple more things. Whether the words
- 1:56stayed correct
- 1:58>> Does this support the new Phantom 5090?
- 2:00>> and whether the audio still sounded
- 2:01intelligible.
- 2:02>> Did you try turning it off and on again?
- 2:04>> Then I scored the videos with some
- 2:05common benchmarks like SSIM, LPIPS, and
- 2:08prompt alignment tests. Plus the test of
- 2:10my actual eyeballs and ear holes. But,
- 2:13enough about my holes. Let's see the
- 2:15next level so we can actually get
- 2:16something to compare it against, shall
- 2:18we? So, here's when FP16, the whole
- 2:20model is 27 GB, quite chunky. And I do
- 2:23see pretty good character consistency.
- 2:26Looks pretty good. Except the very
- 2:27beginning, two frames or so, where it
- 2:29looks like an old, I don't know, TV CRT
- 2:32screen. But, after that it balances out
- 2:34and the character remains pretty
- 2:36consistent throughout. The leaves look
- 2:38normal. Let's see LTX.
- 2:40>> Does this support the new Phantom 5090?
- 2:42>> Oh, yeah. 24 lanes.
- 2:45Wait.
- 2:47THEY LIED TO US AGAIN!
- 2:50>> DOES THIS SUPPORT THE NEW PHANTOM 5090?
- 2:51>> That's crazy. Yeah, my prompt had him
- 2:54doing something funny, but uh that was
- 2:56just nuts. That's me and Dan at Micro
- 2:58Center, but our voices are just
- 3:01not at all the same. Although,
- 3:03the voices are very clear, very
- 3:05consistent. And after analyzing both of
- 3:07the image quality and the audio quality
- 3:10on this one, it's kind of hard to see,
- 3:12but these are our baselines and I will
- 3:13go through these very soon here.
- 3:15>> This tiny thing has quietly become one
- 3:18of the most useful tools I carry. As
- 3:19someone who constantly is bouncing
- 3:21between meetings, conferences, and
- 3:22content, I need a better way to keep
- 3:24track of the details. The hard part for
- 3:25me is I cannot stay fully present,
- 3:28listen, and get footage, and take clean
- 3:30notes all at the same time. So, now I
- 3:32use Plaud Note Pen S like a second brain
- 3:35for meetings, interviews, and event
- 3:37days. Most note-taking setups still
- 3:39leave me doing the worst part
- 3:40afterwards, sorting out who said what
- 3:42and what actually matters. This thing is
- 3:44actually really simple to use. One click
- 3:46starts the recording. The physical
- 3:48button helps mark key moments. And then
- 3:50it organizes everything for me after. It
- 3:53can capture up to 20 hours non-stop and
- 3:56then turns that into transcripts with
- 3:58speaker labels, clean summaries, and
- 4:00actionable to-do lists instead of one
- 4:02giant recording that I never revisit. I
- 4:04also really like ask Plod. It lets me
- 4:06pull up past information or think
- 4:07through next steps without digging
- 4:09through everything manually. The big win
- 4:11for me is less mental load and better
- 4:13follow through. And I'm saying that as
- 4:15someone who's been using Plod since
- 4:172024. I went through all the different
- 4:19versions. Obviously, use it where
- 4:21recording is appropriate and always get
- 4:23permission, but as a workflow tool, this
- 4:25thing has really been useful for me. Use
- 4:27my code Alex15off for 15% off. There's
- 4:30also a Prime Day deal plus 30-day free
- 4:32return policy. Check the link in the
- 4:34description.
- 4:37Now we cut to half the bits, and this is
- 4:39where it gets weird. Eight bits per
- 4:40weight, but specifically the FP8 format,
- 4:44so floating points. Half the bytes of
- 4:46FP16. FP8 is natively supported on this
- 4:49GPU, so that should be a slam dunk. FP8
- 4:52is 14 GB on disk. How's the quality? So
- 4:54here I started with the simplest thing
- 4:56that I rendered, and this is a just
- 4:57glass of water just sitting there.
- 4:59Camera's moving around just a little
- 5:01bit. Whatever changes you see here is
- 5:03just the model basically rendering
- 5:05things differently, even though the seed
- 5:06is exactly the same. The FP16 version
- 5:08looks just a little bit more realistic
- 5:10to me except for what are these lines
- 5:12walking across the table? I don't know,
- 5:14it might be raining or something or
- 5:15something else is moving in the room. On
- 5:17the right, it's a little choppier and a
- 5:19little bit more blurry, but if it wasn't
- 5:21next to the FP16 image, it would just
- 5:23pass. Here's how things change. FP8 is
- 5:25already off the baseline here. This is
- 5:27the static detail video. 0.19 on LPIPS,
- 5:30and LPIPS is basically like my eyeballs
- 5:34in software. If it was zero, that would
- 5:35mean it's identical, and the higher the
- 5:37numbers, it means it's more different
- 5:39from the original. Now here's the
- 5:40surprise. The very next level down, it's
- 5:43not really down, it's parallel, I'd say.
- 5:46It's Q80. So, it's still eight bits per
- 5:49weight, but in a different format. It's
- 5:52integer eight with K quant grouping. On
- 5:54the glass video, Q8 lands at 0.07. It's
- 5:58almost identical to full precision,
- 6:00while FP8 is about twice that much. And
- 6:02yeah, look at that. The glass actually
- 6:05looks almost identical to the FP16 in
- 6:08the video itself, according to my balls
- 6:10in the eye. Q8, pretty much the same
- 6:12disk size as FP8, but better fidelity.
- 6:15And guess what? The average across all
- 6:18the five prompts here, FP8 is still
- 6:21about twice as far from baseline as Q8.
- 6:24So, FP8, the format with hardware
- 6:26acceleration that everybody said was the
- 6:27future, it drifts further and further
- 6:30from full precision than the older int8
- 6:33approach. Same number of bits, the
- 6:34difference is just where it spends that
- 6:37precision. And here is the kicker, the
- 6:39red car, a simple motion prompt. There's
- 6:42FP16, at FP8, the car drives backwards.
- 6:45Not in any other quantization does this
- 6:48happen. Just FP8. And this time, it's
- 6:51not just slightly different rendering,
- 6:53it's just wrong. So, more bits is not
- 6:55the same as better. The format does more
- 6:57work than the bit count does. Remember
- 6:59that, because we're about to watch the
- 7:01same exact surprise happen in a
- 7:03completely different model.
- 7:07FP8 again, same hardware acceleration,
- 7:10same expected slam dunk.
- 7:11>> Does this support the new Phantom 5090?
- 7:13>> Looks about the same, right? Watch the
- 7:15face and listen at the end.
- 7:17>> Oh, yeah, 24 lanes.
- 7:20Wait.
- 7:21THEY LIED TO US AGAIN.
- 7:24>> [laughter]
- 7:24>> DID YOU CATCH THAT?
- 7:27WHISPER picks it up. Word error rate on
- 7:30FP8 is 0.18. And basically, word error
- 7:33rate is just how many words drifted from
- 7:35the baseline transcript. Zero is
- 7:38identical. This is the only quantization
- 7:41between BF16 and Q3KM with a non-zero
- 7:45WER, word error rate. The model added a
- 7:48laugh that does not exist in the
- 7:50baseline audio. SSIM point 87 LPIPS is
- 7:55at point 07, pretty good. By the way,
- 7:57the top three are for video quality
- 8:00comparisons, and the bottom two for LTX
- 8:02only are for audio. And here we have mel
- 8:04spectrogram MSE 26.7.
- 8:07And for video, by the way, the SSIM is
- 8:10the pixel match. So, one would be
- 8:13identical. Obviously, BF16 is identical
- 8:16to itself, so that's one. And here on
- 8:18FP8, we're sliding down to point 8. And
- 8:21mel MSC down here is the audio
- 8:23fingerprint. So, zero would be identical
- 8:25in this case. Anything higher than that
- 8:27is drift.
- 8:28>> Does this support the new Phantom 5090?
- 8:30>> Oh, yeah, 24 lanes.
- 8:33Wait.
- 8:36Stealing
- 8:37our words again.
- 8:38>> So, compare that to Q80,
- 8:41one level down on the ladder, sort of
- 8:44more like
- 8:45horizontal. Same eight bits per weight,
- 8:47of course. Different format, though.
- 8:49Look at the difference between BF16 and
- 8:51Q80. There's a huge difference in the
- 8:53amount of space it takes. 46 GB versus
- 8:5623, but they look identical. Even the
- 8:58motion blur in those exact moments, in
- 9:01those frames, is exactly the same. But
- 9:03FP8
- 9:05not. And Q8 is better on every single
- 9:08metric, both video and audio. So, FP8
- 9:11underperformed in one
- 9:14one one, I don't know how to say that
- 9:16properly, don't ask me. And it also
- 9:20underperformed in LTX. Two different
- 9:23model architectures, two different
- 9:24teams, two different training runs.
- 9:27Weird, right? Maybe not. So, if you've
- 9:29got a choice between FP8 and Q8 for a
- 9:32video model right now, take Q8. I'm
- 9:35going to skip over levels between Q6 and
- 9:39Q5 because they do get a little bit
- 9:42worse, but there's no real big jumps
- 9:45here until you get to Q4. In fact, with
- 9:47these glasses, Q4 looks pretty good to
- 9:49me, too. Same thing with the car. I see
- 9:52slight differences with the detail of
- 9:54the car itself, but overall, the motion,
- 9:58the clarity looks pretty good. Let's
- 10:00take a look at the lady here. Q6 and Q5
- 10:03both 12 and 11 GB respectively on disk.
- 10:06It definitely looks like the same exact
- 10:08lady. So, what happens at Q4? Can we
- 10:10save more space and actually get away
- 10:12with it?
- 10:13>> [music]
- 10:15>> Q4, 9 GB on disk. This is the kind of
- 10:19thing you'd run on a 24 GB consumer GPU.
- 10:22And this is where most of the community
- 10:24lives. And if you watch this in
- 10:26isolation, you'd probably think, "Huh,
- 10:28it's fine." Sure, the face is a little
- 10:31bit fuzzier and a little bit not as
- 10:33crisp, but it's passable. The forest
- 10:36still moves, but if you compare it, L
- 10:38pips says, "We are
- 10:41at 0.34." That's further away than FP8.
- 10:45For a complex scene, we're at 0.35. What
- 10:48about that glass of water? Not so bad
- 10:50here, 0.2. For L pips, that character
- 10:52consistency is getting there. Let me
- 10:54show you. Notice anything different
- 10:55about her hair? And also, it kind of
- 10:57doesn't look like the same woman
- 10:58anymore. Here's the complex scene. It's
- 11:00a little bit harder to tell here. Tokyo
- 11:02night market, quarter disk size of FP16,
- 11:06but it still holds together. This is
- 11:08kind of like a sweet spot over here, and
- 11:10you can stop here if you don't have any
- 11:13reason to get any smaller. Red car looks
- 11:15fine, and do comment down below if you
- 11:18notice any weirdnesses that I didn't
- 11:20notice. We save a ton of space in LTX
- 11:232.3. The Q4 is only 14 GB here. There's
- 11:26some extra weirdness going on here with
- 11:28the face there and my eyes and Dan's
- 11:31eyes, and the unexpected reaction makes
- 11:33its reappearance.
- 11:36But look at the mel spectrogram here. We
- 11:38just jumped to 46.9.
- 11:41One level up at Q5, it was just 9.9. So,
- 11:46the audio fidelity just dropped roughly
- 11:48five times in a single step. Now, check
- 11:50the top row, the video metrics. SSIM is
- 11:54.88, basically right where it was. LPIPS
- 11:56is .07, the picture is right about where
- 11:59it was for FP8. So, it's definitely
- 12:02gradually getting worse here. And Q4
- 12:04shows a big change, at least in the
- 12:06objective measurements, but not as much
- 12:08as audio. So, that's finding number two,
- 12:10audio degrades before video here, even
- 12:13though it might not be perceived as such
- 12:16or might not be as noticeable as video.
- 12:18If you're only watching the picture,
- 12:19you're going to miss that. And why does
- 12:21this matter? The picture degrading is
- 12:23something you might catch on a rewatch,
- 12:25but the audio degrading is something the
- 12:27viewer hears immediately. You ever watch
- 12:30a YouTube video with terrible audio? I
- 12:32hope I hope my audio is actually decent
- 12:34here. Can't be perfect, but I try. But,
- 12:37you can put up with pretty bad video.
- 12:40However, if you hear terrible audio,
- 12:42people will just click away right away.
- 12:44Very different tolerance levels. So,
- 12:45models like LTX, the ones that support
- 12:48audio, have to take extra special care.
- 12:50>> [music]
- 12:54>> All right, two bits per weight. This is
- 12:565 GB for when. This is the bottom of the
- 13:00ladder, and it shows. That glass is
- 13:04>> [laughter]
- 13:05>> uh pretty bad looking. There's even
- 13:07differences frame to frame. Look at the
- 13:08lady. Woo,
- 13:10that's terrible. And Q3 is very similar
- 13:12to Q4, where it does mess with the hair
- 13:16quite a bit from the original and
- 13:18doesn't look at all like the original.
- 13:20But, Q2 is a whole different level
- 13:22that's completely unusable here at this
- 13:23point. Look at this character
- 13:25consistency chart. LPIPS is 0.54.
- 13:29Tokyo market, the complex scene, we're
- 13:32at 0.57, even worse. Both up about 60%
- 13:36from Q4. Now, LTX at the same two bits
- 13:39does the exact same thing. Look at that
- 13:41first frame, looks pretty good, right?
- 13:42This is image to video, by the way, in
- 13:44case you didn't know. Small model, 8.3
- 13:46GB,
- 13:47but look what happens if I play this.
- 13:49>> Does this support the new Phantom 5090?
- 13:51Oh, yeah, 24 lanes.
- 13:54Wait.
- 13:57THEY LIED TO US [screaming] AGAIN!
- 13:59>> OH MY GOD, THAT WAS just like tugs at
- 14:01the heartstrings. The audio is very
- 14:03different. It sounds robotic until he
- 14:05screams, then it sounds like more of a
- 14:07somebody trying to get an Oscar. But
- 14:09look how a terribly video quality is at
- 14:12every frame. Anything that's moving is
- 14:14basically completely destroyed and
- 14:16smashed.
- 14:20Whoa. Look at the transcript here. Word
- 14:23error rate here for Q3 and Q2, it went
- 14:26from they lied to us again to they lied
- 14:29to us again. Caps in the middle of the
- 14:31sentence. These are not the prompts, by
- 14:32the way, these are the transcripts for
- 14:35that uh whisper test that are generated
- 14:37from the video. And for LTX, it's not
- 14:39just this one skit. LTX again, this is
- 14:42BF16, FP8 looking like a very different
- 14:45person. Q8, Q6, Q5
- 14:49look like BF16 pretty much, pretty much
- 14:51the same person. But the leaves are
- 14:53wrong, even in BF16. Those are not maple
- 14:56leaves, I don't know what that is, like
- 14:58a starfish in the shape of a leaf or a
- 15:00leaf in the shape of a starfish, I
- 15:02guess. The shirt is also very difficult
- 15:05here. Every single one of these has a
- 15:07different shirt. Q5 and Q4 are kind of
- 15:09the same. And look what happens with Q3.
- 15:12We have a different person here.
- 15:16>> Can you tell which one of me is the real
- 15:18one?
- 15:19>> She's not even looking at the camera.
- 15:20Forget about Q2.
- 15:24>> Can you tell which one of me is the real
- 15:26one?
- 15:27>> Yeah, it's not you. Actually, it's none
- 15:28of them, but some of these could pass
- 15:30for real one.
- 15:34>> Can you tell which one of me is the real
- 15:36one?
- 15:36>> Did you notice the audio difference
- 15:38between BF16 and Q2? Huge difference.
- 15:42This one sounds real. Here's Q2 again.
- 15:44>> Can you tell which one of me is the real
- 15:46one?
- 15:46>> Totally fake at this point, right? This
- 15:48is what her charts look like. Pretty
- 15:50gradual SSIM, LPIPS also pretty gradual,
- 15:53but you know, it gets up there. They all
- 15:55have a little bump on FP8, which makes
- 15:57it equivalent to Q4/Q3.
- 16:00Let's check out the tech support scene.
- 16:02BF16, listen to the audio, too.
- 16:05>> Did you try turning it off and on again?
- 16:09>> I AM THE IT DEPARTMENT.
- 16:15>> NOT SURE WHAT'S GOING ON over there.
- 16:17Something's going on. With FP8, the lady
- 16:19is totally different.
- 16:20>> Try turning it off and on again.
- 16:24>> I AM THE IT DEPARTMENT.
- 16:31>> [laughter]
- 16:32>> LOOK AT HER. SHE'S NOT IMPRESSED. And
- 16:34the lady looks different in pretty much
- 16:36every single one of these. Her shirt
- 16:38color is different, the wall decorations
- 16:41are different, what's on the computer
- 16:43screen is different. But once we get to
- 16:45the 11 GB Q3, things just start melting
- 16:48down. Check this out.
- 16:49>> Did you try turning it off and on again?
- 16:54>> I AM THE IT DEPARTMENT.
- 16:59>> KIND OF LOOKS LIKE MR. Smith from The
- 17:01Matrix. You know, she's trying to push
- 17:03his buttons, and she's winning. Q2.
- 17:06>> Did you try turning it off and on again?
- 17:10>> I am
- 17:12the
- 17:14IT DEPARTMENT.
- 17:16>> DID YOU TRY
- 17:17>> WOW, that one is like way off the rails.
- 17:19First of all, the lady sounds robotic.
- 17:21The video is just totally destroyed.
- 17:26So, faces are still the hardest part.
- 17:28Whisper for Q2 and Mel spectrogram for
- 17:32Q2 all just are off the chart terrible.
- 17:36So, by now it looks like two bits just
- 17:38wrecks everything. But, here's what I
- 17:40did not expect. The glass of water scene
- 17:43SSIM is at .74 at two bits per weight.
- 17:47It doesn't look great when I'm examining
- 17:48with my balls of eye, but the numbers
- 17:51are saying it's okay. So, take these
- 17:54tests also with a grain of salt. I
- 17:56didn't come up with these tests. These
- 17:57are pretty much standard tests. So,
- 18:00yeah. Haven't shown you this one yet,
- 18:02the marble rolling down the ramp. And
- 18:04even though each one of these frames by
- 18:06themselves might look okay, the physics
- 18:09of this thing is just terrible at every
- 18:12single weight level. Even at full
- 18:15precision here, the ball just doesn't
- 18:17roll like a real ball and then two balls
- 18:20get merged. That's just the model, not
- 18:23the quantization. It's just wrong, not
- 18:26necessarily worse. Although, at Q2, you
- 18:29could pretty much say it's worse. So,
- 18:30finding number three I'd say is that
- 18:33there is no universal cliff. And that's
- 18:35true in both models. If the thing you're
- 18:38making is a single hard object in
- 18:40motion, you can probably ride this all
- 18:43the way down. But, if you're making
- 18:44humans in it, the floor is a lot higher.
- 18:47Bottom line,
- 18:48run Q4 unless you've got a reason not
- 18:51to. And the sweet spot doesn't move when
- 18:53you switch model architectures. But,
- 18:54also if your output has audio, listen to
- 18:57it at Q4. The picture will trick you,
- 18:59but the voice will not. And if you
- 19:01remember one thing from this whole
- 19:02video, let it be this. Format matters
- 19:06more than bit count. The full ladders
- 19:08are linked down below, every clip and
- 19:09every chart. And thanks to this rig, I
- 19:11was able to generate these videos pretty
- 19:14fast. If you want to know whether this
- 19:1596 GB RTX Pro 6000 rig was actually
- 19:18worth it, that video is over here.
- 19:21Thanks for watching, [music] and I'll
- 19:22see you next time.
- 19:29>> [music]
About this transcript
This page contains the full transcript of This Is What Happens When You CRUSH An AI Video Model by Alex Ziskind, generated from the public captions YouTube serves with the video. The transcript has 3,328 words across 499 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.