MINIMAX H3 IS THE NEW FREE VIDEO AI KING! — Transcript
Full transcript
- 0:00Minemax H3 is the best free video AI
- 0:03king ever, and it's not even close.
- 0:08>> [screaming]
- 0:09>> Hello humans, my name is Kay Ovrlod, and
- 0:11boy oh boy, do I have some mind-blowing
- 0:12stuff for you today. Because yes, you
- 0:14heard it right, we have a brand new
- 0:16video AI model that was released called
- 0:19Minemax H3. That is simply the best free
- 0:23video AI model ever made, capable of
- 0:26generating videos from text, images, as
- 0:29well as from multiple other references
- 0:32at the same time. It is just incredible.
- 0:35So, today I'll show you how to install
- 0:37it, how to run it locally, and on
- 0:39RunPod, and show you how to get the best
- 0:41results possible. So, that being said,
- 0:43sit back, relax, and let's go. And to
- 0:45install Minemax H3, you have two ways.
- 0:48The first is of course by using my
- 0:49one-click installer that is available
- 0:51for my Patreon supporters. Just
- 0:52double-click on the installer, and then
- 0:54it will automatically install ComfyUI
- 0:56and all the models and nodes that you
- 0:57need to run Minemax on your computer.
- 1:00And the second way is to rent a GPU on a
- 1:02website like RunPod, and use my special
- 1:04RunPod installer, and run Minemax as if
- 1:06this was running on your local computer.
- 1:08And then once you have ComfyUI up and
- 1:10running, for this video I prepared a
- 1:11special Minemax Ultra workflow that you
- 1:14can find on my Patreon. Then you're
- 1:15going to drag and drop inside ComfyUI.
- 1:17And now we can finally have some fun.
- 1:19Now, before we begin, try to explain the
- 1:21workflow, what makes this Minemax H3
- 1:24model so special, and why is everybody
- 1:27and their grandma talking about it.
- 1:29Well, Minemax H3 is a 33 billion
- 1:31parameter model that can generate videos
- 1:33from text or images in high resolution
- 1:36and with native audio, and that can run
- 1:39on your local computer. But that is not
- 1:41all. This model is extremely powerful
- 1:44and versatile thanks to its ability to
- 1:46do video editing and using multiple
- 1:48references at the same time to generate
- 1:51your final video. Now, obviously,
- 1:53Minemax H3 is a very thick boy, and you
- 1:56really need a lot of VRAM to generate
- 1:59videos at high resolution at a decent
- 2:01speed. But don't worry, even if you have
- 2:048 GB of VRAM, the model will work as
- 2:06well. Okay, so enough talking, let's
- 2:08actually show what this model can
- 2:10actually do and see how good the video
- 2:12generations are really are. Okay, so
- 2:14first, let's start with the
- 2:15text-to-video workflow, which is by far
- 2:17the easiest to understand. Here, once
- 2:19again, all of the most important part
- 2:21will be located in the first column
- 2:23right there. And before we start
- 2:25generating, we need to look at the
- 2:27available speed-up options that are
- 2:29available for the model right now, which
- 2:31are stage attention and spectrum. Now,
- 2:34basically, all of these two allows you
- 2:36to generate your video much faster at
- 2:39the expense of a small hit in video
- 2:42quality, especially if you're using
- 2:44spectrum. Now, there's also plenty of
- 2:46other nodes that can do that, but stage
- 2:48attention and spectrum are by far the
- 2:50best. Oh, and by the way, speaking of
- 2:52speed-ups, this is entrepreneur in the
- 2:54future because as I'm editing this
- 2:56video, we actually already have the
- 2:58release of the MiniMax Turbo Laura.
- 3:01That's right. Now, obviously, right now,
- 3:03it is still a Laura in training. It is
- 3:05still not done. As of right now, the
- 3:07only version available is the one at 500
- 3:11steps. And to make it work, you need to
- 3:13use the special comfy UI version. And
- 3:15I've already tried it and it is very
- 3:17good for a Laura that is not even done
- 3:19training. And although in this video,
- 3:21you will only see videos generated
- 3:24without the Laura, just know that if
- 3:26you're watching this video right now,
- 3:27the workflow on Patreon will already be
- 3:30updated to support this Laura. And you
- 3:32can either choose to run it without the
- 3:34Laura. And basically, the differences
- 3:35between the workflow is very simple.
- 3:38Without the Laura, we are basically
- 3:40using the res multi-step sampler with
- 3:43simple scheduler and 20 steps. Whereas,
- 3:45if we are using the MiniMax Turbo Laura,
- 3:49we will be using the Euler sampler with
- 3:52beta scheduler at eight steps. That's
- 3:54it. That is pretty much the only
- 3:56difference. So, instead of generating a
- 3:58video with 20 steps, you are now able to
- 4:00generate your video in around eight
- 4:03steps or even less. And of course, as I
- 4:05said, this Turbo Laura is not even done
- 4:07training, so in the upcoming days, we
- 4:09should have an even better Turbo Laura
- 4:12that we can use to generate absolutely
- 4:14amazing videos. So, yeah, there you go.
- 4:16Okay, so then right here, this is where
- 4:18you're going to input your prompt for
- 4:20the video generation. However, keep in
- 4:22mind that MiniMax H3 has a very, very
- 4:25specific prompting style. And I'm not
- 4:27going to go into details on how it works
- 4:30and everything. If you want more info,
- 4:32you can simply either read the note I
- 4:34have input right there, or if you want
- 4:36to simplify your life, you can simply go
- 4:38there and then copy and paste this
- 4:40entire note inside software like
- 4:42ChatGPT, and there you go. And now you
- 4:44can simply just ask ChatGPT to write you
- 4:47a prompt for whatever video you want.
- 4:49Oh, and also, right before you generate,
- 4:51let's also have a talk about the
- 4:53resolution, because here you have
- 4:55actually two options to choose the
- 4:57resolution of your video. Either you use
- 4:59the resolution selector that will
- 5:01automatically select the resolution
- 5:03following the aspect ratio guide right
- 5:06there, or you can simply just leave this
- 5:08option enabled and then choose yourself
- 5:10the resolution that you want. Now, if
- 5:12you don't want to type anything
- 5:13yourself, it is much easier to just use
- 5:16the resolution selector. And also here,
- 5:18I have highlighted all the resolutions
- 5:20that I recommend you to use, and it is
- 5:22simply either 480p, 720p, and 1080p. So,
- 5:27depending on your GPU, you might want to
- 5:29choose either one of these three,
- 5:31because these resolutions are more
- 5:32standard, so they are much easier to
- 5:34upscale. But obviously, if you have like
- 5:36a 5090, you can definitely go to the
- 5:381080p and then generate your huge video
- 5:41from scratch. And then here, finally,
- 5:43this is where you input the video
- 5:44duration in seconds. And of course, the
- 5:47longer the video, the longer it will
- 5:49take for the video to be generated. So,
- 5:51yeah, I mean, very simple stuff. If
- 5:53you've done it once, you've done it a
- 5:54thousand times. So, let me actually just
- 5:56generate a video. Let me write a prompt.
- 5:59Then I'm going to leave the 0.9
- 6:00megapixel resolution, use 10 seconds
- 6:03instead, and then I'm going to click
- 6:04run, which in the end will give us
- 6:06something like this.
- 6:07>> I'm dead, Jerry. The city says I'm dead.
- 6:10Well, that should clear up the parking
- 6:12tickets.
- 6:13Good news. I got you a box.
- 6:16>> So, yeah, I mean, it's pretty good. Now,
- 6:18as you can see, like this is very, very
- 6:20decent. Now, this was generated in one
- 6:22single shot. This is not cherry-picked
- 6:24or anything. I mean, it is really,
- 6:26really good. Actually, I even generated
- 6:28another version with the same exact
- 6:30prompt, but this is not at 1080p. Take a
- 6:32look.
- 6:34>> I'm dead, Jerry. The city says I'm dead.
- 6:37Well, that should clear up the parking
- 6:38tickets.
- 6:39>> [laughter]
- 6:40>> Good news.
- 6:41I got you a box.
- 6:43>> So, as you can see, in terms of quality,
- 6:45except resolution, there aren't that
- 6:47many differences between the two
- 6:49generations. Hence why it's probably
- 6:51better to generate at a lower
- 6:53resolution, like 720p, and then upscale
- 6:56it to a higher resolution later. And of
- 6:58course, the model can do a lot of very
- 7:00cool stuff, like a Breaking Bad parody,
- 7:02for example, or memes, kind of like
- 7:04this.
- 7:06>> Jesse,
- 7:07we have to cook.
- 7:15>> Yeah, there you go. Like, super, super
- 7:16funny
- 7:18uh Breaking Bad, Baking Bread parody. We
- 7:21all know that. It's not new, but it's
- 7:23still really, really fun. And of course,
- 7:25it can do more than just show parody or
- 7:28real-life stuff. It can generate anime
- 7:30parody as well. Like, take a look.
- 7:34>> I'm going to become the king of the hook
- 7:36again.
- 7:43>> Yeah, I mean, this is really, really
- 7:44incredible. Like this is insane. Like
- 7:47you can make your own anime parody or
- 7:49anime stuff already immediately right
- 7:52now. Like we've just a simple text
- 7:54prompt. Now this video was generated at
- 7:571080p, so you get like the best quality
- 7:59possible. But if you don't want to
- 8:01generate at such a high resolution, you
- 8:03can also generate at 720p. And if you
- 8:06do, you get something like this.
- 8:07>> I'm going to become the king of the
- 8:09hook. Okay.
- 8:17>> So yeah, I mean as you can see, very
- 8:18decent as well. Maybe not as good as the
- 8:211080p resolution, but it's still really,
- 8:24really fantastic. I mean it's really,
- 8:26really cool. So yeah, I mean I could
- 8:27spend hours and hours generating new
- 8:30videos to show you how good it is. But
- 8:32I'm sure that by this time that you're
- 8:34watching this video, you've already seen
- 8:36multiple examples anyway. So let's move
- 8:39on and see what else this model can do.
- 8:41Now the second thing that this model can
- 8:43also do is that instead of generating
- 8:45videos from text, you can generate a
- 8:47video from a image, which is definitely
- 8:49something that I prefer. I do like to
- 8:52generate my images on a separate model
- 8:54like Create 2, and then using an image
- 8:56to video model to bring it to life. I
- 8:59find this much, much better. So here
- 9:01basically the principle is exactly the
- 9:03same, except that now you have the
- 9:05choice of inputting two different
- 9:07images. Either you input one single
- 9:10image and you just make that image into
- 9:13a video, like I'm going to show you
- 9:14right now. Let's say I input this image,
- 9:17this selfie of a warrior on a
- 9:19battlefield, then I'm going to write my
- 9:20prompt. Then here I either have the
- 9:22possibility of using the original image
- 9:25size, or if I don't want that because
- 9:28it's going to take too much time, I can
- 9:29just use the normal resolution selector
- 9:32and then generate at the resolution
- 9:34written right there, which is going to
- 9:35be 720p instead. Input the video
- 9:38duration and now if I click run, which
- 9:40in the end will give us something like
- 9:41this.
- 9:47>> Well, I guess I'm going to be late for
- 9:49bingo tonight.
- 9:52>> And here, I mean as you can see, like
- 9:53this is fantastic. This is really,
- 9:55really good. This is great. I mean,
- 9:57huge, amazing quality. The generation is
- 10:00fantastic. The people behind him
- 10:03walking,
- 10:04you know, sitting down and everything. I
- 10:06mean, this is really super, super
- 10:08impressive. Super realistic, super good.
- 10:11I mean, yeah, I mean, this is really,
- 10:12really good. And of course, this is not
- 10:14the only thing that you can do with this
- 10:16particular workflow. And that is because
- 10:17you also have the ability to input a
- 10:20second image. Because this workflow can
- 10:22also work as an in-between video
- 10:25generation. So, if you put a first frame
- 10:28and a last frame, you will be able to
- 10:30generate a video between those two
- 10:33beginning and end frames. So, like for
- 10:35example, if I enable this last frame and
- 10:38then I input frame one with the warrior
- 10:40sitting on the ground, the last frame
- 10:43where he is shouting, if now I write my
- 10:45prompt and I click run, which in the end
- 10:48will give us something like this.
- 10:58>> [screaming]
- 10:59>> So, yeah, I mean, as you can see, this
- 11:00is really fantastic. I mean, we used
- 11:02only two separate frames, one beginning
- 11:06and one end, and with one single prompt,
- 11:08Minimax made a video in between those
- 11:12two frames. And I mean, I mean, what do
- 11:14you want me to say? It looks really,
- 11:15really good. It's it's fantastic. I
- 11:17mean, LTX could do it, too, but Minimax
- 11:20is really just on another level. It is
- 11:23really super, super good. But of course,
- 11:25once again, this is not the only thing
- 11:27that this model can do. Because not only
- 11:30can do all of that, but more, because
- 11:32Mini Max also has a very specific video
- 11:35model that allows you to use several
- 11:38different references to make your video.
- 11:40And the way it works is very, very
- 11:42similar to the image to video workflow,
- 11:45except that this time you have three
- 11:47different types of references. You can
- 11:49actually use image reference, video
- 11:52reference, and also audio reference,
- 11:55which really makes this workflow and
- 11:58model extremely powerful and extremely
- 12:00versatile. Now, the way you do this is
- 12:02the same exact principle as what we did
- 12:05with the image to video, except that
- 12:07this time it's kind of up to you to
- 12:08decide how many references you want to
- 12:11use. And I think the limit is around
- 12:13nine references that you can use at one
- 12:15single time, which is really just
- 12:17insane. Now, by default, I have input
- 12:20here like two different references for
- 12:22the images, one reference for the video,
- 12:24and one reference for the audio, but if
- 12:26you want to add more or disable
- 12:28different references, you can do so very
- 12:30easily as well. So, like for example,
- 12:32let's say that I want to add an
- 12:33additional reference image, all I have
- 12:35to do is just click on one node, then
- 12:38press control C on your keyboard, then
- 12:40control V to copy and paste that node,
- 12:43and then what you're going to do is that
- 12:44you're going to click and then connect
- 12:46it to that node right there. And you're
- 12:48going to connect it to the reference
- 12:50image node right there. And as you can
- 12:52see, now we have three different
- 12:54reference images that we can use. And
- 12:56you can do the same thing with video,
- 12:58audio, it's the same exact principle.
- 13:01And if you don't want to use a certain
- 13:02reference, you can just click on the
- 13:04node, and then click on this little
- 13:05button to bypass the reference, as well
- 13:08as right there. So, like for example,
- 13:09let's say that I upload an image of this
- 13:12character sheet right there, then
- 13:14another character for reference image
- 13:17two, and then for the third image, the
- 13:19interior of this castle right there. And
- 13:21now, if I write my prompt, which once
- 13:23again, I made it with ChatGPT, I'm not
- 13:26going to write all of that myself. Like,
- 13:28no way. And now, if I click run, which
- 13:30in the end gives us something like this.
- 13:32>> Wow. So, it can generate multiple
- 13:35characters inside the video?
- 13:37>> Yep, it looks like it.
- 13:41>> Amazing.
- 13:42>> Yeah, it sure is, buddy. I mean, this is
- 13:44really just incredible. I mean, it's
- 13:46insane. I mean,
- 13:49>> [laughter]
- 13:49>> I mean, I'm I'm kind of shocked. I'm
- 13:51going to say it's It works so good. It
- 13:54worked better than I thought it's going
- 13:55to be. It really like took the
- 13:57characters from the sheet, the character
- 14:00sheet, all the images, and then made
- 14:03video like this with full dialogue
- 14:06interacting with two different
- 14:07characters, two different voices, and I
- 14:11mean, it looks so good. It's It's
- 14:13amazing. I mean, yes, it was generated
- 14:15at a lower resolution, so you can have
- 14:17even higher quality if you generate at a
- 14:19higher resolution, or if you upscale it.
- 14:22But, I mean, just like that already, it
- 14:25is It is really, really good. I mean,
- 14:27it's It's incredible. I got to say, it's
- 14:29insane. It's fantastic. You don't need
- 14:31any LORAs, you don't need any training.
- 14:34Everything is already done for you. Just
- 14:36Just incredible. And also, of course,
- 14:38there's plenty of other ways of using
- 14:40this, using different types of
- 14:42references. Like, for example, let's say
- 14:44that I want to use a video reference.
- 14:47Let's say I upload this video of this
- 14:49woman kind of like turning around, and I
- 14:51want to like edit this video. I want to
- 14:54replace that woman with a different
- 14:56character. Let's say I want to replace
- 14:57it with that girl. I'm actually going to
- 14:59upload that girl into image zero. Then,
- 15:02I'm going to disable all the other
- 15:04references. I'm going to write my
- 15:06prompt, check the video direction. And
- 15:08now, if I click run, which gives us
- 15:10something like this.
- 15:16I mean, listen. I mean, is it amazing or
- 15:19what?
- 15:20What do you want me to say? It's It's
- 15:21incredible. I mean, we literally just
- 15:24like took a video reference and then
- 15:27replace it with a character from an
- 15:31image. No Laura training, no anything.
- 15:35We just like put everything together and
- 15:38it just worked. It just works. All
- 15:40right, it's
- 15:44Huh? It It's I'm I'm blown away. I'm
- 15:46blown away. If you can tell, I'm blown
- 15:48away. It is incredible. I mean
- 15:51And obviously, this is once again not
- 15:54the only thing that we can do. We can do
- 15:57much, much more than that. We can do so
- 15:59much more. I can spend hours and hours
- 16:02and hours showing you everything that
- 16:05you can do. Like for example, you can
- 16:07even like simply use like one single
- 16:09video as a reference. If I use like this
- 16:11video from, you know, Soprano with
- 16:13Paulie talking, which I'm not going to
- 16:15play because I kind of want to avoid
- 16:18any, you know, copyright issues if there
- 16:21is any. But what I can do instead is
- 16:23just use that video as a reference and
- 16:25then make him say something completely
- 16:26different with my prompt. And because
- 16:28the video not only can reference the
- 16:32video itself, but also the audio, we
- 16:34should also be able to copy his voice
- 16:37fairly well. So now if I click run, so
- 16:40that in the end we get something like
- 16:41this.
- 16:42>> Wow, this model is incredible. If Tony
- 16:45hears about this, he's going to be
- 16:46pissed.
- 16:47>> So yeah, I mean, what what do you want
- 16:49me to say? It's It's great. It's
- 16:51fantastic. It's
- 16:53It's everything that we've ever wanted
- 16:54from a video model. It's It is that
- 16:57simple, guys. It is that simple, okay?
- 17:00Like Why Why Why are you still here?
- 17:03Just Just stop listening to me and, you
- 17:06know, like and and do your thing. Do
- 17:08your thing. Try it out. It It's It's
- 17:10amazing. It's amazing. Oh, and also one
- 17:12little trick. If you want to upscale
- 17:14your video from a lower to a higher
- 17:16resolution, I highly recommend using my
- 17:18LCX 2.3 ultra workflow version 3 because
- 17:22in that workflow I have added a video
- 17:24enhancer upscaler, which is actually
- 17:27really really good at enhancing and
- 17:30upscaling at the same time your video.
- 17:32So, you can literally just upload a
- 17:34video made with Minimax. This is a video
- 17:36that was made and generated at 720p, and
- 17:40then you can like input a resolution of
- 17:42full HD, for example, and then if you
- 17:45click run, it will use LCX 2.3 to take
- 17:48this video and then upscale it at a
- 17:50higher resolution and enhance it at the
- 17:53same time. So, that in the end we go
- 17:55from this, a 720p video that looks like
- 17:58this,
- 18:05there we go, to a 1080p video.
- 18:13Yeah, I mean, there is really a huge
- 18:15difference in quality between those two
- 18:18videos. Hence why this workflow is
- 18:20really insanely good because not only we
- 18:23are upscaling, but we are also enhancing
- 18:27the original video. And obviously, it
- 18:29takes way less time to generate at 720p
- 18:32and then upscale it at 1080p rather than
- 18:35generating the full video at 1080p
- 18:37instead. So, yeah, I definitely
- 18:39recommend you to use this upscaler
- 18:41whenever you want to upscale one of your
- 18:43videos made with Minimax. Now, the one a
- 18:46little bit of a weird stuff with this
- 18:48model is the fact that the license is
- 18:51very very strange. Basically, it says
- 18:53that you cannot really use the model if
- 18:56you are in the European Union, United
- 18:59Kingdom, Korea, or the United States,
- 19:02which is, you know, very very strange.
- 19:04But, that is because Minimax is
- 19:06currently in a lawsuit against Disney,
- 19:09and they kind of want to avoid any
- 19:10issues with the model. Now, if you want
- 19:13to make sure that everything is fine,
- 19:14you can simply click on this little link
- 19:16and just like sign a waiver giving you
- 19:19the right to use this model however you
- 19:21want. It takes like 30 seconds to do and
- 19:24then you're good to go. Now, obviously,
- 19:26if you are like a normal person, you
- 19:28don't necessarily need to do that. I you
- 19:30know, you're fine. But in my case and a
- 19:32lot of people's cases, since we do have
- 19:34companies and we do represent companies
- 19:37ourselves, just to make sure that
- 19:38everything is fine, we need to kind of
- 19:40go through this, sign that waiver, and
- 19:43everything is going to be fine. So, yes,
- 19:45YouTube, if you're watching this video,
- 19:47just know that I did receive
- 19:49confirmation and I do have the right to
- 19:52use this model. Thank you very much. I
- 19:54love you. No sarcasm intended. So, yeah,
- 19:57there you go. This has been Min Max H3,
- 20:00simply the best open weight video AI
- 20:03model ever made. An absolute beast of a
- 20:06model that you can run on your computer
- 20:08right now. And once again, in this
- 20:10video, I barely scratched the surface of
- 20:13what this model can do. And even with my
- 20:15version one of the Min Max workflow, you
- 20:17can pretty much do absolutely insane
- 20:19stuff. But don't worry, this is
- 20:21definitely not the last time that I'll
- 20:22be making an update video about this
- 20:25model. There is really too much to say
- 20:27and too much to do. So, yeah, like once
- 20:30again, stop wasting time, download this
- 20:32workflow, download this model right now
- 20:34either locally or on RunPod, and just
- 20:37use it. It is incredible.
- 20:41Insane. So, that you and I can finally
- 20:44make our dreams come true.
- 20:51>> [music]
- 20:56>> And there we have it, folks. Thank you
- 20:58guys so much for watching. Don't forget
- 20:59to subscribe and smash the like button
- 21:01for the YouTube algorithm. Thanking also
- 21:03so much to my Patreon supporters for
- 21:04supporting my videos. You guys are
- 21:06absolutely awesome. You people are the
- 21:07reason why I'm able to make these
- 21:09videos, so thank you so much, and I'll
- 21:11see you guys next time. Bye-bye.
About this transcript
This page contains the full transcript of MINIMAX H3 IS THE NEW FREE VIDEO AI KING! by Aitrepreneur, generated from the public captions YouTube serves with the video. The transcript has 3,788 words across 549 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.