ComfyUI Tutorial MiniMax H3 on RTX 3060 6GB VRAM – Optimized Low VRAM Workflow + LTX 2 3 Upscaling — Transcript
Full transcript
- 0:00Hello everyone, welcome back to the
- 0:01channel. Since the release of the LTX
- 0:04model, AI, video generation has become
- 0:06much more accessible to the users with
- 0:07consumer grade GPUs. Now we have a
- 0:10powerful new model called Minimax H3
- 0:12which is shaping up to be the strong
- 0:14competitor to the LTX model. Miniax H3
- 0:18supports a wide range of generation
- 0:19model including text to video, image to
- 0:22video, first frame to last frame
- 0:23animation, reference to video, and even
- 0:25video editing, making it one of the most
- 0:28versatile OpenAI video model available.
- 0:30In today tutorial, we will put Minimax
- 0:32H3 to the test to see how it will
- 0:34perform. I will also show you how I
- 0:36managed to make it run on the RTX 3060
- 0:39with only 6 GB of VRAM using an
- 0:41optimized low VRAM workflow and some
- 0:43additional nodes. Finally, I will
- 0:45demonstrate how to upscale the generated
- 0:47video using the LTX 2.3 model to achieve
- 0:50much higher quality results. So, without
- 0:52further ado, let's get started. Okay,
- 0:53the model has day zero support in Comfy
- 0:56Y and uh they claim to be an openweight
- 0:59model with native audio and the 2K video
- 1:02resolution, meaning that you can run
- 1:04this model to make it generate uh 2K
- 1:07resolution on a local uh RTX 3060. So if
- 1:11you scroll down a little bit, you can
- 1:12see that it is a next generation
- 1:14openweight video model that can allows
- 1:16you to go up to 2K and with the video
- 1:19duration of 15 seconds, which
- 1:21practically the same as the LTX model.
- 1:24However, it is a third generation video
- 1:26model following high low 01 and 02. From
- 1:29what I saw, it is the first model that
- 1:31can rival the CES 2.5. So we have here
- 1:34the text to video prompt only the image
- 1:36to video where you will uh animate an
- 1:39image the first and last frame video and
- 1:42reference to video. The model also have
- 1:44context understanding. So it should uh
- 1:46have a better prompt adurance and
- 1:49understanding compared to all previous
- 1:51model that we saw. We have also a better
- 1:53audio quality. And one of the most
- 1:56important thing is the editing and the
- 1:58motion transfer. As you all know, the
- 2:00motion transfer in the edex isn't its uh
- 2:03strong point even if you try with the
- 2:06other workflow or even if you try with
- 2:08Laura and other editing workflow.
- 2:11However, this one can do that easily.
- 2:13So, you can add a reference video that
- 2:15can supply movement, camera move,
- 2:17performance, cutting, rhythm and etc.
- 2:19Here we have some examples and they also
- 2:23provide the prompt so you can test them
- 2:25by yourself once you have the workflow.
- 2:27And let's talk now about the model. As
- 2:30you can see the first main downside of
- 2:32this uh minimax is the size of the
- 2:35model. You can see that the BF16 is
- 2:37waiting 17 GB and it is also the main
- 2:41model not to mention the text encoder
- 2:44and the VAE. And the first step of using
- 2:46this workflow is downloading all the
- 2:48necessary model. So let me take you to
- 2:51the config
- 2:53link. You can see here that we have the
- 2:55diffusion model, the text encoder and
- 2:58the VAE. So let's take a look here. We
- 3:00have a lot of diffusion model. You can
- 3:03also see that we have the first last
- 3:05frame to video which can do both text to
- 3:08video and image to video. And for video
- 3:11editing, you can use this reference to
- 3:13video model. But for now, I did not use
- 3:16those models. You can also see that we
- 3:18have a VAE that has different size. The
- 3:22BF16 is also has also a very uh
- 3:25important weight. And lastly, we will
- 3:27grab here the VAE. Make sure to download
- 3:30the audio and the video VAE first. Once
- 3:33it is done, you can go ahead and use
- 3:35this rebuilt AI uh GG web version. Uh
- 3:39this guide did a good work on this
- 3:41minimax model and it also has YouTube
- 3:44channel that you can check if you want
- 3:46to see some new tutorial or models. So
- 3:49here what we have we have the the text
- 3:51encoder GG version which clearly reduce
- 3:55its size. I directly downloaded this 8.5
- 3:59GB version. You can do the same. And
- 4:01here you can see that we also have the
- 4:04first last frame which can do image and
- 4:07text to video. So you I directly
- 4:10downloaded this Q3 version which has a
- 4:14lower size. Once everything is done, you
- 4:16can place the VAE under the VAE folder.
- 4:20Next, you can use the GG version. Place
- 4:23it under the unit folder. As for the
- 4:25text encoder, place it directly under
- 4:28the clip folder here. Now that we have
- 4:30everything, we need to install some
- 4:32additional node for optimization. And we
- 4:34will start with this Confy Spectrum MX.
- 4:38So, make sure to grab the code here. Go
- 4:40to confy root folder. Search for custom
- 4:43nodes on the search bar. Type in cmd get
- 4:46clone then paste your code here. Click
- 4:49enter and it will install everything for
- 4:51you. So this node is going to help you
- 4:54to boost up your uh video generation by
- 4:5725%. It works by reducing the expensive
- 5:01H3 transformer evaluation during the
- 5:04sampling steps. So it will helps you a
- 5:06lot while generating your video. So once
- 5:10we have everything, make sure to restart
- 5:12Confy and we are good to go. Okay, now
- 5:15let's talk about the workflow. As you
- 5:17can see here, we have three main groups,
- 5:20the text to video, the first last frame
- 5:22video, and finally the upscaling group
- 5:26using the LTX model. So what you can do
- 5:29first is bypassing all those group by
- 5:32using the bypass button. And let's start
- 5:35focusing on the first last frame
- 5:37workflow first. As you can see here, we
- 5:40have our first image and our second
- 5:42image. In order to generate the video, I
- 5:45need first to downscale the resolution
- 5:47of my images. After that, you can see
- 5:49that we having a subgraph here that
- 5:52regroup all the minimax uh workflow. But
- 5:55before that, make sure to select all the
- 5:57necessary model. We have the VAE here,
- 5:59the audio VAE, the GGF or the main model
- 6:03that will generate our video and the
- 6:06clip uh text encoder in order to make
- 6:09the prompt understandable by the model.
- 6:12We also here have some parameters or the
- 6:14video parameters starting with the
- 6:16duration. It is set to 5 seconds. Make
- 6:19sure to leave it as it is if you have
- 6:21the same uh configuration setup as me
- 6:23otherwise it will take you more longer.
- 6:25And we have this resolution selector
- 6:27here. The model can generate and adapt
- 6:30to many resolution. If uh for example
- 6:32here we have the square but you can
- 6:34choose the portrait, the portrait
- 6:36standard, the wide screen and etc. The
- 6:39model will generate the video for you
- 6:41without uh any artifact or without
- 6:43crashing. And that was also the case for
- 6:45the LDX model. So on that point they are
- 6:48practically the same. So let's take a
- 6:51closer look at the subgraph. By clicking
- 6:53on this button, you can see here that we
- 6:56have all the main model and in order to
- 6:58optimize the workflow and avoid
- 7:00crashing, I use this clean VM after each
- 7:04model. So once the model is load, the
- 7:06VRAM is cleared in order to save for you
- 7:09some additional VRAM. And next we have
- 7:12the conditioning. Here you can see that
- 7:14this conditioning will need the clip,
- 7:16the VA, the first frame, the last frame
- 7:19and this green dot over here is uh
- 7:21referring to the prompt section uh that
- 7:24is outside the subgraph. We also have
- 7:26the width and the height and the length
- 7:28of the our video. So once all these data
- 7:31is processed, we will end up with the
- 7:34positive output and our latent for the
- 7:37video here that is going to be used
- 7:39directly by the sampler. So our next
- 7:42step is the sampling steps which use a
- 7:44simpler custom advanced notes with the
- 7:47uler set as simpler because uh combined
- 7:50with the spectrum nodes it gives you the
- 7:53perfect uh boost or the highest boost
- 7:55while generating your video. So for the
- 7:57scheduleuler I am using a simple one. Uh
- 8:00the steps they must be set to 20 or 25
- 8:04between those two but 20 also give me a
- 8:07good results. You can also try with 15
- 8:10in order to test out the model. Then
- 8:12once you are satisfied with the result,
- 8:14you can try to increase it to 20 or 25.
- 8:17Okay. Once everything is done, the
- 8:19output here going to be used in two
- 8:22different way. The first one is going to
- 8:23decode our video using the video decoder
- 8:26and the second one is going to decode
- 8:29our audio using the audio decoder in
- 8:31order to create our audio. After that,
- 8:33we're going to save our video using this
- 8:36notes create a video and get video
- 8:38component. But along the way, I also
- 8:40added some clean VRAM or clean caches in
- 8:43order to avoid the VRM issues or the VRM
- 8:46crashes. Once everything is set up, we
- 8:48will have our video saved. So now let's
- 8:51let's talk more about this spectrum
- 8:53apply minimax nodes. As you can see
- 8:55here, uh using these two nodes allowed
- 8:57me to boost my video generation by 25%.
- 9:01meaning now the model is more faster and
- 9:04use less VRAM compared to the previous
- 9:07or the default workflow. Make sure to
- 9:09place that node after the unit loader or
- 9:12the the fusion model loader and this one
- 9:16is going to be connected directly here
- 9:18to the basic guider and the basic
- 9:20scheduleuler. So once uh it is connected
- 9:23that way it will uh really helps you
- 9:25boosting your video generation. Okay,
- 9:28that's the first subg graph or the
- 9:30workflow. As for the text to video, you
- 9:33can see that it is practically the same
- 9:35workflow. Here we have our resolution
- 9:37selector and the video duration, this
- 9:40the model and practically the same
- 9:41parameters. However, as you can see
- 9:44here, the minimax H3 image to video, we
- 9:47don't necessarily upload first or last
- 9:50frame. You can leave it as it is and it
- 9:52will automatically generate uh video
- 9:54based on text. One more important thing
- 9:57is the prompting and it also has its
- 10:00unique pro way of prompting. As you can
- 10:02see here for this first last frame
- 10:04animation. You can see that we are
- 10:06describing the some motion the style of
- 10:09the video and etc. Next we have the
- 10:12scene overview and the story board. So
- 10:15you can write here a storyboard like we
- 10:17saw in the LTX director by providing the
- 10:20the video duration much uh like here for
- 10:23example we set to much frame or watch
- 10:26the first frame our video and we will
- 10:28describe the storyboard or the video
- 10:31during uh all those uh time-lapses. You
- 10:34will also need to include the camera and
- 10:36the audio. As you can see here, I added
- 10:39clear audio, woman, voice over,
- 10:41speaking. Exactly. And I add my
- 10:43dialogue. As for the camera, you can see
- 10:45that I use some keywords like single
- 10:47smooth, backward tracking, zoom, no cuts
- 10:50and etc. As for the text to video, you
- 10:52can see that uh this one was provided by
- 10:55the default workflow of comfy UI. It
- 10:57also has the same uh principle like you
- 11:00will describe the style, the scene
- 11:01overview, the storyboard, the camera,
- 11:03audio and etc.
- 11:06I also included some prompt for you if
- 11:08you want to use the same and all you
- 11:10have to do is using this string over
- 11:12here and click run. While prompting make
- 11:15sure to use the same uh prompt uh
- 11:18structure otherwise you copy one prompt
- 11:21paste it to sh GBT or any uh other AI
- 11:24and ask him to give you uh other example
- 11:27of using the same prompt structure of
- 11:29this one. So once everything is set, you
- 11:31can start generating your video which
- 11:33will bring us to the results section of
- 11:35this video tutorial. Okay, before
- 11:37showing you the results, I want to talk
- 11:39to you about the resolution which is the
- 11:42most important factor for this video
- 11:44generation. For the first last frame, I
- 11:46used megapixel of 0.4 and you can use
- 11:49this table in order to see the
- 11:51resolution of your final video. Of
- 11:54course, I jumped to 0.5 and it works uh
- 11:58for me, but not the first time. I first
- 12:01run out of RAM and but when I insisted
- 12:04on that by clicking on the run button,
- 12:06it managed to work. It took me more time
- 12:08to generate. As for the text to image,
- 12:10you can directly go ahead and use 0.5
- 12:13MP. it will uh generate your video more
- 12:17easily because the text to video doesn't
- 12:19necessarily need need to encode the
- 12:22image in order to animate it which can
- 12:24grant you to generate the video at a
- 12:26higher resolution compared to the image
- 12:28to video. As for the generated time, I
- 12:31leave this note here where I recorded
- 12:33every generation time. So for the LDX
- 12:37upscaler, it took me around 8 minute in
- 12:40order to upscale or double the
- 12:41resolution of a video generated at 0.5
- 12:44MP. As for the text to video, it took
- 12:47around 20 minutes to generate a video
- 12:49with 0.5 megapixel. And lastly, we have
- 12:52the first last frame which took 35
- 12:54minutes in order to generate a video at
- 12:5704 megapixel. To generate high video
- 13:00resolution, the best solution is using
- 13:02the text video and it takes upscale in
- 13:04order to gain more time and avoid
- 13:07running out of your RAM. As for the
- 13:09video generated, I started with this
- 13:11mouse here and you can clearly see the
- 13:13quality of video and the quality of the
- 13:15movement that looks more smoothly and
- 13:18can even generate the same object at
- 13:20different perspective which is really
- 13:22awesome by the way. Next, I jump to this
- 13:25commercial linking and you can see that
- 13:27it also has some unique way of
- 13:30generating the video. The movement looks
- 13:32more smoother and also the camera is uh
- 13:35pretty awesome compared to what we used
- 13:37to saw on the LTX. After that, I tried
- 13:41with this uh anime superhero video and
- 13:44it also give me this results which also
- 13:47looks pretty good. You can see that uh
- 13:50we can generate a lot of styled and we
- 13:52can also include the text on the video
- 13:55which is new things by the way that
- 13:57looks very awesome too. After that I use
- 14:00this prompt of Chinese woman that looks
- 14:02on the mask and here we have some dragon
- 14:05animation here. The results is not uh
- 14:08pretty awesome on this one. However uh
- 14:11if you try to increase the resolution
- 14:13you will end up with something very
- 14:15stunning. So right now this Miniax model
- 14:18seems to have a very good uh video
- 14:21generation. It has better camera
- 14:23movement and actor movement compared to
- 14:25LTX 2.3 and it can adapt to a lot of
- 14:29styles while generating a text directly
- 14:31incorporated on the video which is a new
- 14:34feature too. However, I'm pretty sure
- 14:37we'll have more optimization regarding
- 14:39uh the steps or also the resolution in
- 14:42order to adapt it more easily to all
- 14:45users. Okay, that was my first review on
- 14:47this Miniax model. We saw that this
- 14:49workflow can run with 6 GB of VRAM and
- 14:53even if it reproduce lower resolution uh
- 14:55video quality, we can upscale it using
- 14:57this ATX upscaler in order to get a
- 15:00decent quality. I will release more
- 15:02video on this model and I will try to
- 15:04give you more impressive and more
- 15:07stunning results next time. If you like
- 15:09this video, please push the like button
- 15:11for me, subscribe to my channel, leave
- 15:13me some comments down below, and don't
- 15:15forget to become a Patreon member of my
- 15:16Patreon page where you can get early
- 15:18access to my workflow and ask me for
- 15:20additional help. You can also go ahead
- 15:22and become a YouTube premium subscriber
- 15:24in order to get early access to my
- 15:26workflow too if you can't access to
- 15:28Patreon page. So thank you.
About this transcript
This page contains the full transcript of ComfyUI Tutorial MiniMax H3 on RTX 3060 6GB VRAM – Optimized Low VRAM Workflow + LTX 2 3 Upscaling by CG Pixel, generated from the public captions YouTube serves with the video. The transcript has 2,722 words across 370 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.