ComfyUI MiniMax H3: Best Video Generation Workflows (Ep29) — Transcript
Full transcript
- 0:01[music]
- 0:02Welcome to another video tutorial. Today
- 0:05I will show you how to use the new
- 0:07Miniax H3 video model so you can create
- 0:11awesome videos.
- 0:13>> Welcome to my Comfy UI setup where every
- 0:15cable has a purpose probably.
- 0:19Perfect. I have now connected everything
- 0:21exactly the way I was not supposed to.
- 0:25So I guess we are buying more VRAM.
- 0:28>> What is your name? I am Pixab Bunny.
- 0:32>> Where are you coming from?
- 0:42>> You always look at me like you
- 0:43understand everything. Are you my little
- 0:45best friend?
- 0:46>> Of course I am. But a best friend also
- 0:48deserves treats.
- 0:49>> So that is what this is about. You are
- 0:51adorable and very clever.
- 0:53>> Clever, cute, and hungry. That is my
- 0:55full talent set.
- 0:57>> Milo, be honest. Do you think I am
- 0:59overthinking everything again?
- 1:01>> Yes, but in a very talented way.
- 1:04>> That is not exactly comforting, but
- 1:06somehow it helps.
- 1:22As you know, I use the Comfy UI easy
- 1:24install like in episode 1 because it has
- 1:27a lot of useful scripts that make things
- 1:28easier. So, before I run Miniax, I have
- 1:31to do some updates. I will update easy
- 1:33install because that will also update my
- 1:36Pixaroma nodes. If you have another
- 1:38Comfy UI version, just update Pixar Roma
- 1:41nodes manually with GitPull or from the
- 1:43manager if it works. I asked for it to
- 1:45be added to the manager again since it
- 1:47was removed for a while because of the
- 1:49problem I had with GitHub in the past.
- 1:51Press enter to exit when it shows the
- 1:53message. Then you need to update Comfy
- 1:56UI. You should have a similar BAT file
- 1:59somewhere, either in the update folder
- 2:01or in the Comfy UI folder that lets you
- 2:04update Comfy UI. This is necessary
- 2:07because the miniax nodes are new and you
- 2:09do not have them in older comfy UI
- 2:11versions. Then to speed things up, it
- 2:14helps to have Sage attention. If you do
- 2:17not have it with easy install, it is
- 2:19quite easy to install. From the add-ons
- 2:22folder, run this BAT file. And if all
- 2:24these things are green, it should work
- 2:26correctly. If you get messages in red,
- 2:28then maybe your video card is not
- 2:30supported. This is for Windows and
- 2:32Nvidia cards. I had some problems with
- 2:3420 series video cards, but on the 30
- 2:37series, 40 series, and 50 series, it
- 2:40should work correctly. So, after you run
- 2:43that, you will have this extra comfy UI
- 2:45startup BAT file with Sage Attention. If
- 2:48we take a quick look inside, we can see
- 2:51it has these arguments. So, it knows to
- 2:53start with Sage Attention. This means we
- 2:56do not need to add any extra Sage
- 2:58Attention nodes to our workflow. In
- 3:00fact, if you have Sage Attention nodes
- 3:03in your workflow, it is better to start
- 3:05Comfy UI normally and let the nodes do
- 3:07the job. But I prefer this method
- 3:09because it gives me smaller workflows
- 3:11and it works the same. So now in the
- 3:14launcher, I can start either the desktop
- 3:17or browser version of Comfy UI with Sage
- 3:20Attention. You can see that when it
- 3:22starts, it says it is using Sage
- 3:24Attention. I added a version check
- 3:27Pixaroma node here so you can see which
- 3:30Comfy UI version and which version of
- 3:32Pixaroma nodes I have at the moment of
- 3:35recording. Make sure you have this
- 3:37version or a newer one, not an older
- 3:39version. Then you can download the free
- 3:42workflows in the Comfy UI folder. Look
- 3:45for the user folder, then default and
- 3:48finally workflows. You can place the
- 3:50downloaded folder with the workflows
- 3:52there. I removed the resources and left
- 3:55only the actual workflows in that
- 3:57folder. So now in Comfy UI, you should
- 4:00see the new workflows. If not, use this
- 4:03refresh button. You have them organized
- 4:05in the same order that I present them in
- 4:07this tutorial. Some of you may have
- 4:09noticed some new buttons here, like a W
- 4:12letter and a question mark. This new
- 4:14button or the alt + W shortcut opens the
- 4:18workflow manager that I built for you.
- 4:20This is a long video, so I will keep
- 4:23this part short. Basically, you can move
- 4:26and arrange workflows, rename them, and
- 4:28do pretty much anything you want with
- 4:30them, so you can stay organized. But the
- 4:33best way to get started is to click this
- 4:35help button and read about everything it
- 4:37can do. It has a large help section, so
- 4:39it should include the basics you need to
- 4:41figure it out. You can click W again to
- 4:43close it. You also have this help file
- 4:46for all Pixaroma nodes so you can easily
- 4:48find out how to use them and what they
- 4:50do. I hope the Comfy UI team creates
- 4:53something similar for their nodes as
- 4:55well to make things easier. So, click
- 4:57around, search, and look at what each
- 4:59node does. You can also find a lot of
- 5:02other useful information there. If you
- 5:04do not want to see those buttons, you
- 5:06can go to settings, then look for
- 5:08Pixarroma and scroll down until you find
- 5:10these settings. From here, you can hide
- 5:14all the buttons or only the ones you
- 5:16want. You can use the shortcut for
- 5:18workflows and the help section can also
- 5:20be opened from each Pixaroma node. For a
- 5:23line, you need to have it enabled for it
- 5:25to work. So, let me move on to the
- 5:27MiniAX H3 model. First, I will open the
- 5:31folder with the workflows for this
- 5:32model. By the way, you can add workflows
- 5:35to favorites, and you can also
- 5:37right-click them to see more options,
- 5:39including one that lets you add a
- 5:41preview cover for that workflow. If you
- 5:43select a workflow, like this texttovideo
- 5:46workflow, it will show you a preview and
- 5:48useful information on the right. Scroll
- 5:51down and you can see which model it uses
- 5:54and a lot of other information. You can
- 5:56also add a note to the workflow so you
- 5:58can search for it by that note later.
- 6:00Double click a workflow to open it or
- 6:03use the open button. I tried to make the
- 6:05workflows simple and easy to use and to
- 6:08reduce the number of nodes where
- 6:09possible. All the things you need for
- 6:11these workflows are in this note. The
- 6:13most important one is the Miniax H3
- 6:16model which is quite big and you can
- 6:18click this button to see more versions.
- 6:20So here is a list of them. There are
- 6:22actually two models with different
- 6:24quantizations. One is first last frame
- 6:27to video and one is reference to video.
- 6:30No matter what video card you have, even
- 6:32if it is powerful, I still recommend
- 6:35that everyone uses this model, the
- 6:37pruned int 8 version. The rest are
- 6:39slower with not much improvement in
- 6:42quality. For the reference model, use
- 6:44this version as well, but I already have
- 6:46the correct model selected in the
- 6:47downloaded workflow, so you do not have
- 6:49to worry about it. Then there is one
- 6:51thing that many people do not seem to
- 6:53talk about, the license. If we read the
- 6:56license, we can see that some
- 6:58territories are excluded from it. The
- 7:00European Union, United Kingdom, Korea
- 7:03and United States of America are not
- 7:06covered by this license. I assume this
- 7:08is happening because of legislation in
- 7:10those countries and instead of having
- 7:12problems, they added this condition to
- 7:14the license. But does that mean no one
- 7:16from those countries can use it? I added
- 7:19another link here where you can read
- 7:21more about the license and they explain
- 7:23it in more detail. They explained that
- 7:25they had to choose between not releasing
- 7:28the model until it complied with the
- 7:29legislation in those countries or
- 7:32releasing it while excluding those
- 7:33countries or something similar. I am
- 7:36from Europe, so how can I use it if
- 7:39Europe is excluded? There is a link here
- 7:41to an application form. It asks for an
- 7:44organization. I am not sure if this also
- 7:47applies to people without an
- 7:48organization, but you probably do not
- 7:50need this license anyway if you only use
- 7:52it for fun and personal use. So, I
- 7:55clicked this application form and
- 7:56completed my details. I have the
- 7:58Pixaroma company, which is a oneperson
- 8:01company where I do everything myself.
- 8:03Then, it asks about annual revenue and
- 8:05obviously I do not make that amount of
- 8:07money. After that, you have to check a
- 8:10lot of terms confirming that you agree
- 8:12to comply with the license. I am not a
- 8:14lawyer but I think this means they give
- 8:17the responsibility to you so they cannot
- 8:19be held accountable for any problems
- 8:21caused by how you use the model. This
- 8:24keeps them safe and I am happy that I
- 8:26can use the model so everyone is happy.
- 8:29That is how I did it so I could use the
- 8:31models on my channel. I received an
- 8:33email confirming that I was approved and
- 8:35that was all. So let me show you how to
- 8:37get the models. Use this button to
- 8:40download the model. It will take a while
- 8:42because of its size. After downloading
- 8:44it, navigate to your Comfy UI folder. Go
- 8:47to models, then diffusion models, and
- 8:50create an H3 folder to keep things
- 8:52organized. Place the model for this
- 8:54workflow inside that folder. I have two
- 8:57models because I also downloaded the
- 8:59reference model for the other workflows.
- 9:01Now, you have the Miniax model. Next, we
- 9:04need a text encoder. Download this Quen
- 9:07model which is also big but very smart
- 9:09with 32 billion parameters. Until now we
- 9:13used models with four and 8 billion
- 9:15parameters. Place it in the text
- 9:17encoders folder. For the VAE we have two
- 9:20models because this model also supports
- 9:23audio. Download both of them into the
- 9:25VAE folder and then you can load them
- 9:27into these nodes. As for the custom
- 9:30nodes, you know me, I try to keep things
- 9:32simple. Only Pixaroma nodes are used. I
- 9:35also included a custom chat GPT link I
- 9:38made for prompting, but I will talk more
- 9:40about that later. After all the models
- 9:42have finished downloading, you can press
- 9:44R to refresh the node definitions, and
- 9:47everything should work if the files have
- 9:49the same names and are placed in the
- 9:51correct folders. If not, you will need
- 9:54to select them manually inside the
- 9:55nodes. Like I said, I like simple
- 9:58workflows, so compared with what you
- 10:00find on the internet, this should be
- 10:02much easier to use. I added some steps
- 10:05in blue for you to follow. And this
- 10:07workflow only needs three steps from
- 10:08you. Let me start with the first one
- 10:10where you can choose the orientation and
- 10:12size. I added some of my favorite sizes
- 10:15here, which I recommend depending on
- 10:17your video card. The more VRAM you have,
- 10:20the larger the size you can generate.
- 10:22People have tested this even on a video
- 10:24card with 8 GB of VRAM, and it worked if
- 10:26they had enough system RAM and dynamic
- 10:28VRAM was enabled. Remember to switch
- 10:31between portrait and landscape since
- 10:33that decides the ratio. Because we are
- 10:35making videos, I did not include a
- 10:37square ratio, but you can add one as
- 10:40well. If you go to settings, you can add
- 10:42the sizes you want here, change their
- 10:44order, remove sizes, or add more. I
- 10:47included all the sizes the model can
- 10:49generate, and I added a star next to the
- 10:52ones that I think work well. For the
- 10:54optimal size, I think this one is the
- 10:56best, but I did many of my tests using
- 10:59this size. Miniax also prefers sizes
- 11:02that are multiples of 32, so I enabled
- 11:05that option. No matter what size you
- 11:07enter, it will round it to a compatible
- 11:10value. So, that is the Pixar sizes node.
- 11:13I only changed the title to make it
- 11:15easier to use. For this video recording,
- 11:17I will select portrait and this size.
- 11:20Then, I created another node for the
- 11:22duration. If we go to settings, you can
- 11:25add your custom durations here in
- 11:26seconds. I also added a few ready-made
- 11:29formulas directly into the node. By
- 11:32default, it uses this Miniax H3 formula,
- 11:36but you can use others for different
- 11:37workflows or create your own formula.
- 11:40The cool part is that you can turn it
- 11:41into a slider if you want and adjust the
- 11:44values or you can just type them if you
- 11:46prefer that. I like this button version
- 11:48because I can add only the values I
- 11:50need. So, I will use a 3-second video
- 11:53for this test. Then comes the long
- 11:56detailed prompt. The more details you
- 11:58add, the better it can understand what
- 12:00you are asking for. You can write all of
- 12:02that by hand, but I cannot write prompts
- 12:04that are so detailed and structured in a
- 12:07way that AI understands perfectly. So,
- 12:09why not use AI to create prompts for AI,
- 12:12right? Even if a short prompt might work
- 12:15well, you give the model too much
- 12:17freedom to invent things. That is why I
- 12:19built this custom chat GPT. You can
- 12:22click on it and use it. It should also
- 12:24work with a free subscription, but you
- 12:26can use another LLM if you want. I will
- 12:29also provide some prompt formulas inside
- 12:31the workflow archive. Since I created
- 12:33this GPT, I can also edit it so you can
- 12:36see how I set it up. In the configure
- 12:38tab, I added instructions telling it to
- 12:41use the attached knowledge. Here I
- 12:44uploaded this text file with the long
- 12:46formula that explains how to write the
- 12:48prompts. The formulas are also inside
- 12:50this folder that you downloaded with the
- 12:52workflows. There is one for the
- 12:54reference model and one for text to
- 12:56video and first last frame to video. You
- 12:59can look at them, change them, and adapt
- 13:01them however you want. So back to the
- 13:04custom GPT, you need to tell it what
- 13:06type of prompt you want. Is it for text
- 13:08to video? Is it for first frame to
- 13:10video? or is it for first frame last
- 13:12frame to video. This helps it choose the
- 13:15correct prompt structure. In this case,
- 13:17I want a textto video prompt. Then I
- 13:20describe in my own words what I want and
- 13:22I hope it understands me. And I got a
- 13:24long detailed prompt for that. Some
- 13:27parts look like instructions, but that
- 13:29is fine because it helps the model
- 13:31understand what I want it to do. You can
- 13:33see how it describes everything and also
- 13:35how it uses English. I tried it with
- 13:38other languages and it works. It says it
- 13:40supports 11 languages, but I also tried
- 13:43Romanian and it worked. Sometimes it
- 13:46mixed up one word, but overall it worked
- 13:48well. I added a run timer here so I can
- 13:51see how long it takes, but more recently
- 13:53I prefer using a run log because I can
- 13:56see the history for each run. Let me run
- 13:58it and see how it works. It loads all
- 14:01the models and then the prompt and
- 14:03settings go into this new miniax node
- 14:05created by the Comfy UI team. After
- 14:08that, it goes into the K sampler where I
- 14:10used 20 steps and set CFG to one because
- 14:13we do not have a negative prompt. For
- 14:15the sampler, I used res multi-step. And
- 14:18for theuler, I used simple. Some people
- 14:21also recommend beta. So, you can try
- 14:23that one as well. Then, I ran into a
- 14:26problem and got an error. What could it
- 14:28be, I wonder? I recorded this because
- 14:30some of you might get the same error,
- 14:32and I will show you what fixed it for
- 14:34me. What I did was close Comfy UI. Then
- 14:39inside the BAT file I used to start
- 14:41Comfy UI, I had this argument that
- 14:44disabled dynamic VRAMm. You could remove
- 14:47it manually and save the file, but to
- 14:49make sure all the BAT files are set
- 14:50correctly and have that argument
- 14:52removed, I will do something else. If
- 14:54you use a GGUF model, you might want
- 14:56dynamic VRAM disabled. But in this case,
- 14:59I want it enabled. So, I go to add-ons,
- 15:02then tools, and toggle it easily using
- 15:05this BAT file. Now, it removed that
- 15:07argument. If I go back and check, you
- 15:10can see that the argument is no longer
- 15:12there. So, everything should work now.
- 15:14But let me test it. I open comfy UI
- 15:17again and run the workflow. And now I
- 15:19got the video I described.
- 15:23>> Do you like my cat?
- 15:28>> Do you like my cat?
- 15:30I will run it again and then I can add a
- 15:33label for each recorded runtime. By the
- 15:35way, you can now rightclick the node, go
- 15:38to node settings and enable this option.
- 15:41Then you can close the settings and look
- 15:44now it shows my video card along with
- 15:46the VRAM and RAM I have. Pretty cool,
- 15:49right? And the second video looks like
- 15:52this.
- 15:53>> Do you like my cat?
- 15:58Do you like my cat?
- 16:00>> I also did some tests while I was not
- 16:02recording so the recording would not
- 16:04affect the generation times and here is
- 16:07what I got. All the tests are for a
- 16:105-second texttovideo workflow. I did 10
- 16:13runs and for each run I increased the
- 16:16resolution to see how it affected the
- 16:18total running time. The higher the
- 16:20resolution, the more time it takes. But
- 16:22the increase is not linear. For example,
- 16:24for this resolution, it took around 3
- 16:26minutes, but for full HD, it took 10
- 16:29minutes. So, I prefer to create a
- 16:3115-second video at a lower resolution
- 16:33instead.
- 16:37[music]
- 16:48[music]
- 16:50Heat. Heat.
- 16:52[music]
- 16:57[music]
- 17:03[music]
- 17:09[music]
- 17:14>> [music]
- 17:25[music]
- 17:26>> So you saw that the higher the
- 17:28resolution the better it looks but it
- 17:30also takes more time to generate. Let me
- 17:33move now to the first frame to video
- 17:35workflow since we also want to control
- 17:37the first frame so we can animate
- 17:38exactly the image we want. This workflow
- 17:41uses the same models and nodes. There is
- 17:44one extra thing here, an image, and that
- 17:47image is connected to the first frame
- 17:49input in the miniax node. After the
- 17:51image is loaded, step two is to choose
- 17:54the duration. The longer the video, the
- 17:56more time it takes, but at least the
- 17:59increase is more linear. So, if you know
- 18:01the generation time for 3 seconds, it
- 18:03will take around twice as long for 6
- 18:05seconds, three times as long for 9
- 18:07seconds, and so on. To make things
- 18:10easier, instead of constantly changing
- 18:12the video size, checking the image
- 18:14ratio, and adjusting everything
- 18:16manually, we can do all of that with a
- 18:18single click. So, I built another node
- 18:21called longest side. It takes the image
- 18:23from the load image node, makes the
- 18:25dimensions multiples of 32, and sets the
- 18:28longest side of the video to the value
- 18:30you choose here while keeping the
- 18:32original ratio. Then, step four is the
- 18:34prompt. For longest side, we have some
- 18:37settings here. You can see that for this
- 18:39workflow I used 32. You can change that
- 18:43from here by clicking the button
- 18:44multiple times. You can choose the
- 18:47longest side of the video from here. I
- 18:49added the recommended values, but you
- 18:51can add more. You can even use a
- 18:54different ratio if you want to crop the
- 18:56image. For example, if I want a square
- 18:59video, I can select that. Then from the
- 19:02settings, I can choose whether the crop
- 19:04should be centered or positioned
- 19:06differently. In this case, because the
- 19:08woman is near the top, a centered crop
- 19:11would remove her face. So, a top crop
- 19:13would work better. But I do not plan to
- 19:15crop the image. So, most of the time, I
- 19:17prefer to keep a ratio that is as close
- 19:20as possible to the original image. For
- 19:22the size tabs, you can also add your own
- 19:24values here. So, I can change them
- 19:27however I want. And then the buttons are
- 19:29updated here. The next time I click that
- 19:31button, it will use the new size. I will
- 19:34put it back to the default settings for
- 19:36this workflow. Then for the prompt, I
- 19:39usually take a screenshot of the load
- 19:41image node. After that, I go to chat GPT
- 19:44and open the custom chat GPT. I ask for
- 19:48a first frame to video prompt, paste the
- 19:50screenshot, and then describe what I
- 19:52want to happen. Chat GPT is quite good
- 19:55at this. Now, I have a long prompt that
- 19:57I can copy and paste here. I will use 5
- 20:00seconds and then run the workflow. I
- 20:03just noticed that I forgot to remove the
- 20:04square crop. So, I will cancel the
- 20:06generation, set it to keep the original
- 20:09ratio because I do not want to crop and
- 20:11run it again. It took around a minute
- 20:13and a half. Let me check the result.
- 20:18Well, hello there little sweetheart.
- 20:23Well, hello there little sweetheart.
- 20:25>> If I check the prompt, you can see that
- 20:27it used the English tag there along with
- 20:29the words in English. But like I said,
- 20:32you can use other languages, too. Here
- 20:35is an example in Romanian.
- 20:49>> [laughter]
- 20:52>> I did tests with 5, 10, and 15-second
- 20:55videos, and the time is around double
- 20:58for each extra 5 seconds, but the
- 21:00maximum for miniax is 15 seconds.
- 21:02Anyway, I tried to increase the
- 21:04resolution to this one. And now for 15
- 21:07seconds, I got 13 minutes. So, it is
- 21:09almost like 1 minute of running time per
- 21:11second of video at this resolution. Here
- 21:14are a few more tests for image to video.
- 21:17And it is almost the same as text to
- 21:19video just with a few extra seconds,
- 21:21maybe around five extra seconds at HD
- 21:23resolution. I now have some volunteers
- 21:26on Discord who test my workflows called
- 21:29bug hunters. And here is an example of
- 21:31tests on 16 GB VR RAM. They tried it
- 21:35with Sage Attention and also with some
- 21:37extra nodes. With easy cache, it reduced
- 21:40the time, but the results were worse. So
- 21:43I prefer the quality. The generation
- 21:45times with Sage attention as arguments
- 21:48and with nodes are similar. So I prefer
- 21:51using the arguments until someone
- 21:52releases better nodes. It is time to
- 21:55move to the next workflows that use
- 21:57besides the first frame also a last
- 21:59frame so you have control over where it
- 22:02ends as well. For this example, I tried
- 22:05two different bunny images to see if it
- 22:07can do a transition. So I started with
- 22:09this bunny sitting as the first frame.
- 22:11Then this ballerina bunny as the second
- 22:13frame. I set it to four seconds. I let
- 22:16it keep the ratio and this longest side
- 22:19and that is taken from the first frame
- 22:21since that usually decides the ratio of
- 22:24the image. So that first frame is
- 22:26connected here and then it goes into the
- 22:28miniax node into the first frame input.
- 22:31The second image is connected to the
- 22:33last frame input. I will run this
- 22:35workflow and just to give you an idea of
- 22:38how I prompted it, I took a screenshot
- 22:40of both nodes and in chat GPT I asked
- 22:43for a prompt for first frame last frame
- 22:45to video. Then you add how you want the
- 22:47transformation to happen. I wanted a
- 22:505-second video and it gave me a long
- 22:52prompt and the result for this workflow
- 22:55looks like this.
- 22:57[music]
- 23:02[music]
- 23:04For the next workflow, you can use it to
- 23:06get the last frame of a video. So, you
- 23:08decide how the video should end. So, for
- 23:10example, I have this ballerina bunny. I
- 23:13set the duration to 4 seconds and I kept
- 23:16the ratio and generated the video at
- 23:18this longest side. With a long prompt,
- 23:20the image should be connected only to
- 23:22the last frame input. And you leave the
- 23:24first frame input empty so the model can
- 23:26invent that part. So, if I run this
- 23:29workflow, I get this result.
- 23:33>> [music]
- 23:39>> So it is very easy to make the video end
- 23:42on a certain image and you do not have
- 23:44to generate two images all the time. So
- 23:46all these four workflows use the first
- 23:48model the first last frame to video
- 23:50model. Now I am moving to one that uses
- 23:53reference instead. I have one workflow
- 23:55here for two images and one for three
- 23:58images. Let me start with the two image
- 24:00one first. Make sure you downloaded this
- 24:02reference video model because it will
- 24:04not work correctly with the first one.
- 24:06Then for chat GPT, I made a different
- 24:09one because reference works a little
- 24:11differently. It does not start with a
- 24:13certain frame or end with it. It looks
- 24:15at the references and generates
- 24:16something new from them. So for the
- 24:19prompt, I will just take a screenshot of
- 24:21these two images and then in chat GPT I
- 24:24paste it and ask for a reference to
- 24:26video prompt for these two images. Then
- 24:28you describe how you want it. I already
- 24:30did that for this long prompt and the
- 24:32result is this.
- 24:36>> Are you ready to escape this summer?
- 24:41>> Are you ready to escape this summer?
- 24:43>> That is just a simple example, but you
- 24:45can do more complex things. Let me show
- 24:48you another example with three images.
- 24:50It can do more than three. You just
- 24:52connect more images there in the miniax
- 24:54node. It can do a maximum of nine
- 24:57images, three short clips of around 5 to
- 24:5915 seconds, but for me that took forever
- 25:02to generate. I am not sure if that is a
- 25:04bug, but I could not get it working
- 25:06properly. It can also use three audio
- 25:09files with a maximum of 12 files
- 25:11combined. I will show audio examples in
- 25:14a few minutes. Since this uses images as
- 25:17references, the ratio of the images does
- 25:20not matter that much. So, I can use
- 25:22custom ratios here. So decide on the
- 25:25maximum resolution you want to use and
- 25:27do not forget to choose portrait or
- 25:29landscape. I tell you this because I
- 25:31forgot many times to switch it. I could
- 25:34make it follow the ratio of the first
- 25:36image like in the other workflows, but
- 25:38then I have to keep in mind that the
- 25:40first image must have the correct ratio.
- 25:42This way I know I can use images with
- 25:45any ratio and I choose the video ratio
- 25:47myself. So, I want this bunny wearing
- 25:50that orange outfit in an exotic location
- 25:53like the last image. I used chat GPT for
- 25:56the prompt and the result is this one.
- 25:58>> Hello from paradise.
- 26:03Hello from paradise.
- 26:08>> Add more load image nodes there. Connect
- 26:10them and save your own workflows based
- 26:12on these. Now let me move to the next
- 26:15group which can also use an audio
- 26:17reference combined with an image. It
- 26:19does not do audio to audio. So I have
- 26:22one workflow here for speaking and one
- 26:24for singing. Let me start with the
- 26:26singing workflow first. Like I said it
- 26:29uses the same reference video model. You
- 26:31set the duration of the video and I made
- 26:34it so that it also affects the uploaded
- 26:36audio which helps you get better
- 26:37results. From this load audio pixeloma
- 26:40node, you can upload an audio file or
- 26:43select one from the input folder. You
- 26:45can see that when I select a different
- 26:47duration here, it also adjusts this
- 26:50orange selection. To make that work, the
- 26:52seconds output is connected here. So,
- 26:55the selected audio section always
- 26:56matches the chosen duration.
- 26:59[music]
- 27:04second selection to the right to start
- 27:06at a different point.
- 27:10>> So everything is controlled. If I
- 27:12disconnect the seconds input now I have
- 27:14more control and can adjust it manually
- 27:16however I want. I am showing you this
- 27:19because it can be useful for other
- 27:20workflows. But for this one it is easier
- 27:23to leave the seconds input connected.
- 27:26Also check the settings and help
- 27:28sections for more information. If you
- 27:30set a duration longer than the actual
- 27:32audio, you can either add silence or
- 27:34loop the audio. So, I will reconnect the
- 27:37seconds input. Let me pick 5 seconds and
- 27:40play it to see which part will be used
- 27:41for singing by looking at that white
- 27:43line.
- 27:45>> [music]
- 27:54[music]
- 27:58>> Maybe I will make it shorter and use 3
- 28:00seconds.
- 28:03[music]
- 28:04>> Okay, that part seems good for the
- 28:06input. I have a universal prompt here
- 28:08that works with almost anything. You can
- 28:10adjust it if needed, but it should work
- 28:12most of the time as it is. Now I will
- 28:15run it. I forgot to mention that you set
- 28:17the size of the final video from here
- 28:19using the longest side. This workflow
- 28:22also has this new H3 audio sync node
- 28:24that I made. If you click more, you can
- 28:27see that it helps keep the audio
- 28:29consistent. It is too complex for me to
- 28:31fully understand what it does, but
- 28:33Claude figured it out and I got good
- 28:35results. And the final result is this
- 28:38one. [music]
- 28:40[singing]
- 28:45You can adjust the prompt if you want
- 28:47some extra movement there. For example,
- 28:50I can add that the bunny puts his hands
- 28:51in the air while dancing or something
- 28:53similar. Then I run it again and I got
- 28:56this. [music]
- 29:01[singing]
- 29:03>> For some seeds, I got barely visible
- 29:05lines in the generated video. I am not
- 29:08sure if it is because I used a low
- 29:10resolution, if the model sometimes does
- 29:13that, or if there is a wrong setting
- 29:15somewhere. It is just something that
- 29:17appears sometimes, usually around the
- 29:19middle of the screen. Moving now to the
- 29:21second workflow, which is similar to the
- 29:23first one, but has a different prompt.
- 29:26So, I add the voice here.
- 29:29I can talk now since we have the MiniAX
- 29:31H3 model. Maybe I will make it shorter
- 29:34and use only 3 seconds. I can talk now
- 29:37since we have the Miniax H3 model. I set
- 29:40this size for the longest side so it
- 29:42generates faster. And now I will run it.
- 29:45You can see how I set up the prompt, but
- 29:47feel free to adjust it. Maybe you can
- 29:49find a better one, but this worked well
- 29:51for me. You also have more settings in
- 29:54the audio sync node. I made it so the
- 29:56audio does not go beyond 15 seconds
- 29:59since Miniax can only generate a maximum
- 30:0115-second video. Anyway, if the audio
- 30:04clip is longer, I set it to warn me, but
- 30:07you can also make it stop the run if you
- 30:09want. You also have the same options to
- 30:11add silence or loop the audio if the
- 30:13clip is too short. And the result is
- 30:16this one. I can talk now since we have
- 30:18the Miniax H3 model. I can talk now
- 30:20since we have the Miniax H3 model.
- 30:24Even if it is a video model, MiniAX can
- 30:26generate images or edit images, but it
- 30:29is a little slow. So I still prefer the
- 30:31Crea 2 model for that. But if you
- 30:34already have the model, at least you
- 30:36know it is an option. So for this
- 30:38portrait landscape node, I added a
- 30:41setting so it is constrained to
- 30:42multiples of 32. And you put your sizes
- 30:45there and just flip it to get the right
- 30:46ratio. You add a long prompt and you can
- 30:49run it. So it took 45 seconds for a full
- 30:52HD image and the result is not bad if
- 30:54you do full HD.
- 30:57If you reduce the size, you get faster
- 30:59generation, like only 17 seconds. But
- 31:02the result is not great, more like a
- 31:05frame from a lowquality video, which in
- 31:07fact it is. It generates like a batch of
- 31:10five frames and picks the first one,
- 31:12index zero, from that. I could not set
- 31:15the length to one because the node gave
- 31:17an error. I think Miniax likes five
- 31:19frames at once. And now let me try the
- 31:22edit workflow. It is similar to that
- 31:24one, but this time we load an image. And
- 31:27I added here what to change, but longer
- 31:30prompts that give details about what
- 31:32should change and what should stay the
- 31:33same could work better. I tried with a
- 31:36smaller size this time for edit using
- 31:38this one. It has the same get image from
- 31:40batch like before. And the loaded image
- 31:43is connected to the first frame. And
- 31:46this time it took around a minute. So it
- 31:48changed the hair like I asked, but it is
- 31:50losing a bit of quality. You can get
- 31:52away with small details, but sometimes
- 31:54it can change more things. For example,
- 31:57sometimes it changed the text to pink
- 31:59also. Not only the hair, but it depends
- 32:02on the seed. In this case, it worked
- 32:04even with a short prompt. So, do not
- 32:07forget to check the help for each
- 32:08Pixaroma node. So, you can learn more
- 32:11about the node. That help is updated
- 32:13when I update the node and change
- 32:15something in it. Search for the things
- 32:17you need. And I hope all these nodes are
- 32:19useful for you. I will leave you with a
- 32:21few more examples of videos generated
- 32:23with Miniax.
- 32:25>> Meet the Pixa time machine because
- 32:27waiting for the weekend is basically
- 32:29unbearable. Destination selected.
- 32:32Medieval era.
- 32:33>> Okay, definitely older than my
- 32:36apartment.
- 32:37Picks a time machine. Travel smarter.
- 32:40Everyone has a smartphone. I wanted
- 32:42something fresher.
- 32:44Hey, Apple phone. What's on my schedule?
- 32:46>> First a call, then a snack. Hopefully
- 32:49not in that order.
- 32:51>> Finally, a phone with appeal.
- 32:53>> I was going to [music] say that.
- 32:55>> Hey guys, welcome with me into this
- 32:56amazing jungle. It's so beautiful here.
- 32:59Look at this flower. The color is
- 33:00incredible. It's even more beautiful in
- 33:02real life. Oh my god, look, a little
- 33:04monkey. He's so cute. This is definitely
- 33:07the best part of the trip. Say bye,
- 33:09little guy. Welcome back to the vlog.
- 33:11Today we are exploring the jungle where
- 33:13everything is beautiful, wet, and
- 33:14probably trying to eat something. Look
- 33:16at these flowers. They are gorgeous,
- 33:17peaceful, and suspiciously perfect.
- 33:19Okay, that flower has teeth and it just
- 33:21caught a dragonfly. I officially take
- 33:22back the word peaceful. Let's leave
- 33:24before it discovers food delivery.
- 33:28[music]
- 33:33[music]
- 33:40>> I knew it. The cheese thief always
- 33:42returns to the scene of the crime.
- 33:44>> Thief is such a harsh word. I prefer
- 33:45Tiny Cheese Consultant.
- 33:47>> And what exactly does a cheese
- 33:48consultant do?
- 33:49>> Quality control. Very brave work.
- 33:51Extremely delicious work.
- 33:52>> Fine, but I am supervising the
- 33:54investigation.
- 33:57[music]
- 34:04[music]
- 34:08[music]
- 34:12>> [snorts]
- 34:18[music]
- 34:27[music]
- 34:31>> Frontline tested. Planet approved. This
- 34:34is the Vorax [music]
- 34:359. Dense alloy frame, stabilized plasma
- 34:38chamber, beautiful recoil. management in
- 34:41the right hands. It ends arguments
- 34:43before they begin. And yes, it comes in
- 34:45matte black.
- 35:01[music]
- 35:11>> [music]
- 35:16>> That is all for today. Leave a like and
- 35:18a comment if you found this useful.
- 35:20Thank you legends and everyone who
- 35:22supports this channel. You are amazing.
- 35:25Have a great day.
About this transcript
This page contains the full transcript of ComfyUI MiniMax H3: Best Video Generation Workflows (Ep29) by pixaroma, generated from the public captions YouTube serves with the video. The transcript has 5,914 words across 842 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.