YouTube2Text

ComfyUI MiniMax H3: Best Video Generation Workflows (Ep29) — Transcript

by pixaroma · 5,914 words · 842 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[music]
  2. 0:02Welcome to another video tutorial. Today
  3. 0:05I will show you how to use the new
  4. 0:07Miniax H3 video model so you can create
  5. 0:11awesome videos.
  6. 0:13>> Welcome to my Comfy UI setup where every
  7. 0:15cable has a purpose probably.
  8. 0:19Perfect. I have now connected everything
  9. 0:21exactly the way I was not supposed to.
  10. 0:25So I guess we are buying more VRAM.
  11. 0:28>> What is your name? I am Pixab Bunny.
  12. 0:32>> Where are you coming from?
  13. 0:42>> You always look at me like you
  14. 0:43understand everything. Are you my little
  15. 0:45best friend?
  16. 0:46>> Of course I am. But a best friend also
  17. 0:48deserves treats.
  18. 0:49>> So that is what this is about. You are
  19. 0:51adorable and very clever.
  20. 0:53>> Clever, cute, and hungry. That is my
  21. 0:55full talent set.
  22. 0:57>> Milo, be honest. Do you think I am
  23. 0:59overthinking everything again?
  24. 1:01>> Yes, but in a very talented way.
  25. 1:04>> That is not exactly comforting, but
  26. 1:06somehow it helps.
  27. 1:22As you know, I use the Comfy UI easy
  28. 1:24install like in episode 1 because it has
  29. 1:27a lot of useful scripts that make things
  30. 1:28easier. So, before I run Miniax, I have
  31. 1:31to do some updates. I will update easy
  32. 1:33install because that will also update my
  33. 1:36Pixaroma nodes. If you have another
  34. 1:38Comfy UI version, just update Pixar Roma
  35. 1:41nodes manually with GitPull or from the
  36. 1:43manager if it works. I asked for it to
  37. 1:45be added to the manager again since it
  38. 1:47was removed for a while because of the
  39. 1:49problem I had with GitHub in the past.
  40. 1:51Press enter to exit when it shows the
  41. 1:53message. Then you need to update Comfy
  42. 1:56UI. You should have a similar BAT file
  43. 1:59somewhere, either in the update folder
  44. 2:01or in the Comfy UI folder that lets you
  45. 2:04update Comfy UI. This is necessary
  46. 2:07because the miniax nodes are new and you
  47. 2:09do not have them in older comfy UI
  48. 2:11versions. Then to speed things up, it
  49. 2:14helps to have Sage attention. If you do
  50. 2:17not have it with easy install, it is
  51. 2:19quite easy to install. From the add-ons
  52. 2:22folder, run this BAT file. And if all
  53. 2:24these things are green, it should work
  54. 2:26correctly. If you get messages in red,
  55. 2:28then maybe your video card is not
  56. 2:30supported. This is for Windows and
  57. 2:32Nvidia cards. I had some problems with
  58. 2:3420 series video cards, but on the 30
  59. 2:37series, 40 series, and 50 series, it
  60. 2:40should work correctly. So, after you run
  61. 2:43that, you will have this extra comfy UI
  62. 2:45startup BAT file with Sage Attention. If
  63. 2:48we take a quick look inside, we can see
  64. 2:51it has these arguments. So, it knows to
  65. 2:53start with Sage Attention. This means we
  66. 2:56do not need to add any extra Sage
  67. 2:58Attention nodes to our workflow. In
  68. 3:00fact, if you have Sage Attention nodes
  69. 3:03in your workflow, it is better to start
  70. 3:05Comfy UI normally and let the nodes do
  71. 3:07the job. But I prefer this method
  72. 3:09because it gives me smaller workflows
  73. 3:11and it works the same. So now in the
  74. 3:14launcher, I can start either the desktop
  75. 3:17or browser version of Comfy UI with Sage
  76. 3:20Attention. You can see that when it
  77. 3:22starts, it says it is using Sage
  78. 3:24Attention. I added a version check
  79. 3:27Pixaroma node here so you can see which
  80. 3:30Comfy UI version and which version of
  81. 3:32Pixaroma nodes I have at the moment of
  82. 3:35recording. Make sure you have this
  83. 3:37version or a newer one, not an older
  84. 3:39version. Then you can download the free
  85. 3:42workflows in the Comfy UI folder. Look
  86. 3:45for the user folder, then default and
  87. 3:48finally workflows. You can place the
  88. 3:50downloaded folder with the workflows
  89. 3:52there. I removed the resources and left
  90. 3:55only the actual workflows in that
  91. 3:57folder. So now in Comfy UI, you should
  92. 4:00see the new workflows. If not, use this
  93. 4:03refresh button. You have them organized
  94. 4:05in the same order that I present them in
  95. 4:07this tutorial. Some of you may have
  96. 4:09noticed some new buttons here, like a W
  97. 4:12letter and a question mark. This new
  98. 4:14button or the alt + W shortcut opens the
  99. 4:18workflow manager that I built for you.
  100. 4:20This is a long video, so I will keep
  101. 4:23this part short. Basically, you can move
  102. 4:26and arrange workflows, rename them, and
  103. 4:28do pretty much anything you want with
  104. 4:30them, so you can stay organized. But the
  105. 4:33best way to get started is to click this
  106. 4:35help button and read about everything it
  107. 4:37can do. It has a large help section, so
  108. 4:39it should include the basics you need to
  109. 4:41figure it out. You can click W again to
  110. 4:43close it. You also have this help file
  111. 4:46for all Pixaroma nodes so you can easily
  112. 4:48find out how to use them and what they
  113. 4:50do. I hope the Comfy UI team creates
  114. 4:53something similar for their nodes as
  115. 4:55well to make things easier. So, click
  116. 4:57around, search, and look at what each
  117. 4:59node does. You can also find a lot of
  118. 5:02other useful information there. If you
  119. 5:04do not want to see those buttons, you
  120. 5:06can go to settings, then look for
  121. 5:08Pixarroma and scroll down until you find
  122. 5:10these settings. From here, you can hide
  123. 5:14all the buttons or only the ones you
  124. 5:16want. You can use the shortcut for
  125. 5:18workflows and the help section can also
  126. 5:20be opened from each Pixaroma node. For a
  127. 5:23line, you need to have it enabled for it
  128. 5:25to work. So, let me move on to the
  129. 5:27MiniAX H3 model. First, I will open the
  130. 5:31folder with the workflows for this
  131. 5:32model. By the way, you can add workflows
  132. 5:35to favorites, and you can also
  133. 5:37right-click them to see more options,
  134. 5:39including one that lets you add a
  135. 5:41preview cover for that workflow. If you
  136. 5:43select a workflow, like this texttovideo
  137. 5:46workflow, it will show you a preview and
  138. 5:48useful information on the right. Scroll
  139. 5:51down and you can see which model it uses
  140. 5:54and a lot of other information. You can
  141. 5:56also add a note to the workflow so you
  142. 5:58can search for it by that note later.
  143. 6:00Double click a workflow to open it or
  144. 6:03use the open button. I tried to make the
  145. 6:05workflows simple and easy to use and to
  146. 6:08reduce the number of nodes where
  147. 6:09possible. All the things you need for
  148. 6:11these workflows are in this note. The
  149. 6:13most important one is the Miniax H3
  150. 6:16model which is quite big and you can
  151. 6:18click this button to see more versions.
  152. 6:20So here is a list of them. There are
  153. 6:22actually two models with different
  154. 6:24quantizations. One is first last frame
  155. 6:27to video and one is reference to video.
  156. 6:30No matter what video card you have, even
  157. 6:32if it is powerful, I still recommend
  158. 6:35that everyone uses this model, the
  159. 6:37pruned int 8 version. The rest are
  160. 6:39slower with not much improvement in
  161. 6:42quality. For the reference model, use
  162. 6:44this version as well, but I already have
  163. 6:46the correct model selected in the
  164. 6:47downloaded workflow, so you do not have
  165. 6:49to worry about it. Then there is one
  166. 6:51thing that many people do not seem to
  167. 6:53talk about, the license. If we read the
  168. 6:56license, we can see that some
  169. 6:58territories are excluded from it. The
  170. 7:00European Union, United Kingdom, Korea
  171. 7:03and United States of America are not
  172. 7:06covered by this license. I assume this
  173. 7:08is happening because of legislation in
  174. 7:10those countries and instead of having
  175. 7:12problems, they added this condition to
  176. 7:14the license. But does that mean no one
  177. 7:16from those countries can use it? I added
  178. 7:19another link here where you can read
  179. 7:21more about the license and they explain
  180. 7:23it in more detail. They explained that
  181. 7:25they had to choose between not releasing
  182. 7:28the model until it complied with the
  183. 7:29legislation in those countries or
  184. 7:32releasing it while excluding those
  185. 7:33countries or something similar. I am
  186. 7:36from Europe, so how can I use it if
  187. 7:39Europe is excluded? There is a link here
  188. 7:41to an application form. It asks for an
  189. 7:44organization. I am not sure if this also
  190. 7:47applies to people without an
  191. 7:48organization, but you probably do not
  192. 7:50need this license anyway if you only use
  193. 7:52it for fun and personal use. So, I
  194. 7:55clicked this application form and
  195. 7:56completed my details. I have the
  196. 7:58Pixaroma company, which is a oneperson
  197. 8:01company where I do everything myself.
  198. 8:03Then, it asks about annual revenue and
  199. 8:05obviously I do not make that amount of
  200. 8:07money. After that, you have to check a
  201. 8:10lot of terms confirming that you agree
  202. 8:12to comply with the license. I am not a
  203. 8:14lawyer but I think this means they give
  204. 8:17the responsibility to you so they cannot
  205. 8:19be held accountable for any problems
  206. 8:21caused by how you use the model. This
  207. 8:24keeps them safe and I am happy that I
  208. 8:26can use the model so everyone is happy.
  209. 8:29That is how I did it so I could use the
  210. 8:31models on my channel. I received an
  211. 8:33email confirming that I was approved and
  212. 8:35that was all. So let me show you how to
  213. 8:37get the models. Use this button to
  214. 8:40download the model. It will take a while
  215. 8:42because of its size. After downloading
  216. 8:44it, navigate to your Comfy UI folder. Go
  217. 8:47to models, then diffusion models, and
  218. 8:50create an H3 folder to keep things
  219. 8:52organized. Place the model for this
  220. 8:54workflow inside that folder. I have two
  221. 8:57models because I also downloaded the
  222. 8:59reference model for the other workflows.
  223. 9:01Now, you have the Miniax model. Next, we
  224. 9:04need a text encoder. Download this Quen
  225. 9:07model which is also big but very smart
  226. 9:09with 32 billion parameters. Until now we
  227. 9:13used models with four and 8 billion
  228. 9:15parameters. Place it in the text
  229. 9:17encoders folder. For the VAE we have two
  230. 9:20models because this model also supports
  231. 9:23audio. Download both of them into the
  232. 9:25VAE folder and then you can load them
  233. 9:27into these nodes. As for the custom
  234. 9:30nodes, you know me, I try to keep things
  235. 9:32simple. Only Pixaroma nodes are used. I
  236. 9:35also included a custom chat GPT link I
  237. 9:38made for prompting, but I will talk more
  238. 9:40about that later. After all the models
  239. 9:42have finished downloading, you can press
  240. 9:44R to refresh the node definitions, and
  241. 9:47everything should work if the files have
  242. 9:49the same names and are placed in the
  243. 9:51correct folders. If not, you will need
  244. 9:54to select them manually inside the
  245. 9:55nodes. Like I said, I like simple
  246. 9:58workflows, so compared with what you
  247. 10:00find on the internet, this should be
  248. 10:02much easier to use. I added some steps
  249. 10:05in blue for you to follow. And this
  250. 10:07workflow only needs three steps from
  251. 10:08you. Let me start with the first one
  252. 10:10where you can choose the orientation and
  253. 10:12size. I added some of my favorite sizes
  254. 10:15here, which I recommend depending on
  255. 10:17your video card. The more VRAM you have,
  256. 10:20the larger the size you can generate.
  257. 10:22People have tested this even on a video
  258. 10:24card with 8 GB of VRAM, and it worked if
  259. 10:26they had enough system RAM and dynamic
  260. 10:28VRAM was enabled. Remember to switch
  261. 10:31between portrait and landscape since
  262. 10:33that decides the ratio. Because we are
  263. 10:35making videos, I did not include a
  264. 10:37square ratio, but you can add one as
  265. 10:40well. If you go to settings, you can add
  266. 10:42the sizes you want here, change their
  267. 10:44order, remove sizes, or add more. I
  268. 10:47included all the sizes the model can
  269. 10:49generate, and I added a star next to the
  270. 10:52ones that I think work well. For the
  271. 10:54optimal size, I think this one is the
  272. 10:56best, but I did many of my tests using
  273. 10:59this size. Miniax also prefers sizes
  274. 11:02that are multiples of 32, so I enabled
  275. 11:05that option. No matter what size you
  276. 11:07enter, it will round it to a compatible
  277. 11:10value. So, that is the Pixar sizes node.
  278. 11:13I only changed the title to make it
  279. 11:15easier to use. For this video recording,
  280. 11:17I will select portrait and this size.
  281. 11:20Then, I created another node for the
  282. 11:22duration. If we go to settings, you can
  283. 11:25add your custom durations here in
  284. 11:26seconds. I also added a few ready-made
  285. 11:29formulas directly into the node. By
  286. 11:32default, it uses this Miniax H3 formula,
  287. 11:36but you can use others for different
  288. 11:37workflows or create your own formula.
  289. 11:40The cool part is that you can turn it
  290. 11:41into a slider if you want and adjust the
  291. 11:44values or you can just type them if you
  292. 11:46prefer that. I like this button version
  293. 11:48because I can add only the values I
  294. 11:50need. So, I will use a 3-second video
  295. 11:53for this test. Then comes the long
  296. 11:56detailed prompt. The more details you
  297. 11:58add, the better it can understand what
  298. 12:00you are asking for. You can write all of
  299. 12:02that by hand, but I cannot write prompts
  300. 12:04that are so detailed and structured in a
  301. 12:07way that AI understands perfectly. So,
  302. 12:09why not use AI to create prompts for AI,
  303. 12:12right? Even if a short prompt might work
  304. 12:15well, you give the model too much
  305. 12:17freedom to invent things. That is why I
  306. 12:19built this custom chat GPT. You can
  307. 12:22click on it and use it. It should also
  308. 12:24work with a free subscription, but you
  309. 12:26can use another LLM if you want. I will
  310. 12:29also provide some prompt formulas inside
  311. 12:31the workflow archive. Since I created
  312. 12:33this GPT, I can also edit it so you can
  313. 12:36see how I set it up. In the configure
  314. 12:38tab, I added instructions telling it to
  315. 12:41use the attached knowledge. Here I
  316. 12:44uploaded this text file with the long
  317. 12:46formula that explains how to write the
  318. 12:48prompts. The formulas are also inside
  319. 12:50this folder that you downloaded with the
  320. 12:52workflows. There is one for the
  321. 12:54reference model and one for text to
  322. 12:56video and first last frame to video. You
  323. 12:59can look at them, change them, and adapt
  324. 13:01them however you want. So back to the
  325. 13:04custom GPT, you need to tell it what
  326. 13:06type of prompt you want. Is it for text
  327. 13:08to video? Is it for first frame to
  328. 13:10video? or is it for first frame last
  329. 13:12frame to video. This helps it choose the
  330. 13:15correct prompt structure. In this case,
  331. 13:17I want a textto video prompt. Then I
  332. 13:20describe in my own words what I want and
  333. 13:22I hope it understands me. And I got a
  334. 13:24long detailed prompt for that. Some
  335. 13:27parts look like instructions, but that
  336. 13:29is fine because it helps the model
  337. 13:31understand what I want it to do. You can
  338. 13:33see how it describes everything and also
  339. 13:35how it uses English. I tried it with
  340. 13:38other languages and it works. It says it
  341. 13:40supports 11 languages, but I also tried
  342. 13:43Romanian and it worked. Sometimes it
  343. 13:46mixed up one word, but overall it worked
  344. 13:48well. I added a run timer here so I can
  345. 13:51see how long it takes, but more recently
  346. 13:53I prefer using a run log because I can
  347. 13:56see the history for each run. Let me run
  348. 13:58it and see how it works. It loads all
  349. 14:01the models and then the prompt and
  350. 14:03settings go into this new miniax node
  351. 14:05created by the Comfy UI team. After
  352. 14:08that, it goes into the K sampler where I
  353. 14:10used 20 steps and set CFG to one because
  354. 14:13we do not have a negative prompt. For
  355. 14:15the sampler, I used res multi-step. And
  356. 14:18for theuler, I used simple. Some people
  357. 14:21also recommend beta. So, you can try
  358. 14:23that one as well. Then, I ran into a
  359. 14:26problem and got an error. What could it
  360. 14:28be, I wonder? I recorded this because
  361. 14:30some of you might get the same error,
  362. 14:32and I will show you what fixed it for
  363. 14:34me. What I did was close Comfy UI. Then
  364. 14:39inside the BAT file I used to start
  365. 14:41Comfy UI, I had this argument that
  366. 14:44disabled dynamic VRAMm. You could remove
  367. 14:47it manually and save the file, but to
  368. 14:49make sure all the BAT files are set
  369. 14:50correctly and have that argument
  370. 14:52removed, I will do something else. If
  371. 14:54you use a GGUF model, you might want
  372. 14:56dynamic VRAM disabled. But in this case,
  373. 14:59I want it enabled. So, I go to add-ons,
  374. 15:02then tools, and toggle it easily using
  375. 15:05this BAT file. Now, it removed that
  376. 15:07argument. If I go back and check, you
  377. 15:10can see that the argument is no longer
  378. 15:12there. So, everything should work now.
  379. 15:14But let me test it. I open comfy UI
  380. 15:17again and run the workflow. And now I
  381. 15:19got the video I described.
  382. 15:23>> Do you like my cat?
  383. 15:28>> Do you like my cat?
  384. 15:30I will run it again and then I can add a
  385. 15:33label for each recorded runtime. By the
  386. 15:35way, you can now rightclick the node, go
  387. 15:38to node settings and enable this option.
  388. 15:41Then you can close the settings and look
  389. 15:44now it shows my video card along with
  390. 15:46the VRAM and RAM I have. Pretty cool,
  391. 15:49right? And the second video looks like
  392. 15:52this.
  393. 15:53>> Do you like my cat?
  394. 15:58Do you like my cat?
  395. 16:00>> I also did some tests while I was not
  396. 16:02recording so the recording would not
  397. 16:04affect the generation times and here is
  398. 16:07what I got. All the tests are for a
  399. 16:105-second texttovideo workflow. I did 10
  400. 16:13runs and for each run I increased the
  401. 16:16resolution to see how it affected the
  402. 16:18total running time. The higher the
  403. 16:20resolution, the more time it takes. But
  404. 16:22the increase is not linear. For example,
  405. 16:24for this resolution, it took around 3
  406. 16:26minutes, but for full HD, it took 10
  407. 16:29minutes. So, I prefer to create a
  408. 16:3115-second video at a lower resolution
  409. 16:33instead.
  410. 16:37[music]
  411. 16:48[music]
  412. 16:50Heat. Heat.
  413. 16:52[music]
  414. 16:57[music]
  415. 17:03[music]
  416. 17:09[music]
  417. 17:14>> [music]
  418. 17:25[music]
  419. 17:26>> So you saw that the higher the
  420. 17:28resolution the better it looks but it
  421. 17:30also takes more time to generate. Let me
  422. 17:33move now to the first frame to video
  423. 17:35workflow since we also want to control
  424. 17:37the first frame so we can animate
  425. 17:38exactly the image we want. This workflow
  426. 17:41uses the same models and nodes. There is
  427. 17:44one extra thing here, an image, and that
  428. 17:47image is connected to the first frame
  429. 17:49input in the miniax node. After the
  430. 17:51image is loaded, step two is to choose
  431. 17:54the duration. The longer the video, the
  432. 17:56more time it takes, but at least the
  433. 17:59increase is more linear. So, if you know
  434. 18:01the generation time for 3 seconds, it
  435. 18:03will take around twice as long for 6
  436. 18:05seconds, three times as long for 9
  437. 18:07seconds, and so on. To make things
  438. 18:10easier, instead of constantly changing
  439. 18:12the video size, checking the image
  440. 18:14ratio, and adjusting everything
  441. 18:16manually, we can do all of that with a
  442. 18:18single click. So, I built another node
  443. 18:21called longest side. It takes the image
  444. 18:23from the load image node, makes the
  445. 18:25dimensions multiples of 32, and sets the
  446. 18:28longest side of the video to the value
  447. 18:30you choose here while keeping the
  448. 18:32original ratio. Then, step four is the
  449. 18:34prompt. For longest side, we have some
  450. 18:37settings here. You can see that for this
  451. 18:39workflow I used 32. You can change that
  452. 18:43from here by clicking the button
  453. 18:44multiple times. You can choose the
  454. 18:47longest side of the video from here. I
  455. 18:49added the recommended values, but you
  456. 18:51can add more. You can even use a
  457. 18:54different ratio if you want to crop the
  458. 18:56image. For example, if I want a square
  459. 18:59video, I can select that. Then from the
  460. 19:02settings, I can choose whether the crop
  461. 19:04should be centered or positioned
  462. 19:06differently. In this case, because the
  463. 19:08woman is near the top, a centered crop
  464. 19:11would remove her face. So, a top crop
  465. 19:13would work better. But I do not plan to
  466. 19:15crop the image. So, most of the time, I
  467. 19:17prefer to keep a ratio that is as close
  468. 19:20as possible to the original image. For
  469. 19:22the size tabs, you can also add your own
  470. 19:24values here. So, I can change them
  471. 19:27however I want. And then the buttons are
  472. 19:29updated here. The next time I click that
  473. 19:31button, it will use the new size. I will
  474. 19:34put it back to the default settings for
  475. 19:36this workflow. Then for the prompt, I
  476. 19:39usually take a screenshot of the load
  477. 19:41image node. After that, I go to chat GPT
  478. 19:44and open the custom chat GPT. I ask for
  479. 19:48a first frame to video prompt, paste the
  480. 19:50screenshot, and then describe what I
  481. 19:52want to happen. Chat GPT is quite good
  482. 19:55at this. Now, I have a long prompt that
  483. 19:57I can copy and paste here. I will use 5
  484. 20:00seconds and then run the workflow. I
  485. 20:03just noticed that I forgot to remove the
  486. 20:04square crop. So, I will cancel the
  487. 20:06generation, set it to keep the original
  488. 20:09ratio because I do not want to crop and
  489. 20:11run it again. It took around a minute
  490. 20:13and a half. Let me check the result.
  491. 20:18Well, hello there little sweetheart.
  492. 20:23Well, hello there little sweetheart.
  493. 20:25>> If I check the prompt, you can see that
  494. 20:27it used the English tag there along with
  495. 20:29the words in English. But like I said,
  496. 20:32you can use other languages, too. Here
  497. 20:35is an example in Romanian.
  498. 20:49>> [laughter]
  499. 20:52>> I did tests with 5, 10, and 15-second
  500. 20:55videos, and the time is around double
  501. 20:58for each extra 5 seconds, but the
  502. 21:00maximum for miniax is 15 seconds.
  503. 21:02Anyway, I tried to increase the
  504. 21:04resolution to this one. And now for 15
  505. 21:07seconds, I got 13 minutes. So, it is
  506. 21:09almost like 1 minute of running time per
  507. 21:11second of video at this resolution. Here
  508. 21:14are a few more tests for image to video.
  509. 21:17And it is almost the same as text to
  510. 21:19video just with a few extra seconds,
  511. 21:21maybe around five extra seconds at HD
  512. 21:23resolution. I now have some volunteers
  513. 21:26on Discord who test my workflows called
  514. 21:29bug hunters. And here is an example of
  515. 21:31tests on 16 GB VR RAM. They tried it
  516. 21:35with Sage Attention and also with some
  517. 21:37extra nodes. With easy cache, it reduced
  518. 21:40the time, but the results were worse. So
  519. 21:43I prefer the quality. The generation
  520. 21:45times with Sage attention as arguments
  521. 21:48and with nodes are similar. So I prefer
  522. 21:51using the arguments until someone
  523. 21:52releases better nodes. It is time to
  524. 21:55move to the next workflows that use
  525. 21:57besides the first frame also a last
  526. 21:59frame so you have control over where it
  527. 22:02ends as well. For this example, I tried
  528. 22:05two different bunny images to see if it
  529. 22:07can do a transition. So I started with
  530. 22:09this bunny sitting as the first frame.
  531. 22:11Then this ballerina bunny as the second
  532. 22:13frame. I set it to four seconds. I let
  533. 22:16it keep the ratio and this longest side
  534. 22:19and that is taken from the first frame
  535. 22:21since that usually decides the ratio of
  536. 22:24the image. So that first frame is
  537. 22:26connected here and then it goes into the
  538. 22:28miniax node into the first frame input.
  539. 22:31The second image is connected to the
  540. 22:33last frame input. I will run this
  541. 22:35workflow and just to give you an idea of
  542. 22:38how I prompted it, I took a screenshot
  543. 22:40of both nodes and in chat GPT I asked
  544. 22:43for a prompt for first frame last frame
  545. 22:45to video. Then you add how you want the
  546. 22:47transformation to happen. I wanted a
  547. 22:505-second video and it gave me a long
  548. 22:52prompt and the result for this workflow
  549. 22:55looks like this.
  550. 22:57[music]
  551. 23:02[music]
  552. 23:04For the next workflow, you can use it to
  553. 23:06get the last frame of a video. So, you
  554. 23:08decide how the video should end. So, for
  555. 23:10example, I have this ballerina bunny. I
  556. 23:13set the duration to 4 seconds and I kept
  557. 23:16the ratio and generated the video at
  558. 23:18this longest side. With a long prompt,
  559. 23:20the image should be connected only to
  560. 23:22the last frame input. And you leave the
  561. 23:24first frame input empty so the model can
  562. 23:26invent that part. So, if I run this
  563. 23:29workflow, I get this result.
  564. 23:33>> [music]
  565. 23:39>> So it is very easy to make the video end
  566. 23:42on a certain image and you do not have
  567. 23:44to generate two images all the time. So
  568. 23:46all these four workflows use the first
  569. 23:48model the first last frame to video
  570. 23:50model. Now I am moving to one that uses
  571. 23:53reference instead. I have one workflow
  572. 23:55here for two images and one for three
  573. 23:58images. Let me start with the two image
  574. 24:00one first. Make sure you downloaded this
  575. 24:02reference video model because it will
  576. 24:04not work correctly with the first one.
  577. 24:06Then for chat GPT, I made a different
  578. 24:09one because reference works a little
  579. 24:11differently. It does not start with a
  580. 24:13certain frame or end with it. It looks
  581. 24:15at the references and generates
  582. 24:16something new from them. So for the
  583. 24:19prompt, I will just take a screenshot of
  584. 24:21these two images and then in chat GPT I
  585. 24:24paste it and ask for a reference to
  586. 24:26video prompt for these two images. Then
  587. 24:28you describe how you want it. I already
  588. 24:30did that for this long prompt and the
  589. 24:32result is this.
  590. 24:36>> Are you ready to escape this summer?
  591. 24:41>> Are you ready to escape this summer?
  592. 24:43>> That is just a simple example, but you
  593. 24:45can do more complex things. Let me show
  594. 24:48you another example with three images.
  595. 24:50It can do more than three. You just
  596. 24:52connect more images there in the miniax
  597. 24:54node. It can do a maximum of nine
  598. 24:57images, three short clips of around 5 to
  599. 24:5915 seconds, but for me that took forever
  600. 25:02to generate. I am not sure if that is a
  601. 25:04bug, but I could not get it working
  602. 25:06properly. It can also use three audio
  603. 25:09files with a maximum of 12 files
  604. 25:11combined. I will show audio examples in
  605. 25:14a few minutes. Since this uses images as
  606. 25:17references, the ratio of the images does
  607. 25:20not matter that much. So, I can use
  608. 25:22custom ratios here. So decide on the
  609. 25:25maximum resolution you want to use and
  610. 25:27do not forget to choose portrait or
  611. 25:29landscape. I tell you this because I
  612. 25:31forgot many times to switch it. I could
  613. 25:34make it follow the ratio of the first
  614. 25:36image like in the other workflows, but
  615. 25:38then I have to keep in mind that the
  616. 25:40first image must have the correct ratio.
  617. 25:42This way I know I can use images with
  618. 25:45any ratio and I choose the video ratio
  619. 25:47myself. So, I want this bunny wearing
  620. 25:50that orange outfit in an exotic location
  621. 25:53like the last image. I used chat GPT for
  622. 25:56the prompt and the result is this one.
  623. 25:58>> Hello from paradise.
  624. 26:03Hello from paradise.
  625. 26:08>> Add more load image nodes there. Connect
  626. 26:10them and save your own workflows based
  627. 26:12on these. Now let me move to the next
  628. 26:15group which can also use an audio
  629. 26:17reference combined with an image. It
  630. 26:19does not do audio to audio. So I have
  631. 26:22one workflow here for speaking and one
  632. 26:24for singing. Let me start with the
  633. 26:26singing workflow first. Like I said it
  634. 26:29uses the same reference video model. You
  635. 26:31set the duration of the video and I made
  636. 26:34it so that it also affects the uploaded
  637. 26:36audio which helps you get better
  638. 26:37results. From this load audio pixeloma
  639. 26:40node, you can upload an audio file or
  640. 26:43select one from the input folder. You
  641. 26:45can see that when I select a different
  642. 26:47duration here, it also adjusts this
  643. 26:50orange selection. To make that work, the
  644. 26:52seconds output is connected here. So,
  645. 26:55the selected audio section always
  646. 26:56matches the chosen duration.
  647. 26:59[music]
  648. 27:04second selection to the right to start
  649. 27:06at a different point.
  650. 27:10>> So everything is controlled. If I
  651. 27:12disconnect the seconds input now I have
  652. 27:14more control and can adjust it manually
  653. 27:16however I want. I am showing you this
  654. 27:19because it can be useful for other
  655. 27:20workflows. But for this one it is easier
  656. 27:23to leave the seconds input connected.
  657. 27:26Also check the settings and help
  658. 27:28sections for more information. If you
  659. 27:30set a duration longer than the actual
  660. 27:32audio, you can either add silence or
  661. 27:34loop the audio. So, I will reconnect the
  662. 27:37seconds input. Let me pick 5 seconds and
  663. 27:40play it to see which part will be used
  664. 27:41for singing by looking at that white
  665. 27:43line.
  666. 27:45>> [music]
  667. 27:54[music]
  668. 27:58>> Maybe I will make it shorter and use 3
  669. 28:00seconds.
  670. 28:03[music]
  671. 28:04>> Okay, that part seems good for the
  672. 28:06input. I have a universal prompt here
  673. 28:08that works with almost anything. You can
  674. 28:10adjust it if needed, but it should work
  675. 28:12most of the time as it is. Now I will
  676. 28:15run it. I forgot to mention that you set
  677. 28:17the size of the final video from here
  678. 28:19using the longest side. This workflow
  679. 28:22also has this new H3 audio sync node
  680. 28:24that I made. If you click more, you can
  681. 28:27see that it helps keep the audio
  682. 28:29consistent. It is too complex for me to
  683. 28:31fully understand what it does, but
  684. 28:33Claude figured it out and I got good
  685. 28:35results. And the final result is this
  686. 28:38one. [music]
  687. 28:40[singing]
  688. 28:45You can adjust the prompt if you want
  689. 28:47some extra movement there. For example,
  690. 28:50I can add that the bunny puts his hands
  691. 28:51in the air while dancing or something
  692. 28:53similar. Then I run it again and I got
  693. 28:56this. [music]
  694. 29:01[singing]
  695. 29:03>> For some seeds, I got barely visible
  696. 29:05lines in the generated video. I am not
  697. 29:08sure if it is because I used a low
  698. 29:10resolution, if the model sometimes does
  699. 29:13that, or if there is a wrong setting
  700. 29:15somewhere. It is just something that
  701. 29:17appears sometimes, usually around the
  702. 29:19middle of the screen. Moving now to the
  703. 29:21second workflow, which is similar to the
  704. 29:23first one, but has a different prompt.
  705. 29:26So, I add the voice here.
  706. 29:29I can talk now since we have the MiniAX
  707. 29:31H3 model. Maybe I will make it shorter
  708. 29:34and use only 3 seconds. I can talk now
  709. 29:37since we have the Miniax H3 model. I set
  710. 29:40this size for the longest side so it
  711. 29:42generates faster. And now I will run it.
  712. 29:45You can see how I set up the prompt, but
  713. 29:47feel free to adjust it. Maybe you can
  714. 29:49find a better one, but this worked well
  715. 29:51for me. You also have more settings in
  716. 29:54the audio sync node. I made it so the
  717. 29:56audio does not go beyond 15 seconds
  718. 29:59since Miniax can only generate a maximum
  719. 30:0115-second video. Anyway, if the audio
  720. 30:04clip is longer, I set it to warn me, but
  721. 30:07you can also make it stop the run if you
  722. 30:09want. You also have the same options to
  723. 30:11add silence or loop the audio if the
  724. 30:13clip is too short. And the result is
  725. 30:16this one. I can talk now since we have
  726. 30:18the Miniax H3 model. I can talk now
  727. 30:20since we have the Miniax H3 model.
  728. 30:24Even if it is a video model, MiniAX can
  729. 30:26generate images or edit images, but it
  730. 30:29is a little slow. So I still prefer the
  731. 30:31Crea 2 model for that. But if you
  732. 30:34already have the model, at least you
  733. 30:36know it is an option. So for this
  734. 30:38portrait landscape node, I added a
  735. 30:41setting so it is constrained to
  736. 30:42multiples of 32. And you put your sizes
  737. 30:45there and just flip it to get the right
  738. 30:46ratio. You add a long prompt and you can
  739. 30:49run it. So it took 45 seconds for a full
  740. 30:52HD image and the result is not bad if
  741. 30:54you do full HD.
  742. 30:57If you reduce the size, you get faster
  743. 30:59generation, like only 17 seconds. But
  744. 31:02the result is not great, more like a
  745. 31:05frame from a lowquality video, which in
  746. 31:07fact it is. It generates like a batch of
  747. 31:10five frames and picks the first one,
  748. 31:12index zero, from that. I could not set
  749. 31:15the length to one because the node gave
  750. 31:17an error. I think Miniax likes five
  751. 31:19frames at once. And now let me try the
  752. 31:22edit workflow. It is similar to that
  753. 31:24one, but this time we load an image. And
  754. 31:27I added here what to change, but longer
  755. 31:30prompts that give details about what
  756. 31:32should change and what should stay the
  757. 31:33same could work better. I tried with a
  758. 31:36smaller size this time for edit using
  759. 31:38this one. It has the same get image from
  760. 31:40batch like before. And the loaded image
  761. 31:43is connected to the first frame. And
  762. 31:46this time it took around a minute. So it
  763. 31:48changed the hair like I asked, but it is
  764. 31:50losing a bit of quality. You can get
  765. 31:52away with small details, but sometimes
  766. 31:54it can change more things. For example,
  767. 31:57sometimes it changed the text to pink
  768. 31:59also. Not only the hair, but it depends
  769. 32:02on the seed. In this case, it worked
  770. 32:04even with a short prompt. So, do not
  771. 32:07forget to check the help for each
  772. 32:08Pixaroma node. So, you can learn more
  773. 32:11about the node. That help is updated
  774. 32:13when I update the node and change
  775. 32:15something in it. Search for the things
  776. 32:17you need. And I hope all these nodes are
  777. 32:19useful for you. I will leave you with a
  778. 32:21few more examples of videos generated
  779. 32:23with Miniax.
  780. 32:25>> Meet the Pixa time machine because
  781. 32:27waiting for the weekend is basically
  782. 32:29unbearable. Destination selected.
  783. 32:32Medieval era.
  784. 32:33>> Okay, definitely older than my
  785. 32:36apartment.
  786. 32:37Picks a time machine. Travel smarter.
  787. 32:40Everyone has a smartphone. I wanted
  788. 32:42something fresher.
  789. 32:44Hey, Apple phone. What's on my schedule?
  790. 32:46>> First a call, then a snack. Hopefully
  791. 32:49not in that order.
  792. 32:51>> Finally, a phone with appeal.
  793. 32:53>> I was going to [music] say that.
  794. 32:55>> Hey guys, welcome with me into this
  795. 32:56amazing jungle. It's so beautiful here.
  796. 32:59Look at this flower. The color is
  797. 33:00incredible. It's even more beautiful in
  798. 33:02real life. Oh my god, look, a little
  799. 33:04monkey. He's so cute. This is definitely
  800. 33:07the best part of the trip. Say bye,
  801. 33:09little guy. Welcome back to the vlog.
  802. 33:11Today we are exploring the jungle where
  803. 33:13everything is beautiful, wet, and
  804. 33:14probably trying to eat something. Look
  805. 33:16at these flowers. They are gorgeous,
  806. 33:17peaceful, and suspiciously perfect.
  807. 33:19Okay, that flower has teeth and it just
  808. 33:21caught a dragonfly. I officially take
  809. 33:22back the word peaceful. Let's leave
  810. 33:24before it discovers food delivery.
  811. 33:28[music]
  812. 33:33[music]
  813. 33:40>> I knew it. The cheese thief always
  814. 33:42returns to the scene of the crime.
  815. 33:44>> Thief is such a harsh word. I prefer
  816. 33:45Tiny Cheese Consultant.
  817. 33:47>> And what exactly does a cheese
  818. 33:48consultant do?
  819. 33:49>> Quality control. Very brave work.
  820. 33:51Extremely delicious work.
  821. 33:52>> Fine, but I am supervising the
  822. 33:54investigation.
  823. 33:57[music]
  824. 34:04[music]
  825. 34:08[music]
  826. 34:12>> [snorts]
  827. 34:18[music]
  828. 34:27[music]
  829. 34:31>> Frontline tested. Planet approved. This
  830. 34:34is the Vorax [music]
  831. 34:359. Dense alloy frame, stabilized plasma
  832. 34:38chamber, beautiful recoil. management in
  833. 34:41the right hands. It ends arguments
  834. 34:43before they begin. And yes, it comes in
  835. 34:45matte black.
  836. 35:01[music]
  837. 35:11>> [music]
  838. 35:16>> That is all for today. Leave a like and
  839. 35:18a comment if you found this useful.
  840. 35:20Thank you legends and everyone who
  841. 35:22supports this channel. You are amazing.
  842. 35:25Have a great day.

About this transcript

This page contains the full transcript of ComfyUI MiniMax H3: Best Video Generation Workflows (Ep29) by pixaroma, generated from the public captions YouTube serves with the video. The transcript has 5,914 words across 842 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.