YouTube2Text

ComfyUI Tutorial MiniMax H3 on RTX 3060 6GB VRAM – Optimized Low VRAM Workflow + LTX 2 3 Upscaling — Transcript

by CG Pixel · 2,722 words · 370 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hello everyone, welcome back to the
  2. 0:01channel. Since the release of the LTX
  3. 0:04model, AI, video generation has become
  4. 0:06much more accessible to the users with
  5. 0:07consumer grade GPUs. Now we have a
  6. 0:10powerful new model called Minimax H3
  7. 0:12which is shaping up to be the strong
  8. 0:14competitor to the LTX model. Miniax H3
  9. 0:18supports a wide range of generation
  10. 0:19model including text to video, image to
  11. 0:22video, first frame to last frame
  12. 0:23animation, reference to video, and even
  13. 0:25video editing, making it one of the most
  14. 0:28versatile OpenAI video model available.
  15. 0:30In today tutorial, we will put Minimax
  16. 0:32H3 to the test to see how it will
  17. 0:34perform. I will also show you how I
  18. 0:36managed to make it run on the RTX 3060
  19. 0:39with only 6 GB of VRAM using an
  20. 0:41optimized low VRAM workflow and some
  21. 0:43additional nodes. Finally, I will
  22. 0:45demonstrate how to upscale the generated
  23. 0:47video using the LTX 2.3 model to achieve
  24. 0:50much higher quality results. So, without
  25. 0:52further ado, let's get started. Okay,
  26. 0:53the model has day zero support in Comfy
  27. 0:56Y and uh they claim to be an openweight
  28. 0:59model with native audio and the 2K video
  29. 1:02resolution, meaning that you can run
  30. 1:04this model to make it generate uh 2K
  31. 1:07resolution on a local uh RTX 3060. So if
  32. 1:11you scroll down a little bit, you can
  33. 1:12see that it is a next generation
  34. 1:14openweight video model that can allows
  35. 1:16you to go up to 2K and with the video
  36. 1:19duration of 15 seconds, which
  37. 1:21practically the same as the LTX model.
  38. 1:24However, it is a third generation video
  39. 1:26model following high low 01 and 02. From
  40. 1:29what I saw, it is the first model that
  41. 1:31can rival the CES 2.5. So we have here
  42. 1:34the text to video prompt only the image
  43. 1:36to video where you will uh animate an
  44. 1:39image the first and last frame video and
  45. 1:42reference to video. The model also have
  46. 1:44context understanding. So it should uh
  47. 1:46have a better prompt adurance and
  48. 1:49understanding compared to all previous
  49. 1:51model that we saw. We have also a better
  50. 1:53audio quality. And one of the most
  51. 1:56important thing is the editing and the
  52. 1:58motion transfer. As you all know, the
  53. 2:00motion transfer in the edex isn't its uh
  54. 2:03strong point even if you try with the
  55. 2:06other workflow or even if you try with
  56. 2:08Laura and other editing workflow.
  57. 2:11However, this one can do that easily.
  58. 2:13So, you can add a reference video that
  59. 2:15can supply movement, camera move,
  60. 2:17performance, cutting, rhythm and etc.
  61. 2:19Here we have some examples and they also
  62. 2:23provide the prompt so you can test them
  63. 2:25by yourself once you have the workflow.
  64. 2:27And let's talk now about the model. As
  65. 2:30you can see the first main downside of
  66. 2:32this uh minimax is the size of the
  67. 2:35model. You can see that the BF16 is
  68. 2:37waiting 17 GB and it is also the main
  69. 2:41model not to mention the text encoder
  70. 2:44and the VAE. And the first step of using
  71. 2:46this workflow is downloading all the
  72. 2:48necessary model. So let me take you to
  73. 2:51the config
  74. 2:53link. You can see here that we have the
  75. 2:55diffusion model, the text encoder and
  76. 2:58the VAE. So let's take a look here. We
  77. 3:00have a lot of diffusion model. You can
  78. 3:03also see that we have the first last
  79. 3:05frame to video which can do both text to
  80. 3:08video and image to video. And for video
  81. 3:11editing, you can use this reference to
  82. 3:13video model. But for now, I did not use
  83. 3:16those models. You can also see that we
  84. 3:18have a VAE that has different size. The
  85. 3:22BF16 is also has also a very uh
  86. 3:25important weight. And lastly, we will
  87. 3:27grab here the VAE. Make sure to download
  88. 3:30the audio and the video VAE first. Once
  89. 3:33it is done, you can go ahead and use
  90. 3:35this rebuilt AI uh GG web version. Uh
  91. 3:39this guide did a good work on this
  92. 3:41minimax model and it also has YouTube
  93. 3:44channel that you can check if you want
  94. 3:46to see some new tutorial or models. So
  95. 3:49here what we have we have the the text
  96. 3:51encoder GG version which clearly reduce
  97. 3:55its size. I directly downloaded this 8.5
  98. 3:59GB version. You can do the same. And
  99. 4:01here you can see that we also have the
  100. 4:04first last frame which can do image and
  101. 4:07text to video. So you I directly
  102. 4:10downloaded this Q3 version which has a
  103. 4:14lower size. Once everything is done, you
  104. 4:16can place the VAE under the VAE folder.
  105. 4:20Next, you can use the GG version. Place
  106. 4:23it under the unit folder. As for the
  107. 4:25text encoder, place it directly under
  108. 4:28the clip folder here. Now that we have
  109. 4:30everything, we need to install some
  110. 4:32additional node for optimization. And we
  111. 4:34will start with this Confy Spectrum MX.
  112. 4:38So, make sure to grab the code here. Go
  113. 4:40to confy root folder. Search for custom
  114. 4:43nodes on the search bar. Type in cmd get
  115. 4:46clone then paste your code here. Click
  116. 4:49enter and it will install everything for
  117. 4:51you. So this node is going to help you
  118. 4:54to boost up your uh video generation by
  119. 4:5725%. It works by reducing the expensive
  120. 5:01H3 transformer evaluation during the
  121. 5:04sampling steps. So it will helps you a
  122. 5:06lot while generating your video. So once
  123. 5:10we have everything, make sure to restart
  124. 5:12Confy and we are good to go. Okay, now
  125. 5:15let's talk about the workflow. As you
  126. 5:17can see here, we have three main groups,
  127. 5:20the text to video, the first last frame
  128. 5:22video, and finally the upscaling group
  129. 5:26using the LTX model. So what you can do
  130. 5:29first is bypassing all those group by
  131. 5:32using the bypass button. And let's start
  132. 5:35focusing on the first last frame
  133. 5:37workflow first. As you can see here, we
  134. 5:40have our first image and our second
  135. 5:42image. In order to generate the video, I
  136. 5:45need first to downscale the resolution
  137. 5:47of my images. After that, you can see
  138. 5:49that we having a subgraph here that
  139. 5:52regroup all the minimax uh workflow. But
  140. 5:55before that, make sure to select all the
  141. 5:57necessary model. We have the VAE here,
  142. 5:59the audio VAE, the GGF or the main model
  143. 6:03that will generate our video and the
  144. 6:06clip uh text encoder in order to make
  145. 6:09the prompt understandable by the model.
  146. 6:12We also here have some parameters or the
  147. 6:14video parameters starting with the
  148. 6:16duration. It is set to 5 seconds. Make
  149. 6:19sure to leave it as it is if you have
  150. 6:21the same uh configuration setup as me
  151. 6:23otherwise it will take you more longer.
  152. 6:25And we have this resolution selector
  153. 6:27here. The model can generate and adapt
  154. 6:30to many resolution. If uh for example
  155. 6:32here we have the square but you can
  156. 6:34choose the portrait, the portrait
  157. 6:36standard, the wide screen and etc. The
  158. 6:39model will generate the video for you
  159. 6:41without uh any artifact or without
  160. 6:43crashing. And that was also the case for
  161. 6:45the LDX model. So on that point they are
  162. 6:48practically the same. So let's take a
  163. 6:51closer look at the subgraph. By clicking
  164. 6:53on this button, you can see here that we
  165. 6:56have all the main model and in order to
  166. 6:58optimize the workflow and avoid
  167. 7:00crashing, I use this clean VM after each
  168. 7:04model. So once the model is load, the
  169. 7:06VRAM is cleared in order to save for you
  170. 7:09some additional VRAM. And next we have
  171. 7:12the conditioning. Here you can see that
  172. 7:14this conditioning will need the clip,
  173. 7:16the VA, the first frame, the last frame
  174. 7:19and this green dot over here is uh
  175. 7:21referring to the prompt section uh that
  176. 7:24is outside the subgraph. We also have
  177. 7:26the width and the height and the length
  178. 7:28of the our video. So once all these data
  179. 7:31is processed, we will end up with the
  180. 7:34positive output and our latent for the
  181. 7:37video here that is going to be used
  182. 7:39directly by the sampler. So our next
  183. 7:42step is the sampling steps which use a
  184. 7:44simpler custom advanced notes with the
  185. 7:47uler set as simpler because uh combined
  186. 7:50with the spectrum nodes it gives you the
  187. 7:53perfect uh boost or the highest boost
  188. 7:55while generating your video. So for the
  189. 7:57scheduleuler I am using a simple one. Uh
  190. 8:00the steps they must be set to 20 or 25
  191. 8:04between those two but 20 also give me a
  192. 8:07good results. You can also try with 15
  193. 8:10in order to test out the model. Then
  194. 8:12once you are satisfied with the result,
  195. 8:14you can try to increase it to 20 or 25.
  196. 8:17Okay. Once everything is done, the
  197. 8:19output here going to be used in two
  198. 8:22different way. The first one is going to
  199. 8:23decode our video using the video decoder
  200. 8:26and the second one is going to decode
  201. 8:29our audio using the audio decoder in
  202. 8:31order to create our audio. After that,
  203. 8:33we're going to save our video using this
  204. 8:36notes create a video and get video
  205. 8:38component. But along the way, I also
  206. 8:40added some clean VRAM or clean caches in
  207. 8:43order to avoid the VRM issues or the VRM
  208. 8:46crashes. Once everything is set up, we
  209. 8:48will have our video saved. So now let's
  210. 8:51let's talk more about this spectrum
  211. 8:53apply minimax nodes. As you can see
  212. 8:55here, uh using these two nodes allowed
  213. 8:57me to boost my video generation by 25%.
  214. 9:01meaning now the model is more faster and
  215. 9:04use less VRAM compared to the previous
  216. 9:07or the default workflow. Make sure to
  217. 9:09place that node after the unit loader or
  218. 9:12the the fusion model loader and this one
  219. 9:16is going to be connected directly here
  220. 9:18to the basic guider and the basic
  221. 9:20scheduleuler. So once uh it is connected
  222. 9:23that way it will uh really helps you
  223. 9:25boosting your video generation. Okay,
  224. 9:28that's the first subg graph or the
  225. 9:30workflow. As for the text to video, you
  226. 9:33can see that it is practically the same
  227. 9:35workflow. Here we have our resolution
  228. 9:37selector and the video duration, this
  229. 9:40the model and practically the same
  230. 9:41parameters. However, as you can see
  231. 9:44here, the minimax H3 image to video, we
  232. 9:47don't necessarily upload first or last
  233. 9:50frame. You can leave it as it is and it
  234. 9:52will automatically generate uh video
  235. 9:54based on text. One more important thing
  236. 9:57is the prompting and it also has its
  237. 10:00unique pro way of prompting. As you can
  238. 10:02see here for this first last frame
  239. 10:04animation. You can see that we are
  240. 10:06describing the some motion the style of
  241. 10:09the video and etc. Next we have the
  242. 10:12scene overview and the story board. So
  243. 10:15you can write here a storyboard like we
  244. 10:17saw in the LTX director by providing the
  245. 10:20the video duration much uh like here for
  246. 10:23example we set to much frame or watch
  247. 10:26the first frame our video and we will
  248. 10:28describe the storyboard or the video
  249. 10:31during uh all those uh time-lapses. You
  250. 10:34will also need to include the camera and
  251. 10:36the audio. As you can see here, I added
  252. 10:39clear audio, woman, voice over,
  253. 10:41speaking. Exactly. And I add my
  254. 10:43dialogue. As for the camera, you can see
  255. 10:45that I use some keywords like single
  256. 10:47smooth, backward tracking, zoom, no cuts
  257. 10:50and etc. As for the text to video, you
  258. 10:52can see that uh this one was provided by
  259. 10:55the default workflow of comfy UI. It
  260. 10:57also has the same uh principle like you
  261. 11:00will describe the style, the scene
  262. 11:01overview, the storyboard, the camera,
  263. 11:03audio and etc.
  264. 11:06I also included some prompt for you if
  265. 11:08you want to use the same and all you
  266. 11:10have to do is using this string over
  267. 11:12here and click run. While prompting make
  268. 11:15sure to use the same uh prompt uh
  269. 11:18structure otherwise you copy one prompt
  270. 11:21paste it to sh GBT or any uh other AI
  271. 11:24and ask him to give you uh other example
  272. 11:27of using the same prompt structure of
  273. 11:29this one. So once everything is set, you
  274. 11:31can start generating your video which
  275. 11:33will bring us to the results section of
  276. 11:35this video tutorial. Okay, before
  277. 11:37showing you the results, I want to talk
  278. 11:39to you about the resolution which is the
  279. 11:42most important factor for this video
  280. 11:44generation. For the first last frame, I
  281. 11:46used megapixel of 0.4 and you can use
  282. 11:49this table in order to see the
  283. 11:51resolution of your final video. Of
  284. 11:54course, I jumped to 0.5 and it works uh
  285. 11:58for me, but not the first time. I first
  286. 12:01run out of RAM and but when I insisted
  287. 12:04on that by clicking on the run button,
  288. 12:06it managed to work. It took me more time
  289. 12:08to generate. As for the text to image,
  290. 12:10you can directly go ahead and use 0.5
  291. 12:13MP. it will uh generate your video more
  292. 12:17easily because the text to video doesn't
  293. 12:19necessarily need need to encode the
  294. 12:22image in order to animate it which can
  295. 12:24grant you to generate the video at a
  296. 12:26higher resolution compared to the image
  297. 12:28to video. As for the generated time, I
  298. 12:31leave this note here where I recorded
  299. 12:33every generation time. So for the LDX
  300. 12:37upscaler, it took me around 8 minute in
  301. 12:40order to upscale or double the
  302. 12:41resolution of a video generated at 0.5
  303. 12:44MP. As for the text to video, it took
  304. 12:47around 20 minutes to generate a video
  305. 12:49with 0.5 megapixel. And lastly, we have
  306. 12:52the first last frame which took 35
  307. 12:54minutes in order to generate a video at
  308. 12:5704 megapixel. To generate high video
  309. 13:00resolution, the best solution is using
  310. 13:02the text video and it takes upscale in
  311. 13:04order to gain more time and avoid
  312. 13:07running out of your RAM. As for the
  313. 13:09video generated, I started with this
  314. 13:11mouse here and you can clearly see the
  315. 13:13quality of video and the quality of the
  316. 13:15movement that looks more smoothly and
  317. 13:18can even generate the same object at
  318. 13:20different perspective which is really
  319. 13:22awesome by the way. Next, I jump to this
  320. 13:25commercial linking and you can see that
  321. 13:27it also has some unique way of
  322. 13:30generating the video. The movement looks
  323. 13:32more smoother and also the camera is uh
  324. 13:35pretty awesome compared to what we used
  325. 13:37to saw on the LTX. After that, I tried
  326. 13:41with this uh anime superhero video and
  327. 13:44it also give me this results which also
  328. 13:47looks pretty good. You can see that uh
  329. 13:50we can generate a lot of styled and we
  330. 13:52can also include the text on the video
  331. 13:55which is new things by the way that
  332. 13:57looks very awesome too. After that I use
  333. 14:00this prompt of Chinese woman that looks
  334. 14:02on the mask and here we have some dragon
  335. 14:05animation here. The results is not uh
  336. 14:08pretty awesome on this one. However uh
  337. 14:11if you try to increase the resolution
  338. 14:13you will end up with something very
  339. 14:15stunning. So right now this Miniax model
  340. 14:18seems to have a very good uh video
  341. 14:21generation. It has better camera
  342. 14:23movement and actor movement compared to
  343. 14:25LTX 2.3 and it can adapt to a lot of
  344. 14:29styles while generating a text directly
  345. 14:31incorporated on the video which is a new
  346. 14:34feature too. However, I'm pretty sure
  347. 14:37we'll have more optimization regarding
  348. 14:39uh the steps or also the resolution in
  349. 14:42order to adapt it more easily to all
  350. 14:45users. Okay, that was my first review on
  351. 14:47this Miniax model. We saw that this
  352. 14:49workflow can run with 6 GB of VRAM and
  353. 14:53even if it reproduce lower resolution uh
  354. 14:55video quality, we can upscale it using
  355. 14:57this ATX upscaler in order to get a
  356. 15:00decent quality. I will release more
  357. 15:02video on this model and I will try to
  358. 15:04give you more impressive and more
  359. 15:07stunning results next time. If you like
  360. 15:09this video, please push the like button
  361. 15:11for me, subscribe to my channel, leave
  362. 15:13me some comments down below, and don't
  363. 15:15forget to become a Patreon member of my
  364. 15:16Patreon page where you can get early
  365. 15:18access to my workflow and ask me for
  366. 15:20additional help. You can also go ahead
  367. 15:22and become a YouTube premium subscriber
  368. 15:24in order to get early access to my
  369. 15:26workflow too if you can't access to
  370. 15:28Patreon page. So thank you.

About this transcript

This page contains the full transcript of ComfyUI Tutorial MiniMax H3 on RTX 3060 6GB VRAM – Optimized Low VRAM Workflow + LTX 2 3 Upscaling by CG Pixel, generated from the public captions YouTube serves with the video. The transcript has 2,722 words across 370 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.