YouTube2Text

[ComfyUI]Breakthrough Release: MiniMax H3, the New Leading Open-Source Video Model — Transcript

by jingchen573 · 1,014 words · 169 segments · language en · Watch on YouTube

Full transcript

  1. 0:05[whistles]
  2. 0:11[whistles]
  3. 0:13Wow.
  4. 0:18[music]
  5. 0:21Fore!
  6. 0:23[music]
  7. 0:32[snorts]
  8. 0:38[snorts]
  9. 0:40Foreign! Foreign!
  10. 1:00Hi everyone, I'm Jing Chen. It is like
  11. 1:03New Year's in the open source community.
  12. 1:05Although the release was delayed by a
  13. 1:06few hours, that doesn't take away from
  14. 1:08the fact that it is open source. Haha.
  15. 1:11Today I am bringing you Miniaax H3, a
  16. 1:14multimodal video model. It was released
  17. 1:16just 2 days ago, so I haven't had time
  18. 1:18to test it in depth yet. From this
  19. 1:20initial round of testing, I feel it can
  20. 1:23already comfortably beat some closed
  21. 1:25source models. Breaking the dominance of
  22. 1:27closed source models. The consistency is
  23. 1:29great and the motion is great, too. This
  24. 1:32is the kind of result we never would
  25. 1:33have imagined before. So, I'm really
  26. 1:36happy about it. Let's start with the
  27. 1:38most important thing and what everyone
  28. 1:40cares about most, how to speed it up.
  29. 1:43Right now, I most recommend the pruned
  30. 1:46intate model combined with the MVFP4
  31. 1:48clip. This is the fastest combination
  32. 1:50and it works well with most people
  33. 1:52setups. Even lower machines have a
  34. 1:55chance of running it. You only need 8 GB
  35. 1:58of RAM. As long as you have enough
  36. 2:00system memory, then set aside some RAM
  37. 2:02in the startup parameters and give it a
  38. 2:04try. Then add such attention after the
  39. 2:06model. If you haven't installed Sage
  40. 2:08yet, I really recommend installing it.
  41. 2:10There is also this prediction node. If
  42. 2:12the video doesn't have a lot of motion,
  43. 2:14you can turn it on to reduce the amount
  44. 2:16of actual sampling in the middle
  45. 2:18section. In my test with no acceleration
  46. 2:21enabled around 1 megapixel, which is 768
  47. 2:25by 1376,
  48. 2:27generating a 10-second video took about
  49. 2:3030 to 34 minutes on running hub plus
  50. 2:32with a 48 GB 490 setup. After adding
  51. 2:36such attention, the time was cut in
  52. 2:39half. It took around 14 to 16 minutes to
  53. 2:42generate. Then with the prediction node
  54. 2:44added 768p10.
  55. 2:47Second video took about 10 minutes
  56. 2:49online. That is already extremely fast.
  57. 2:52I also tested the easy cache node
  58. 2:54mentioned by other creators.
  59. 2:57Although both methods speed things up by
  60. 2:59reducing full model computations. After
  61. 3:01comparing them, I personally feel the
  62. 3:04prediction node gives better quality.
  63. 3:06You can try them separately as well.
  64. 3:08Let's look at a same seed comparison.
  65. 3:11The middle one uses saggot attention
  66. 3:13only. The one on the left uses a
  67. 3:15detention plus the prediction node.
  68. 3:18There actually isn't much difference
  69. 3:19between the two. For this level of
  70. 3:21motion, the prediction acceleration
  71. 3:24performs quite well. The one on the
  72. 3:26right uses easy. You can feel that the
  73. 3:28distant people look noticeably
  74. 3:29blurriier. If you're making a high
  75. 3:31motion video, I recommend turning on
  76. 3:33prediction acceleration to quickly
  77. 3:35search for a good seed. Once you find a
  78. 3:37result with suitable motion,
  79. 3:38composition, and seed, turn off the
  80. 3:40prediction node and run it again fully
  81. 3:43with the same seed to preserve as much
  82. 3:45final detail as possible. This time I've
  83. 3:48prepared three workflows, text to video,
  84. 3:51first and last frame video, and full
  85. 3:53reference video. The text to video and
  86. 3:56first and last frame workflows are
  87. 3:57fairly simple. You can download them
  88. 3:59locally and explore them yourselves.
  89. 4:01Today, I'll mainly cover the full
  90. 4:03reference workflow, which is the most
  91. 4:04useful one.
  92. 4:06Full reference workflow uses the
  93. 4:08reference to video model.
  94. 4:11It can take reference images, reference
  95. 4:13videos, and reference audio at the same
  96. 4:15time, which can control the characters,
  97. 4:18scene, style, motion, camera movement,
  98. 4:21and voice. The size and duration
  99. 4:23settings are standard.
  100. 4:26Just compare them with the table on the
  101. 4:27side. In general, 0.4 megapixels should
  102. 4:31work on most machines. That is roughly
  103. 4:33480p. The 0.8 and 1.0 megapixel settings
  104. 4:38look sharper, but they also require more
  105. 4:40resources and take longer. You can keep
  106. 4:42adding image, video, and audio inputs on
  107. 4:45the reference node. The node expands
  108. 4:47automatically. According to the official
  109. 4:50documentation, you can add up to nine
  110. 4:52images, three audio clips, and three
  111. 4:54videos. That is more than enough for
  112. 4:56normal use. Every additional reference
  113. 4:59makes the generation take longer,
  114. 5:01especially reference videos. They can
  115. 5:03increase the runtime significantly. Just
  116. 5:06add references according to your actual
  117. 5:08needs.
  118. 5:09The workflow itself isn't complicated.
  119. 5:12The really difficult part is the prompt.
  120. 5:14The full reference workflow handles
  121. 5:15images, videos, audio, character
  122. 5:18relationships, motion, camera work, and
  123. 5:21sound all at once. If you only write a
  124. 5:23few simple sentences, it is easy to make
  125. 5:25the roles of the references unclear. It
  126. 5:28might even say things you can't
  127. 5:29understand. So this time I've prepared a
  128. 5:31system prompt template. It automatically
  129. 5:34identifies whether you're providing
  130. 5:36text, images, videos, or audio. Then
  131. 5:39completes it into a prompt suitable for
  132. 5:41H3.
  133. 5:43When using it, just put this system
  134. 5:46prompt into a web- based vision model.
  135. 5:48Then send it your video requirements and
  136. 5:50all your reference materials. Briefly
  137. 5:52describe the story and result you want,
  138. 5:54and it will organize a prompt you can
  139. 5:56copy directly into the H3 workflow.
  140. 6:03This template was organized and refined
  141. 6:05based on several official prompt guides.
  142. 6:08If you're interested, you can also open
  143. 6:10the official documentation to learn
  144. 6:12more. I've put the links in the
  145. 6:14description. The download package will
  146. 6:16include the pruned model and the
  147. 6:18smallest clip. For local use, I
  148. 6:21recommend starting with the smaller
  149. 6:22clip. If you have plenty of RAM online,
  150. 6:25you can also try the intake clip. ASA
  151. 6:28I'll share the three workflows with
  152. 6:29acceleration nodes already connected
  153. 6:32along with the system prompt template
  154. 6:33and the prediction acceleration node.
  155. 6:36I'll share all of them with you. That is
  156. 6:38all for now. Go ahead and start playing
  157. 6:40with it. Open sourcing a video model at
  158. 6:43this level significantly speeds up the
  159. 6:45open-source community's efforts to catch
  160. 6:47up with closed source models and clearly
  161. 6:49lowers the cost of video generation. I
  162. 6:52think this already counts as a major
  163. 6:53generational leap for open source. I'll
  164. 6:56continue testing the models techniques
  165. 6:57and prompt writing methods in depth and
  166. 7:00share more with you later. That is it
  167. 7:01for this video. If you enjoyed it,
  168. 7:04remember to like and subscribe. See you
  169. 7:06next time.

About this transcript

This page contains the full transcript of [ComfyUI]Breakthrough Release: MiniMax H3, the New Leading Open-Source Video Model by jingchen573, generated from the public captions YouTube serves with the video. The transcript has 1,014 words across 169 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.