[ComfyUI]Breakthrough Release: MiniMax H3, the New Leading Open-Source Video Model — Transcript
Full transcript
- 0:05[whistles]
- 0:11[whistles]
- 0:13Wow.
- 0:18[music]
- 0:21Fore!
- 0:23[music]
- 0:32[snorts]
- 0:38[snorts]
- 0:40Foreign! Foreign!
- 1:00Hi everyone, I'm Jing Chen. It is like
- 1:03New Year's in the open source community.
- 1:05Although the release was delayed by a
- 1:06few hours, that doesn't take away from
- 1:08the fact that it is open source. Haha.
- 1:11Today I am bringing you Miniaax H3, a
- 1:14multimodal video model. It was released
- 1:16just 2 days ago, so I haven't had time
- 1:18to test it in depth yet. From this
- 1:20initial round of testing, I feel it can
- 1:23already comfortably beat some closed
- 1:25source models. Breaking the dominance of
- 1:27closed source models. The consistency is
- 1:29great and the motion is great, too. This
- 1:32is the kind of result we never would
- 1:33have imagined before. So, I'm really
- 1:36happy about it. Let's start with the
- 1:38most important thing and what everyone
- 1:40cares about most, how to speed it up.
- 1:43Right now, I most recommend the pruned
- 1:46intate model combined with the MVFP4
- 1:48clip. This is the fastest combination
- 1:50and it works well with most people
- 1:52setups. Even lower machines have a
- 1:55chance of running it. You only need 8 GB
- 1:58of RAM. As long as you have enough
- 2:00system memory, then set aside some RAM
- 2:02in the startup parameters and give it a
- 2:04try. Then add such attention after the
- 2:06model. If you haven't installed Sage
- 2:08yet, I really recommend installing it.
- 2:10There is also this prediction node. If
- 2:12the video doesn't have a lot of motion,
- 2:14you can turn it on to reduce the amount
- 2:16of actual sampling in the middle
- 2:18section. In my test with no acceleration
- 2:21enabled around 1 megapixel, which is 768
- 2:25by 1376,
- 2:27generating a 10-second video took about
- 2:3030 to 34 minutes on running hub plus
- 2:32with a 48 GB 490 setup. After adding
- 2:36such attention, the time was cut in
- 2:39half. It took around 14 to 16 minutes to
- 2:42generate. Then with the prediction node
- 2:44added 768p10.
- 2:47Second video took about 10 minutes
- 2:49online. That is already extremely fast.
- 2:52I also tested the easy cache node
- 2:54mentioned by other creators.
- 2:57Although both methods speed things up by
- 2:59reducing full model computations. After
- 3:01comparing them, I personally feel the
- 3:04prediction node gives better quality.
- 3:06You can try them separately as well.
- 3:08Let's look at a same seed comparison.
- 3:11The middle one uses saggot attention
- 3:13only. The one on the left uses a
- 3:15detention plus the prediction node.
- 3:18There actually isn't much difference
- 3:19between the two. For this level of
- 3:21motion, the prediction acceleration
- 3:24performs quite well. The one on the
- 3:26right uses easy. You can feel that the
- 3:28distant people look noticeably
- 3:29blurriier. If you're making a high
- 3:31motion video, I recommend turning on
- 3:33prediction acceleration to quickly
- 3:35search for a good seed. Once you find a
- 3:37result with suitable motion,
- 3:38composition, and seed, turn off the
- 3:40prediction node and run it again fully
- 3:43with the same seed to preserve as much
- 3:45final detail as possible. This time I've
- 3:48prepared three workflows, text to video,
- 3:51first and last frame video, and full
- 3:53reference video. The text to video and
- 3:56first and last frame workflows are
- 3:57fairly simple. You can download them
- 3:59locally and explore them yourselves.
- 4:01Today, I'll mainly cover the full
- 4:03reference workflow, which is the most
- 4:04useful one.
- 4:06Full reference workflow uses the
- 4:08reference to video model.
- 4:11It can take reference images, reference
- 4:13videos, and reference audio at the same
- 4:15time, which can control the characters,
- 4:18scene, style, motion, camera movement,
- 4:21and voice. The size and duration
- 4:23settings are standard.
- 4:26Just compare them with the table on the
- 4:27side. In general, 0.4 megapixels should
- 4:31work on most machines. That is roughly
- 4:33480p. The 0.8 and 1.0 megapixel settings
- 4:38look sharper, but they also require more
- 4:40resources and take longer. You can keep
- 4:42adding image, video, and audio inputs on
- 4:45the reference node. The node expands
- 4:47automatically. According to the official
- 4:50documentation, you can add up to nine
- 4:52images, three audio clips, and three
- 4:54videos. That is more than enough for
- 4:56normal use. Every additional reference
- 4:59makes the generation take longer,
- 5:01especially reference videos. They can
- 5:03increase the runtime significantly. Just
- 5:06add references according to your actual
- 5:08needs.
- 5:09The workflow itself isn't complicated.
- 5:12The really difficult part is the prompt.
- 5:14The full reference workflow handles
- 5:15images, videos, audio, character
- 5:18relationships, motion, camera work, and
- 5:21sound all at once. If you only write a
- 5:23few simple sentences, it is easy to make
- 5:25the roles of the references unclear. It
- 5:28might even say things you can't
- 5:29understand. So this time I've prepared a
- 5:31system prompt template. It automatically
- 5:34identifies whether you're providing
- 5:36text, images, videos, or audio. Then
- 5:39completes it into a prompt suitable for
- 5:41H3.
- 5:43When using it, just put this system
- 5:46prompt into a web- based vision model.
- 5:48Then send it your video requirements and
- 5:50all your reference materials. Briefly
- 5:52describe the story and result you want,
- 5:54and it will organize a prompt you can
- 5:56copy directly into the H3 workflow.
- 6:03This template was organized and refined
- 6:05based on several official prompt guides.
- 6:08If you're interested, you can also open
- 6:10the official documentation to learn
- 6:12more. I've put the links in the
- 6:14description. The download package will
- 6:16include the pruned model and the
- 6:18smallest clip. For local use, I
- 6:21recommend starting with the smaller
- 6:22clip. If you have plenty of RAM online,
- 6:25you can also try the intake clip. ASA
- 6:28I'll share the three workflows with
- 6:29acceleration nodes already connected
- 6:32along with the system prompt template
- 6:33and the prediction acceleration node.
- 6:36I'll share all of them with you. That is
- 6:38all for now. Go ahead and start playing
- 6:40with it. Open sourcing a video model at
- 6:43this level significantly speeds up the
- 6:45open-source community's efforts to catch
- 6:47up with closed source models and clearly
- 6:49lowers the cost of video generation. I
- 6:52think this already counts as a major
- 6:53generational leap for open source. I'll
- 6:56continue testing the models techniques
- 6:57and prompt writing methods in depth and
- 7:00share more with you later. That is it
- 7:01for this video. If you enjoyed it,
- 7:04remember to like and subscribe. See you
- 7:06next time.
About this transcript
This page contains the full transcript of [ComfyUI]Breakthrough Release: MiniMax H3, the New Leading Open-Source Video Model by jingchen573, generated from the public captions YouTube serves with the video. The transcript has 1,014 words across 169 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.