MiniMax H3 is native in ComfyUI, day zero with the weights.
Model highlights:
→ Text-to-video, prompt only
→ Image-to-video
→ First-and-last-frame - control the opening frame, the closing frame, or both
→ Reference-to-video - carry a subject, a motion, or a voice through the clip
→ In-place editing - modify a shot you already have instead of regenerating it
The capability
@MiniMax_AI leads with is multimodal context understanding, and it's what collapses those five tasks into one model.
H3 takes images, audio, and video together and resolves them against a prompt that explains how they relate.
Download the workflows & learn more in our blog 👇