MiniMax has released H3, a multimodal generation model that can take in text, images, video, and sound and then generate or edit video from a single natural-language instruction.
Users can have H3 learn camera movement from one video, use a character from a second image, and reference the voice from a third audio clip. What used to require separate tools for motion transfer, character reference, audio reference, and video editing can now be packed into the same job.
H3 generates up to 15 seconds of video at 2K resolution with native stereo sound. Through the official API, 2K costs $0.13 per second — roughly $1.95 for a 15-second clip — while 768P runs $0.09 per second.
A single task accepts up to nine images, three videos, and three audio clips, capping out at 12 files total.
MiniMax plans to open up the model weights in the coming days.