
MiniMax has released H3, an omni-modal generative model that takes in text, images, video, and audio, then generates or edits video from a single natural-language instruction.
With H3, you can have the model copy the camera movement from one video, use the character from a second image, and reference the audio from a third clip. Motion transfer, character reference, audio reference, and video editing — tasks that used to require separate pipelines — can now happen in a single task.
H3 can generate up to 15 seconds of video at 2K resolution with native dual-channel audio. The official API charges $0.13 per second for 2K output, putting a 15-second clip at roughly $1.95.