Moonshot AILaunchAI ModelsAugust 3, 2026

MiniMax H3 omni-modal model generates 2K video with native stereo audio

H3 understands text, image, video, and audio inputs and generates video with native stereo audio at up to 2K resolution over 15 seconds. It combines three modules — H3-Context-IR, H3-Base (768p), and H3-Regenerate-2K — and is available via API, app, and open weights on Hugging Face.

1 source

Moonshot AI by email

Get an email when Moonshot AI has news

No news that day, no email.

More stories today

Open the live feed