MiniMax H3 omni-modal model generates 2K video with native stereo audio

H3 understands text, image, video, and audio inputs and generates video with native stereo audio at up to 2K resolution over 15 seconds. It combines three modules — H3-Context-IR, H3-Base (768p), and H3-Regenerate-2K — and is available via API, app, and open weights on Hugging Face.
1 source
Moonshot AI by email
Get an email when Moonshot AI has news
No news that day, no email.
More stories today
- Mollick: ChatGPT Work, Claude Cowork should explain choices like a PM
- Deedy Das: AI-written prose can evade detection
- Podcast discusses the risks of AI-driven team velocity
- Alex Kantrowitz examines why Big Tech is falling behind in AI
- Stanford researchers change how AI agents access files