MiniMax releases H3 omni-modal video model with native stereo audio

MiniMax H3 generates 15-second 2K clips with native stereo audio, reading text, images, video, and audio as one unified context. It is positioned as a general-purpose multimodal generation model, not a text-to-video model with add-ons.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AWS Quick and fal enable agentic creative workflows
- Anthropic opens 10,000 free Claude seats for scientists
- Researcher breaks Claude Code Opus 5 auto mode with 80% success
- Nvidia CEO Jensen Huang: I wish I had invested more in AI frontier labs
- Apple introduces rubric-based alignment for grounded QA