MiniMax releases H3 omni-modal generative system with open weights

The 33B parameter H3 model supports text, image, audio, and video inputs, generating video with native stereo audio at up to 2K resolution and 15-second durations. It features a modular architecture including a context-processing system and a two-stage generation pipeline for high-fidelity output.
How this story unfolded
12 days · 8 reports · 136 community posts · 144 of 150 shown
- Jul 29
- Jul 31
- Aug 1
- Aug 2
- Aug 3
- Aug 4
- Aug 5
- Aug 7
- Aug 8
- Aug 9
- Aug 10
MiniMax by email
Get an email when MiniMax has news
No news that day, no email.
More stories today
- Open-source course teaches phone agent call center with FastRTC and Twilio
- The Rise of the 1 am Job Interview
- Viseron offers self-hosted AI NVR for object and face detection
- Making Knowledge Distillation Cheap Enough to Run at Scale
- How to turn any Claude agent into a 24/7 employee with MCP