MiniMax H3 generates coherent 30-second music from text prompts

A r/StableDiffusion demo shows MiniMax H3 producing up to 30 seconds of coherent audio with custom lyrics, composition structure, instruments, and genres. Prompts use a MEDIA/SCENE/TIMELINE structure; the example generates a 1990s upbeat hip-hop rap track at 32x32 resolution, 20 steps.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Sequoia Capital invests in AI-native video platform Preview
- US Launches Effort to Speed Trade in AI Goods Between Allies
- DeepMind launches SL2T sign language-to-text model
- Liquid AI releases LFM2.5-VL-3B vision-language model for edge
- Grok and Meta's release discussed on ETN podcast episode