AnalysisAI ModelsAugust 13, 2026

MiDashengLM-Gen generates unified audio scenes via LLM-driven flow matching

MiDashengLM-Gen is an end-to-end framework using a pre-trained LLM and audio tokenizer as backbone, with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. It generates coherent 16 kHz audio scenes blending speech, music, and sound effects.

2 sources

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed