AnalysisAI ModelsAugust 13, 2026

MiDashengLM-Gen generates unified audio scenes via LLM-driven flow matching

Read original source →arxiv.org

MiDashengLM-Gen is an end-to-end framework using a pre-trained LLM and audio tokenizer as backbone, with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. It generates coherent 16 kHz audio scenes blending speech, music, and sound effects.

2 sources

More stories today

Open the live feed