Scaling Properties of Continuous Diffusion Spoken Language Models

Apple ML Research paper introduces the phoneme Jensen-Shannon divergence (pJSD) metric to quantify linguistic quality of spoken language models. Scaling continuous diffusion SLMs to 16B parameters with tens of millions of hours of data enables emotive, prosodic, multi-speaker multilingual speech, though long-form coherence remains challenging.
1 source
Apple by email
Get an email when Apple has news
No news that day, no email.
More stories today
- Anthropic's Mythos-class models to launch this fall with enterprise data controls
- GPT-Image-2 adds transparent background support in API preview
- Palantir called 'the sovereign AI company'
- MiniMax's Hailuo AI and Runway announce collaboration
- Google expands Antigravity AI coding agent beyond its IDE