Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Paper introduces a Fréchet distance loss on speech representations for training few-step diffusion and flow-matching TTS models. The loss complements local objectives like conditional flow matching, improving global naturalness of generated speech.
1 source
Music by email
Get an email when there's news on Music
No news that day, no email.
More stories today
- Lyte closes $165M round at $1.6B valuation
- Meta settlement could clear way for new AI product launches
- Z.ai opens first Tmall store for AI subscriptions
- Fable 5.1 Max users share setup tips and warnings
- Opinion: Next DSM should assess algorithms' role in eating disorders