NVIDIA releases NemotronLabs VoiceChat 11B speech-to-speech model

NVIDIA released NemotronLabs VoiceChat 11B, an open 11B end-to-end speech-to-speech model with ~450 ms turn-taking and live tool calling. It performs streaming speech understanding and generation in one unified network, supporting transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech.
How this story unfolded
3 weeks · 2 reports · 2 community posts · from Jul 21
- Jul 21
- Jul 22
- Aug 4
- Aug 10
NVIDIA by email
Get an email when NVIDIA has news
No news that day, no email.
More stories today
- GOP panics over Big Tech ties as Trump shifts on AI regulation
- Ethan Mollick: Claude's skill creator beats ChatGPT for reusable skills
- Aident Loadout gives agents 27,000+ tools and logs every action
- Etched gains sizable fan base for AI inferencing computers
- Corbell generates technical specs from repository knowledge graphs