Google releases Gemini 3.5 Transcribe speech-to-text model

Gemini 3.5 Transcribe reports 2.6% average WER across 85+ languages. It offers two endpoints: gemini-3.5-transcribe for pre-recorded audio via the Interactions API and gemini-3.5-transcribe-live for real-time streaming via the Live API, with sub-second latency.
2 sources
Google DeepMind by email
Get an email when Google DeepMind has news
No news that day, no email.
More stories today
- David Lowery, Jason Isbell sue Suno over likeness rights
- OpenAI to launch next model soon, Altman says
- Koray Kavukcuoglu discusses AGI path and Gemini 3.7 Flash in podcast
- Palo Alto CEO: $1T of cybersecurity infrastructure isn't ready for AI
- Perplexity CEO teases 'Private, Personal, Powerful AI'