Google launches Gemini 3.5 Transcribe speech-to-text model

Gemini 3.5 Transcribe ranks #5 on AA-WER at 2.6% non-streaming and 4.0% streaming, processing ~84 seconds of audio per second at ~$5 per 1,000 minutes. It supports 85+ languages, multi-speaker attribution, custom vocab, and is available via Live and Interactions APIs in Google AI Studio.
16 sources
Intelligent transcription with Gemini 3.5 Transcribedeepmind.google
Introducing Gemini 3.5 Transcribe, our new speech to text model with smart transcription, function...x.com
Gemini 3.5 Transcribe now available on AI Gatewayvercel.com
Google announces Gemini 3.5 Transcribe for AI-powered speech-to-textarstechnica.com
Google’s new AI transcription edits out your ‘ums’ and ‘ahs’theverge.com
How to build with Gemini 3.5 Transcribeyoutube.com
Introducing Gemini 3.5 Transcribereddit.com
Google DeepMind by email
Get an email when Google DeepMind has news
No news that day, no email.
More stories today
- AWS Quick and fal enable agentic creative workflows
- Anthropic opens 10,000 free Claude seats for scientists
- Researcher breaks Claude Code Opus 5 auto mode with 80% success
- Nvidia CEO Jensen Huang: I wish I had invested more in AI frontier labs
- Apple introduces rubric-based alignment for grounded QA