Intelligent transcription with Gemini 3.5 Transcribe
Intelligent transcription with Gemini 3.5 Transcribe
Our latest speech-to-text model designed for precise and intelligent real-time transcription.
Senior Director, Engineering, Gemini Audio
Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team
Your browser does not support the audio element.
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:
Get more precise and intelligent transcription
Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.
Experience smart transcription and advanced dictation
In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.
By leveraging the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.
Companies like Vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.
Related Stories
Grand Rounds August 21, 2026: Off
47 minutes ago
AI News
Evaluate any agent framework with Amazon Bedrock AgentCore Evaluations
47 minutes ago
AI News
Bill Gates Issues Stark Warning About Artificial Intelligence: 'We Are Not Preparing Adequately'
47 minutes ago
AI News
India’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call
1 hour ago
AI News
Journalism, media, and technology trends and predictions 2025
1 hour ago
AI News
Journalism, media, and technology trends and predictions 2024
1 hour ago
AI News
Artificial Intelligence Brings TILs Into the Digital Era
1 hour ago
AI News
OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm
1 hour ago