Google launches Gemini 3.5 Transcribe for over 85 languages
The new speech-to-text model identifies speakers and turns disfluent speech into clearer, more readable text
Google officially launched Gemini 3.5 Transcribe on August 26, 2026. The speech-to-text model supports more than 85 languages and is designed to transcribe both live speech and recorded audio.
Key features include Speaker Diarization, which identifies who said what, and Smart Transcription, which automatically formats transcripts for readability. Beyond transcribing speech verbatim, the model can handle stutters, filler words, and rambling sentences, then rewrite them into clearer text.
Gemini 3.5 Transcribe is the same model behind Rambler on Android, which Google unveiled at Google I/O 2026 and first introduced on Pixel 11. Rather than producing only word-for-word transcripts, it can turn natural speech into text ready to read or reuse.
Developers can access the model through the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. The Live API supports real-time streaming transcription, while the Interactions API processes recorded audio files.
Support for more than 85 languages could broaden its use for meetings, interviews, and captioning. However, its Thai transcription quality in real-world settings still needs further testing.