Google Gemini Transcribe Cleans Speech-To-Text Output
Google’s Gemini 3.5 Transcribe adds automatic formatting, filler-word removal, vocabulary adaptation, speaker labels and word-level timestamps across more than 85 languages.

Google has added a Gemini 3.5 Transcribe model that can clean up speech-to-text output by formatting transcripts, adapting to specialized vocabulary and removing filler words, The Verge reported.
The update gives Gemini Audio a product-specific transcription path rather than only a voice interaction layer.
Gemini 3.5 Transcribe can work across more than 85 languages, identify specialized jargon and let users provide custom vocabulary so names, technical terms and unusual spellings are not corrected by hand after recording.
Google framed the model as an advance over Chirp 3, its earlier transcription system, especially for multilingual performance and wording error rates.
The product claim remains Google-owned: the company says users can “edit naturally with just your voice,” while the transcription layer can automatically format text and remove words such as “um” and “uh.”
Speaker handling is part of the workflow.
The model can attribute speech for up to three speakers in pre-recorded audio and provide word-level timestamps, turning the service into a tool for meetings, interviews and recorded production work where searchable timing and speaker labels matter as much as a raw transcript.
The launch also shows how Google is splitting Gemini Audio into several surfaces.
Gemini 3.5 Transcribe follows 3.5 Live Translate, while the broader Gemini 3.5 Pro model that Google had promised for June still has not arrived.
The delayed live models were described as a separate real-time layer.
Gemini 3.5 Live was presented as better at handling mid-sentence interruptions, language recognition and live visual processing, while Gemini 3.5 Live Experimental would narrate its progress step by step as it worked through a task.
Google initially paired the transcription release with updates to Gemini 3.5 Live and 3.5 Live Experimental, but later told The Verge those additional models were not launching yet and did not provide a new date.
For now, the usable change is the transcription model, with the live-interaction upgrades still outside release.
That sequencing leaves transcription available while the live audio updates remain delayed.




















