Google released Gemini 3.5 Transcribe, a speech-to-text model that returns cleaned and formatted transcripts rather than raw audio—stripping filler words, resolving self-corrections, and applying custom vocabulary. Google cites word error rates of 4.0% streaming and 2.6% non-streaming against its own Chirp 3, with time-to-final-transcription reduced by 70%. It runs as two separate APIs—a sub-second Live API and a pre-recorded Interactions API with word-level timestamps and up to three-speaker attribution—supporting more than 85 languages in public preview via Google AI Studio and the Gemini Enterprise Agent Platform.





