Back to news
launchGoogle2026-08-26

Google launches Gemini 3.5 Transcribe with 2.6% word error rate across 85+ languages

On August 26, Google DeepMind released Gemini 3.5 Transcribe, posting 2.6% WER non-streaming and 4.0% WER streaming on Artificial Analysis benchmarks, with support for 85+ languages, custom vocabulary, and attribution for up to three speakers.

Google DeepMind released Gemini 3.5 Transcribe on August 26. According to Artificial Analysis benchmarks, the model achieves a 2.6% word error rate (WER) on pre-recorded audio and 4.0% on real-time streaming, reducing time to final transcription by 70% compared to the previous Chirp 3 system.

Gemini 3.5 Transcribe supports over 85 languages with automatic language detection and handles multilingual mixed-language scenarios. The model recognizes speaker self-corrections (e.g., "Tuesday — no, Wednesday"), automatically removes filler words like "um" and "uh," and converts unstructured speech directly into formatted text.

The model supports custom vocabulary for specialized jargon, special spellings, and alphanumeric combinations such as postal codes and order numbers. Pre-recorded audio scenarios support attribution for up to three speakers (3+ speakers is experimental) with word-level timestamps. The model can delegate tasks (such as image generation and file analysis) to other Gemini models via function calls.

Gemini 3.5 Transcribe is available in preview through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Real-time streaming runs via the Live API (identifier gemini-3.5-transcribe-live), pre-recorded audio via the Interactions API (identifier gemini-3.5-transcribe). It has launched in the Android Gboard Rambler dictation feature and in the Mac Gemini app, with Chrome support coming soon.

Early enterprise customers include Vivo, Intellitek Health, and Lingopal.

speech-to-textaudiomultimodaltranscriptionfrontier-model