Google DeepMind launched Gemini 3.5 Transcribe, its most precise speech-to-text model for intelligent voice interactions. Artificial Analysis reports 4.0% streaming WER and 2.6% non-streaming WER, with 85+ languages and screen-aware dictation in Antigravity that grounds filenames, variables, and agent thoughts.

Key Takeaways

  • โœ“4.0% streaming WER and 2.6% non-streaming WER; ~70% faster time-to-final-transcription vs Chirp 3, with 5.50% FLEURS streaming WER
  • โœ“Live API model gemini-3.5-transcribe-live for sub-second bidirectional streaming; Interactions API gemini-3.5-transcribe adds speaker attribution and word-level timestamps
  • โœ“Live in Google AI Studio, Antigravity with screen/chat grounding, Gemini on macOS, and Gboard Rambler; custom vocabulary plus auto-detect for 85+ languages