Google rolls out Gemini 3.5 Transcribe to improve real-time dialogue and speech recognition
Google announced that Gemini 3.5 Transcribe replaces the older Chirp 3 engine and delivers a 4 % word‑error‑rate on live audio and 2.6 % on pre‑recorded clips, even in noisy or conversational settings. The model strips filler words, auto‑formats output, and can be instructed to invoke other Gemini capabilities—such as image generation—through function calls. It also supports custom vocabularies, speaker‑level timestamps for up to three participants, and seamless language detection for more than 85 languages and dialects. Google is embedding the service in Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini macOS app, Gboard, and will expose it via the Gemini API, AI Studio, Antigravity, and its Enterprise Agent Platform.
The launch arrives as the AI‑speech market tightens around a few heavyweight players. OpenAI’s Whisper, Microsoft’s Azure Speech, and Amazon Transcribe each claim high accuracy, but Google is leveraging its massive data pipeline and tight integration with its productivity suite to differentiate. By bundling transcription directly into tools like Docs and Gmail, Google blurs the line between a standalone API and a built‑in feature, echoing the broader industry push to embed generative AI into everyday workflows. The rapid succession from Gemini 3.5 to 3.6 and 3.7 shows Google’s commitment to iterative improvement, while the addition of function calling hints at a future where speech input can trigger multimodal actions without developer glue code.
The real impact will be measured by adoption in enterprise and developer ecosystems. Custom vocabularies and multi‑speaker timestamps make the model attractive for meeting‑notes, call‑center analytics, and industry‑specific jargon, potentially expanding Google’s share of the transcription market. However, the service’s reliance on cloud processing raises privacy and latency concerns, especially for regulated sectors. Watch for the timing and pricing of the API rollout, the stability of speaker‑identification beyond three voices, and how competing platforms respond with their own integrated transcription features.
Key Takeaways
Gemini 3.5 Transcribe claims the lowest publicly disclosed WER for Google’s speech models, at 4 % streaming and 2.6 % pre‑recorded.
The model is being embedded across Google’s consumer apps and enterprise platforms, not just offered as a standalone API.
Function‑calling capability enables voice‑driven multimodal tasks, extending transcription beyond plain text.
Adoption will hinge on API pricing, privacy safeguards, and the model’s performance with more than three simultaneous speakers.
About the Source
This analysis is based on reporting by Android Authority. Here is a short excerpt for context:
Gemini will now offer better precision, deeper context awareness, and more.Read the original at Android Authority