Google's latest voice model does more than simply convert speech into text. It scrubs the audio of stumbles, false starts and filler words before outputting a clean transcript. The company announced Gemini 3.5 Transcribe on Tuesday, positioning it as a faster and more accurate successor to the Chirp 3 engine that previously handled speech-to-text tasks.

What You Need to Know

Gemini 3.5 Transcribe is an AI model designed to process natural, messy speech and output polished text. Google says it runs 70% faster than its predecessor and lowers the live-speech error rate to 5.5%. The feature already powers the Gboard Rambler on the Pixel 11 and will expand across other Google products in the coming months.

Faster and Cleaner Transcription

The core improvement over the Chirp 3 engine lies in both speed and accuracy. Google claims a 70% reduction in the time between a user speaking and seeing finalized text. The error rate for live speech dropped about 1.8 percentage points, from 7.32% with Chirp 3 to 5.5% with the new model.

  • Speed gain: Google measures a 70% improvement in processing from voice to final text.
  • Error reduction: The live-speech error rate fell to 5.5 percent, down from 7.32 percent with Chirp 3.
  • Polishing ability: The model automatically removes filler words like 'um' and corrects self-interruptions.

Where It Appears

Gemini 3.5 Transcribe is already active on the Pixel 11 through the Gboard Rambler feature. Rambler lets users dictate text and receive a version stripped of hesitations and corrections. Google plans to bring the model to other surfaces in its ecosystem, including Google Docs, the Google Assistant and third-party apps that integrate with Google's speech APIs.

Why This Matters

The arrival of Gemini 3.5 Transcribe signals a shift in how voice interfaces handle natural speech. Earlier systems simply converted words verbatim, forcing users to manually edit out mistakes. By removing disfluencies automatically, the new model makes voice dictation feel more like polished typing. For professionals who dictate long emails or documents, this could save significant editing time. The technology also lowers the barrier for hands-free communication, benefiting users with mobility impairments or those multitasking. As artificial intelligence models become better at understanding human speech patterns, the gap between speaking and writing continues to narrow. Google's introduction of this model places pressure on competitors like Apple and Amazon to improve their own dictation engines.