Logo
FrontierNews.ai

Google's New Gemini 3.5 Transcribe Removes 'Ums' and 'Ahs' From Your Voice Notes

Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically removes filler words like "um" and "uh" while formatting your voice notes as you dictate. The model rolled out today to the Gemini app on macOS and to Rambler, Google's dictation feature in Gboard on Android, with Chrome support coming later.

What Makes Gemini 3.5 Transcribe Different From Standard Voice Typing?

Unlike standard voice typing that transcribes word-for-word, including mistakes and filler words, Gemini 3.5 Transcribe cleans up as it goes. The model replaces Chirp 3, Google's previous transcription tool, and delivers significantly better accuracy. According to the official announcement, it achieves a word error rate of 4.0 percent for streaming audio and 2.6 percent for pre-recorded audio, meaning it captures what you say with near-perfect precision.

The model also cuts the time to final transcription by 70 percent, making it substantially faster than previous versions. Beyond removing filler words, Gemini 3.5 Transcribe handles self-corrections automatically, so if you say "let's meet Tuesday, no, Wednesday," the model understands your intent and outputs only the corrected version.

How to Use Gemini 3.5 Transcribe Across Different Devices?

  • macOS Gemini App: Available today in English, allowing journalists, creators, and professionals to dictate clean, formatted text directly into the Gemini application without opening a browser.
  • Android Rambler Feature: Powers the Gboard dictation tool on Android, currently requiring a Pixel 11 device and rolling out in select countries and languages, with Arabic support included.
  • Chrome Browser: Coming soon to any web text field, though Google has not yet announced a specific launch date for this expansion.
  • Developer Access: Developers can access the model through public preview via the Gemini API in AI Studio, Antigravity, and the Gemini Enterprise Agent Platform.

The model recognizes more than 85 languages, including regional accents and dialects, making it useful for global audiences. For pre-recorded audio, it can attribute speech to up to three speakers with word-level timestamps, a feature particularly valuable for transcribing interviews, meetings, or podcasts.

What Advanced Features Does the Model Include?

Gemini 3.5 Transcribe accepts custom vocabularies, allowing specialized jargon and unusual spellings to survive the transcription process intact. This is especially useful for technical professionals, medical practitioners, and industry experts who use terminology that standard models might misinterpret or correct.

The model also handles formatting automatically, converting raw dictation into properly punctuated and grammatically structured text. This means you can speak naturally, including pauses and corrections, and receive polished output ready to share or publish without manual editing.

When Will It Be Available in Your Region?

Availability varies by device and location. The Gemini app on macOS works in English immediately, making it accessible to most English-speaking users worldwide. On Android, the bottleneck is hardware; Rambler currently requires a Pixel 11 device, and Google has not confirmed a rollout timeline for other regions or older Android phones.

For users outside the initial rollout countries, Google's offline dictation app remains available on iPhone with zero internet connection required, though it operates as a separate tool from Gemini 3.5 Transcribe. Chrome support, already in development, will eventually bring the model to any web field, but no launch date has been announced.

The release of Gemini 3.5 Transcribe comes as Google's flagship Gemini 3.5 Pro model, originally promised for June, still awaits a release date. The new transcription model launches alongside Gemini 3.5 Live and 3.5 Live Experimental, which handle voice chat and step-by-step reasoning respectively, expanding Google's voice-first AI capabilities.