Google has officially introduced Gemini 3.5 Transcribe, positioning it as the company's most precise speech-to-text model yet. The technology converts raw audio into polished, formatted text by automatically handling background noise, complex jargon, and disfluency cleanup. It is already powering several first-party products, including the new Gboard Rambler feature on Android, the Gemini macOS app, and is slated for release on the Chrome browser.
Google announced Rambler during its Android Show: I/O Edition 2026 event on a Tuesday morning. The launch puts Google in direct competition with AI dictation startups like Wispr Flow and Typeless, which have yet to establish a strong foothold on Android.
Rethinking Voice Dictation
Unlike traditional speech recognition models that struggle with pauses and mistakes, Gemini 3.5 Transcribe is designed to understand natural speaking styles and intent. The model powers Rambler, a Gboard feature that actively removes filler words like "ums" and "ahs" while accommodating mid-sentence self-corrections.
Because Rambler is integrated directly into Gboard, it works anywhere throughout Android. Users can access the tool to convert speech to text in a Google Messages conversation, a Notion document, or a Slack thread.
The model relies on multilingual capabilities that automatically detect and transcribe over 85 languages, handling regional accents and diverse dialects. It also natively handles code-switching, allowing bilingual users to transition seamlessly between languages—such as English and Hindi—without breaking the transcription context. To ensure privacy, Google noted that Gboard clearly indicates when Rambler is active and uses audio exclusively for transcription without storing voice recordings.
Performance and Latency Upgrades
Gemini 3.5 Transcribe delivers a significant latency and accuracy leap over Google's 2025 Chirp 3 model. According to Artificial Analysis, time to final transcription improved by 70%.
The system sets new benchmarks for precision:
- Achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases.
- Reaches a 5.50% WER in streaming mode and 5.04% in non-streaming mode on the FLEURS benchmark.
- Features multi-speaker identification with timestamps for up to three distinct speakers, with support for more currently experimental.
- Adapts to unique spellings and specialized industry jargon via provided custom vocabulary.
The underlying Gemini technology builds upon a foundation that previously saw models like Gemini Ultra achieve a 59.4% score on the MMMU multimodal benchmark. Google has stated that its models have historically outperformed GPT-4 in text-based reasoning, math, and code benchmarks.
Ecosystem Expansion and Availability
Beyond basic text entry, function calling enables Gemini 3.5 Transcribe to execute complex commands. This allows users to delegate tasks like file analysis and image generation to other Google Gemini models. Google confirmed the model will soon integrate directly into the Chrome browser, enabling users to "talk to type in any web field".
Additionally, Google is bringing its Gemini Nano model to Chrome 126 on desktop to speed up AI-powered features like the "Help me write" tool. The company is also introducing the Speculation Rules API to dramatically speed up browsing by pre-fetching pages, and the View Transitions API for seamless page switching in Chrome Canary 126. Furthermore, Gemini is arriving in Chrome DevTools Console insights to provide debugging solutions for errors.
Currently, Gemini 3.5 Transcribe is available in public preview for developers via the Gemini API in Google AI Studio and Google Antigravity. Enterprise users can access it through the Gemini Enterprise Agent Platform, with support for Gemini Enterprise for Customer Experience arriving soon.