Tech

Google rolls out Gemini 3.5 Audio models with advanced transcription capabilities

The new suite supports over 85 languages, automatically detects specialised jargon, and removes filler words, marking a significant upgrade from the previous Chirp 3 model.

Editorial persona
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: The Verge · View original source
Google’s new AI transcription edits out your ‘ums’ and ‘ahs’
Technology

Google has expanded its Gemini Audio suite with the release of new Gemini 3.5 models, introducing a transcription tool that supports more than 85 languages and automatically detects specialised jargon. The update arrives as users continue to wait for the Gemini 3.5 Pro model, which Google previously promised to roll out in June.

The new Gemini 3.5 Transcribe model is a distinct addition to the Gemini family, replacing the previous Chirp 3 model. Google describes the update as a major advancement in multilingual performance and wording error rates. The tool allows users to edit text naturally using their voice, automatically formatting output and removing filler words such as “um” and “uh”.

To improve accuracy, the model can accept customised vocabulary inputs, enabling it to adapt to unique spelling requirements and industry-specific terminology without manual correction. It is also capable of attributing speech for up to three speakers in pre-recorded audio, providing word-level timestamps for greater precision.

Alongside the transcription tool, Google has launched Gemini 3.5 Live and 3.5 Live Experimental. These models build on existing speech recognition technology to enhance voice-chat performance. Gemini 3.5 Live is designed to handle mid-sentence interruptions, language recognition, and live visual processing more effectively.

Gemini 3.5 Live Experimental extends these capabilities by narrating its reasoning progress step by step in real time while tackling complex tasks. The models are intended to provide better precision for Google’s voice-controlled AI features, reducing issues related to background noise or interrupted speech.

The updates are rolling out immediately to macOS Gemini app users in English and to the Rambler dictation feature on Android in select countries and languages. Developer access is available in public preview via the Gemini API, AI Studio, and Antigravity. Google has indicated that support for the Chrome browser is expected to arrive in the near future.

Continue reading

More from Tech

Read next: GitHub project documents reproducible CUDA compatibility setup for AMD GPUs on Windows
Read next: Automakers pull back from CarPlay as control of vehicle software takes priority
Read next: Budget HDMI extenders offer longer reach, with trade-offs