Reported by 1 source

The short version

  • Gemini 3.5 Transcribe launches with support for over 85 languages and automatic removal of filler words like 'um' and 'uh'.
  • The tool allows users to upload custom vocabularies to ensure accurate spelling of specialized jargon and unique terms.
  • Google has delayed the release of Gemini 3.5 Live and Experimental models, citing no new timeline for their availability.

Google has introduced a significant update to its artificial intelligence audio capabilities with the launch of Gemini 3.5 Transcribe. This new model is designed to handle complex transcription tasks with greater accuracy than its predecessor, Chirp 3. The primary focus of this release is on improving multilingual performance and reducing errors in word recognition. Users can now rely on the system to detect specialized terminology and process audio in more than eighty-five different languages. This expansion marks a shift toward broader accessibility for non-English speakers and professionals working with technical or niche subject matter.

A key feature of the new transcription tool is its ability to clean up spoken language automatically. The model identifies and removes common filler words such as 'um' and 'uh' from the final text output. This functionality aims to produce cleaner, more readable transcripts without requiring manual editing. Additionally, the system supports natural voice-based editing, allowing users to make corrections through speech rather than typing. This approach streamlines the workflow for individuals who prefer dictation over traditional text input methods.

News Journal

To address the challenges of transcribing specialized content, Google has implemented a customizable vocabulary feature. Users can provide specific lists of terms, names, or jargon that are unique to their field or organization. The model uses this information to adapt its spelling and recognition patterns, preventing common errors where technical terms might be misinterpreted as standard words. This capability is particularly useful for industries with dense terminology, such as medicine, law, or engineering, where precision in documentation is critical.

The tool also includes features for managing multi-speaker audio files. It can attribute speech to up to three different speakers within a pre-recorded file, helping to distinguish between participants in interviews or meetings. Furthermore, the system provides word-level timestamps, which allow users to navigate directly to specific parts of the conversation. These tools are intended to make it easier to review and reference long audio recordings without having to listen through entire files.

Availability for Gemini 3.5 Transcribe is currently limited but expanding. The feature is rolling out today for all users of the Gemini app on macOS. Android users in select countries and languages can access the Rambler dictation feature, which utilizes the same underlying technology. Developers are also able to test the model through a public preview in the Gemini API via AI Studio and Antigravity platforms. Google has indicated that support for the Chrome browser is planned for the near future, though no specific date has been provided.

Prior to publication, reports suggested that Google would also launch two additional models: Gemini 3.5 Live and Gemini 3.5 Live Experimental. These were expected to enhance real-time voice chat capabilities, with improvements in handling mid-sentence interruptions and live visual processing. The experimental version was described as having the ability to narrate its reasoning process step-by-step while tackling complex tasks. However, Google has since clarified that these live audio models are not being released at this time.

The company did not provide a new launch date for the delayed live features. This reversal highlights the uncertainty surrounding the rollout schedule for Gemini 3.5’s broader ecosystem. While the transcription tool is available now, users waiting for real-time voice interaction improvements will have to wait further. The delay affects both consumer-facing applications and developer integrations that were anticipating these updates.

This release comes as Google continues to build out its Gemini 3.5 family of models. The company had previously promised the rollout of the Gemini 3.5 Pro model in June, which has not yet occurred. The introduction of 3.5 Transcribe follows the earlier launch of 3.5 Live Translate, indicating a staggered approach to releasing new AI capabilities. Each update targets specific use cases, from translation to transcription, rather than delivering all features simultaneously.

The focus on transcription accuracy and multilingual support reflects broader trends in AI development. As voice interfaces become more common, the ability to process diverse languages and dialects accurately is essential for global adoption. By allowing users to customize vocabularies and clean up filler words, Google aims to reduce the friction between spoken input and written output. This could make AI assistants more practical for professional environments where clear documentation is required.

Looking ahead, the availability of these tools will likely expand as Google integrates them into more platforms. The current limitation to macOS and select Android regions suggests a phased rollout strategy. Developers using the API preview can begin testing the model’s performance with their own data sets. As Chrome support arrives, the tool may become accessible to a wider audience without requiring dedicated applications. The delay in live voice features remains an open question for users expecting real-time interaction improvements.

Sources behind this briefing

Go to the original reporting

  • The Verge↗Google’s new AI transcription edits out your ‘ums’ and ‘ahs’