Google has expanded its Gemini family with a new tool focused on audio transcription, aimed at making it easier for users to capture and refine spoken content. Gemini 3.5 Transcribe is described by the company as a “major advancement” over its previous model, Chirp 3, particularly in multilingual performance and reduced word error rates.
According to a blog post from Google, the model allows users to “edit naturally with just your voice,” hinting at interactive capabilities where speakers can correct or polish transcripts in real time. The system is designed to pick up on technical vocabulary across various fields, a feature that could prove valuable for professionals in medicine, law, or engineering who rely on accurate, jargon-aware transcripts.
The release of 3.5 Transcribe comes weeks after Google introduced Gemini 3.5 Live Translate, a feature that uses the phone’s speaker and microphone to provide live translations when held to the ear. Both updates are part of Google’s broader push to integrate its Gemini AI into everyday productivity tools.
However, the launch also underscores a notable gap in Google’s roadmap. The company announced Gemini 3.5 Pro — a more powerful, multi-modal model — in early 2025, promising a rollout in June. That release has yet to materialize, raising questions about development timelines and resource allocation within Google’s AI division.
Industry observers note that transcription AI has become increasingly competitive, with offerings from OpenAI (Whisper), Apple, and a host of startups. While Google’s model claims superior language coverage and jargon handling, independent benchmarks will be needed to verify those assertions. Privacy concerns also remain: transcription services often process sensitive audio data, and users will be watching how Google handles data security and on-device processing.