Google has introduced Gemini 3.5 Transcribe, a new AI transcription model aimed at making spoken audio easier to capture, edit, and reuse. The tool can automatically detect specialized terminology and supports more than 85 languages, expanding access for global users and multilingual teams.
One standout feature is its ability to produce cleaner transcripts by editing out filler words such as “ums” and “ahs”. That means interviews, podcasts, meetings, lectures, and voice notes can become more readable and professional with less manual cleanup.
Why this matters
- Accessibility: Better transcripts help people who are deaf, hard of hearing, non-native speakers, or reviewing content in noisy environments.
- Productivity: Cleaner AI-generated text can reduce hours of manual editing for creators, journalists, educators, and businesses.
- Global reach: Support for 85+ languages helps make spoken information searchable and shareable across borders.
Google says the model represents a major improvement over its previous transcription system, particularly in multilingual performance and wording accuracy. If widely deployed across Google’s AI and productivity tools, Gemini 3.5 Transcribe could make everyday audio workflows faster, cleaner, and more inclusive.