Alphabet's Google introduced Gemini 3.5 Transcribe, its most advanced speech-to-text AI model.
This model is part of Gemini Audio, a trio of new AI audio models.
Gemini 3.5 Transcribe provides highly accurate transcriptions from raw audio, even in noisy environments. It detects over 85 languages. It identifies up to three different speakers in pre-recorded audio.
The model creates polished text. It removes filler words such as "um" and "ah." It also handles user self-corrections.
This technology already powers the "Rambler" feature on Android devices. Google integrates it into products like Chrome, Docs, and Gmail.
Developers access the model through the Gemini API. It supports both real-time streaming and processing recorded files. The model offers significant improvements in speed and accuracy over its predecessor.