Google has introduced Gemini 3.5 Transcribe, a speech-to-text model designed to turn raw audio into cleaner, context-aware text for voice agents, captions, meetings, and call-analysis tools.
Unlike conventional transcription systems that largely reproduce what was said, Google says Gemini 3.5 Transcribe can handle natural speech patterns such as filler words and self-corrections. For example, a sentence like “let’s meet Tuesday, no, Wednesday” can be interpreted as Wednesday in the polished output.
The model is available to developers in public preview through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Google has not announced Philippine-specific pricing or availability beyond those developer platforms.
Google is offering Gemini 3.5 Transcribe through two separate API paths:
gemini-3.5-transcribe-livethrough the Live API for continuous, bidirectional streaming with sub-second latency.gemini-3.5-transcribethrough the Interactions API for recorded audio such as meetings, interviews, and call logs.
The recorded-audio model supports speaker attribution and word-level timestamps. In practical terms, this means a transcript can identify who said each segment and where that segment occurred in the recording, which is useful for meeting notes and post-call analysis.
Gemini 3.5 Transcribe also supports custom vocabulary, allowing developers to provide specialized terms, names, product codes, or industry jargon that may otherwise be transcribed incorrectly. Google says the model can automatically detect more than 85 languages.
Google is positioning the model for intelligent voice interactions rather than captions alone. Its smart-transcription features can remove filler words, clean up disfluencies, and apply formatting to the output. The company also describes function calling, where the transcription workflow can pass a task to another Gemini model, such as image generation or file analysis.
That approach can make voice interfaces easier to use, but it also creates an important distinction between a verbatim record and an edited interpretation of what someone said. Developers building legal, medical, journalism, or compliance tools may need to retain the original audio and a literal transcript alongside any cleaned-up version.
Google reports an average word error rate of 4.0% for streaming use cases and 2.6% for non-streaming use cases, based on its cited Artificial Analysis measurements. These figures are vendor-reported and can vary depending on language, audio quality, background noise, speakers, and vocabulary.
The model is already being used in some Google products, including Rambler on Android and voice features in the Gemini app on macOS. Google says Gemini 3.5 Transcribe is also coming to Chrome, while broader product availability and local pricing details remain unannounced.
Source: Google DeepMind · Google AI pricing
Availability
Gemini 3.5 Transcribe is available in public preview for developers through Google AI Studio and the Gemini Enterprise Agent Platform.
Community
0 reader comments
No comments yet. Join the conversation with a useful, respectful response.