Gemini 3.5 Transcribe发布,语音转文字模型升级

Today, we’re launching Gemini 3.5 Transcribe, our new speech-to-text model with sub-second streaming...

精选理由

GoogleAIStudio和GeminiApp已上线,秒级流式识别,比Chirp 3快70%,体验更佳。

AI 摘要

Gemini 3.5 Transcribe发布,支持85+语言,非流式识别错误率2.6%,流式识别错误率4.0%,比Chirp 3快70%。支持秒级双向流,提供词级时间戳和说话人识别。

图片来源 · Philipp Schmid
原文 · Philipp Schmid

Today, we’re launching Gemini 3.5 Transcribe, our new speech-to-text model with sub-second streaming...

Today, we’re launching Gemini 3.5 Transcribe, our new speech-to-text model with sub-second streaming and intelligent post-processing for agent interfaces. 2.6% WER on non-streaming and 4.0% on streaming. Supports 85+ languages, cleans up conversational disfluencies ("um", "ah", and mid-sentence self-corrections), handles alphanumeric tokens (postal codes or IDs), 70% reduction time for the final transcription compared to Chirp 3. - `gemini-3.5-transcribe-live`: Sub-second, bidirectional streaming via the Gemini Live API - `gemini-3.5-transcribe`: Recorded processing via the Interactions API with word-level timestamps and speaker attribution (up to 3 speakers). Having access to Transcribe has made me use voice input way more than I ever did before. The native post-processing is incredible. It knows you mean `.json` and not a person named "Jason". I sometimes speak 5 minutes into and it perfectly refactors my instruction to my context. Available today in public preview in Gemini API, @GoogleAIStudio and already in the @GeminiApp and @antigravity . Try here: ai.studio/apps/bundled/g… Your browser does not support the video tag. 🔗 View on Twitter 💬 2 🔄 3 ❤️ 29 👀 1722 📊 6 ⚡