Google发布Gemini 3.5 Transcribe语音转文本模型

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

精选理由

Google新发布的Gemini 3.5 Transcribe支持85种语言,错误率低至2.6%,还提供两种端点选择,适合语音助手和转录管道使用。

AI 摘要

Google推出Gemini 3.5 Transcribe语音转文本模型,分为流式和非流式两个端点。流式端点实现亚秒级转录,但省去说话人分离和词级时间戳功能,错误率为4.0%。非流式端点保留所有功能,错误率为2.6%,成本仅为流式的一半。该模型在85种以上语言中表现优异,最终处理速度比Chirp 3快70%。

图片来源 · marktechpost
原文 · marktechpost

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages

Google has released Gemini 3.5 Transcribe, a speech-to-text model that ships as two separate endpoints rather than one. The streaming endpoint delivers sub-second transcription but drops speaker diarization and word timestamps. The batch endpoint keeps both, at half the cost. Google reports 4.0% word error rate streaming and 2.6% non-streaming, with 70% faster finalization than Chirp 3. Here is what the split means for anyone building voice agents or transcription pipelines. The post Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages appeared first on MarkTechPost .

Google发布Gemini 3.5 Transcribe语音转文本模型 · AI 热点