Google 发布 Gemini 3.8 Live 语音模型,端到端响应省去转写中转
Google 把语音助手那套转写接力砍掉了,Gemini 3.8 Live 一个模型直接听直接答,还便宜到每小时 0.84 美元。
Google 发布 Gemini 3.8 Live 系列,用单一系统直接完成听、推理和语音回复,跳过传统语音转文本再转模型的接力流程。其中 Extended Thinking 版本在 Artificial Analysis 的 Speech to Speech Index 排名第一,标准版在人工盲测对话中排名第二。标准版输入音频定价为每小时 0.84 美元,是该指数中最低价格。两个版本均支持图像和视频输入,可根据屏幕内容直接回答问题。
Most voice assistants work like a relay race: speech becomes text, text goes to a model, the answer becomes speech again. Every handoff adds a pause. ⏱️ Speech-to-speech models skip the relay. Google's new Gemini 3.8 Live models listen, reason, and respond in one system. 🎯 The Extended Thinking version ranks first on Artificial Analysis' Speech to Speech Index 🗣️ The standard version ranks second in blind live conversations judged by people 💰 The standard version costs $0.84 per hour of input audio, the lowest in the index Both models also take in image and video input. Picture asking an assistant for help with whatever is on your screen, and it simply answers. 📱 Read the full story in The Batc hubs.la/Q04ztmlL0 zb #DeepLearningAI i #VoiceAgents g #AI s #AI 💬 3 🔄 0 ❤️ 8 👀 1134 📊 3 ⚡