全双工实时交互解决了 AI 对话中“轮流说话”的延迟痛点,做语音助手或实时交互系统的开发者可以直接看演示和设计思路。
Thinking Machines 展示了其模型 MiniCPM-o 4.5 的全双工交互能力,能同时处理音频、视觉和文本流数据。模型将连续数据流分割为固定长度片段,并按时间戳精确对齐融合,实现实时看、听、说。该设计模仿人类同时对话、观察和思考的方式,交互体验接近真人。早期结果和演示视频已公开,展示了 AI 与人类实时协作的新范式。
This feels really close to ‘real human’ interaction. Full-duplex with a model which is seeing, hear...
This feels really close to ‘real human’ interaction. Full-duplex with a model which is seeing, hearing, and speaking, at the same time is REALLY cool. In short, the model handles continuous streaming data from different sources (audio, visual, and textual content) by dividing it into small, fixed-length segments. And these segments are then perfectly aligned and merged based on the exact moment in time they occurred. This design logic is highly consistent with MiniCPM-o 4.5 🧐 Thinking Machines @thinkymachines People talk, listen, watch, think, and collaborate at the same time, in real time. We've designed an AI that works with people the same way. We share our approach, early results, and a quick look at our model in action. thinkingmachines.ai/blog/interacti… Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 3 🔄 3 ❤️ 12 👀 3613 📊 4 ⚡