提出首个多边对话参与者评估基准 MP-Bench
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
研究团队提出了首个专门评估语音代理在多边对话中表现的基准 MP-Bench,通过12个语音代理的测试,揭示了当前实时语音代理在多边对话中的不足。
研究团队为评估语音代理在多边对话中的表现,提出了首个专门设计的基准 MP-Bench。该基准评估了12个语音代理在多边对话中的行为,发现实时语音代理在多边理解方面得分低于22%,在多边轮流发言方面接近随机水平,暴露了该场景下的开放挑战。
MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.