论文

研究显示全双工语音模型互相对话时轮换时机比人类慢得多

Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue

精选理由

两个 PersonaPlex-7B 模型互相打电话,结果轮到说话比人类慢三四倍,因为它们在傻等对方说完,不会预判话轮结束,挺有意思的。

arXiv 论文让两个 PersonaPlex-7B 实例在共享时钟上交换音频 token,进行无剧本对话,并与人类电话语料库 Switchboard 对比轮换时机。模型间的话轮交接中位数为 400-560 毫秒,而人类仅为 137 毫秒。人类对话者会在对方话轮最后 120 毫秒内完成约十分之一的话轮交接,模型只有 1%。实验还显示单向通道延迟会一对一推移响应时间,说明模型是在感知对方说完后被动等待,而非像人类那样预测话轮结束。

原文 · arXiv cs.AI

Coupled but Late: Turn-Taking Between Full-Duplex Speech Models in Unscripted Dialogue

Full-duplex speech models are trained to converse with a person, but they are increasingly made to converse with each other, in self-play data generation, agent societies, and model-based evaluation. In that loop no human absorbs a timing error: each model's turn-taking is the other's input. We ask what timing the loop settles into. Two PersonaPlex-7B instances exchange audio tokens on a shared clock in unscripted conversation, and one floor-transfer rule is applied to them and to Switchboard. Their timing is coupled: re-pairing speakers across conversations destroys it. But the floor changes hands late, at a median of 400-560 ms against 137 ms for humans, and the last 120 ms of the partner's turn, where human projection places a tenth of its transfers, holds 1% of theirs. Delaying one direction of the channel shifts the response one-for-one and leaves the run-up to it empty, consistent with a reactive wait after the perceived end rather than the turn-end projection human timing requires.