论文多源确认

Acoustic-to-Text KV压缩:让全双工语音模型长对话内存降64.6%

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

精选理由

全双工语音模型聊久了内存会爆,这篇用转录旁路压缩KV,十分钟会话峰值缓存直接砍掉64.6%,性能还不掉。

论文提出 acoustic-to-text KV compression,在全双工语音语言模型的 listening-time slack 区间内,通过转录旁路把语音转成紧凑的文本记忆。推理时 KV 缓存超过预算即淘汰旧声学状态,保留转录与近期声学上下文。旁路用 LoRA 训练,并以知识蒸馏保持原始模型的听说行为。在 MiniCPM-o 4.5 上的实现使十分钟 LongSpeech 会话的峰值流式 KV 缓存降低 64.6%,转录、时序问答和摘要指标同时优于基线,Full-Duplex-Bench 上的停顿、轮次切换和打断表现持平。

原文 · arXiv cs.AI

Acoustic-to-Text KV Compression for Full-Duplex Speech Models

Full-duplex speech language models continuously accumulate acoustic key-value (KV) states, making long-running interactions memory-intensive. During listening, the model can finish processing an audio unit before the next arrives; we term the remaining interval listening-time slack. We propose acoustic-to-text KV compression, which introduces a transcription side channel to convert incoming speech into compact textual memory within this interval. When the cache exceeds a target budget during inference, older acoustic states are evicted while transcripts and recent acoustic context remain. We train the side channel with LoRA using cross-entropy on transcription segments. To preserve listening and speaking behavior, we apply knowledge distillation to the original model's token-level output distributions at native prediction positions. On ten-minute LongSpeech sessions, our MiniCPM-o 4.5 implementation reduces peak streaming KV-cache size by 64.6% compared with the same model without eviction. The proposed method also improves transcription, temporal question answering, and summarization over the baseline. Full-Duplex-Bench evaluations further show comparable pause-handling, turn-taking, and interruption performance.