VoiceMem:实时交互的流式双脑记忆系统

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

精选理由

VoiceMem提供了实时、个性化、情感感知的语音交互的实用记忆基础,在准确性、情感与个性化方面优于Mem0等经典系统,且检索速度快,成本较低,值得一看。

AI 摘要

VoiceMem是一种简单的记忆架构,包含并行信息左脑、情感右脑和流式内存I/O机制。实验和实际部署显示,VoiceMem在准确性、情感与个性化以及实时与低成本方面具有优势。左脑在top-5检索中比Mem0等经典系统高出近30分;右脑在三个角色基准测试中达到最先进的性能,比之前最佳系统提高了4.29分;VoiceMem在134毫秒内完成检索,在标准VAD延迟内,不增加额外的对话延迟,同时保持高准确性和低成本。

原文 · arXiv cs.AI

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.