RUMBA:俄语用户记忆基准,细粒度评估长对话记忆

RUMBA: Russian User Memory Benchmark

精选理由

想测LLM的长对话记忆力?RUMBA这个俄语基准能细粒度诊断模型在时间推理上的强项和短板。

AI 摘要

RUMBA是一个面向长对话记忆的基准,包含时间戳用户-助手对话和QA对(检索、组合、跨会话推理)。基准提供细粒度分类,覆盖语义类型、会话范围、时间推理和表达显式性。评估了当代记忆系统和长上下文模型(如GPT-4、Claude),识别出不同机制在时间线索跨会话聚合上的失败模式。同时提供英语对齐子集用于跨语言比较。

原文 · arXiv cs.AI

RUMBA: Russian User Memory Benchmark

The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate retrieval metrics, failing to capture interactions between long-range context, temporal information, and reasoning. To address this, we introduce RUMBA (Russian User Memory BenchmArk) - a new benchmark for long-term conversational memory that provides a fine-grained taxonomy of memory-centric question types and a unified methodology accounting for semantic type, session scope, temporal reasoning, and the explicitness of temporal expressions. RUMBA consists of timestamped user-assistant dialogues with QA pairs requiring retrieval, combination, and reasoning across sessions. While designed for Russian, we also provide an aligned English subset under the same methodology. We evaluate contemporary memory systems and long-context models, and show how RUMBA serves as a diagnostic tool to analyze model behavior across benchmark slices and identify strengths and failure modes of different memory mechanisms.