Setoka:评估个性化智能体层次化用户理解的新基准

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

精选理由

如果你想了解现有AI智能体到底能多深地理解用户,这个新基准Setoka用心理学理论和方法测出了它们的短板——不只是翻聊天记录,还要推断行为和性格。

AI 摘要

Setoka基准基于认知与人格心理学理论,定义了用户理解的四个层次:语义记忆、情景记忆、行为模式和人格特质。该基准通过心理测量学流程合成异构用户数据与查询,用于评估记忆增强型个性化智能体。研究评估了3种语言模型结合5种记忆系统在10个合成用户上的表现,发现现有系统在语义记忆检索上表现良好,但在情景记忆上性能下降;在处理行为模式和人格特质理解任务时,由于需要整合碎片化信息,性能进一步下滑。

原文 · arXiv cs.AI

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.