搞 agent 记忆的人该看看:ContextWeave 拿真实办公流做了1005个任务,能看出召回记忆到底帮不帮得上忙。
ContextWeave 是一项评估语言智能体长期记忆的新基准,基于14名参与者的多月经脱敏工作流,构建了1005个可执行任务,其中568个为核心评测任务。固定模型下,最强记忆配置将 Workspace Score 从68.08提升至78.20,Preference Score 从41.50提升至70.60。固定记忆组件时,召回记忆改善了全部5个测试基础模型的结果。分析还发现,可操作且经验丰富的记忆比紧凑摘要更能支持工作流延续,但也更容易受误导性召回影响。
ContextWeave: A Real-World Workflow Benchmark
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams. ContextWeave reconstructs privacy-preserved, multi-month workflows of 14 participants into 1,005 executable tasks, including 568 core evaluation tasks, with instructions, containerized environments, trajectories, and task-specific rubrics. It measures workspace quality and alignment with participant-specific preferences, complemented by diagnostics of relevance, continuity, solvability, and robustness to misleading recall. Across six memory components under a fixed model, the strongest configuration raises Workspace Score from 68.08 to 78.20 and Preference Score from 41.50 to 70.60. With a fixed memory component, recall improves both outcomes for all five tested base models, although gains vary substantially. Our analysis shows that actionable, experience-rich memory supports workflow continuation and reduces redundant exploration more effectively than compact summaries, while it can also be more susceptible to misleading recall. These findings motivate memory systems that optimize not only retrieval relevance but also reliable use during execution.