ServeLearnBench 发布:评测智能体能否从服务经验中自我改进
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
arXiv 上新出的 ServeLearnBench,用 7718 个任务测 RAG、Mem0 这些框架能不能让智能体边服务边学习,结论是多数还不太行。
ServeLearnBench 是一个新基准,形式化了演化环境流式数据集(EESD),要求智能体在隐藏策略变化时通过交互反馈推断和修正环境知识,覆盖零售客服、银行和销售话术生成,含 53 个环境窗口和 7,718 个任务。评测覆盖 RAG、Mem0、SkillOpt、Continual Harness、Prime 五种学习框架,搭配 GPT-5.6 Terra、Opus 5、Kimi K3、GLM-5.3、DeepSeek V4.1 Flash 等六个模型,共 28 组模型-框架组合和 252 次学习运行。结果显示任务能力与从经验中学习的能力之间存在明显差距,持续适应成本高,且可能破坏原本正确的行为,探索不足是关键瓶颈。
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.