论文

MemPilot:用强化学习按需编排多模态记忆管理

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

精选理由

一篇讲 agent 记忆管理的论文,用 RL 学出策略在“查现成记忆”和“现整理原始历史”之间动态切换,还能按成本和延迟偏好调,实验跑在五个多模态基准上。

论文提出 MemPilot 框架,针对现有 agent 记忆系统与查询无关、预处理成本高的问题,让策略在检索现成记忆和按查询即时整理原始多模态历史之间做选择。该策略通过强化学习训练,可联合控制证据数量、整理指令、模型选择和视觉访问,并允许按性能、成本、延迟的不同偏好做优化。训练上采用按目标分离估计优势的 objective-wise advantage decoupling,以及基于前缀的边际效用估计来做多步 credit assignment。在五个多模态 agent 记忆基准上的实验显示,其性能-成本-延迟权衡前沿优于现有 trade-off-aware 基线。

原文 · arXiv cs.AI

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored. To address this challenge, we present \textbf{MemPilot}, a flexible framework that orchestrates on-demand memory curation under different performance--cost--latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective's advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance--cost--latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.