AGM让冻结的VLA策略变成闭环系统,物理验证子目标,机器人任务表现提升明显。
AGM是一种轻量级闭环框架,用于冻结的视觉-语言-行动(VLA)策略。该框架将任务表示为带有进度指针的子目标序列,仅在当前子目标通过物理证据验证后才推进记忆。在RoboMME Counting基准测试中,AGM在PickXTimes和BinFill任务上表现优异,平均得分超越最强记忆增强基线。AGM通过单参数验证头实现,无需测试时大模型推理。
AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies
Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so the agent cannot reliably decide whether to continue, retry, or terminate. External memory is a natural remedy, yet it can be harmful when attempted actions are treated as completed progress, turning local execution errors into persistent task-state errors. We propose Achievement-Grounded Memory (AGM), a lightweight closed-loop framework for frozen VLA policies that represents a task as a subgoal sequence with a progress pointer and advances this memory only after the current subgoal is verified by physical evidence. Proprioceptive interaction cues decide when to verify, while coherent point tracking and language-conditioned cross-view comparison, sourced from frozen foundation models through a single 2.43M-parameter verification head, decide what was achieved. AGM thereby converts open-loop execution into a closed loop of execution, verification, and progress, keeping the policy frozen without test-time large-model inference. On the RoboMME Counting benchmark, AGM reaches on PickXTimes and on BinFill, surpassing the strongest memory-augmented baseline by points on average, and the framework yields equally decisive gains on a physical robot. Reliable embodied memory thus depends more on disciplined state updates than on memory capacity.