TRACE:解决多智能体系统中过期记忆准入问题
TRACE: Governing Memory Validity in Evolving Multi-Agent Systems
多智能体协作里队友改了共享状态,回来的智能体怎么办?这篇论文给了个免训练方案,两个指标同时拉到98%以上,做法挺巧。
论文提出 temporal memory admission 问题:智能体离席期间共享状态被其他四个队友修改,返回后哪些记忆仍可用于行动。TRACE 是免训练的准入层,将重新入场视为资格判定,对齐离席检查点与缺席期更新,仅释放覆盖未决义务的有限 Return View。在 ManBench-Return 上,TRACE 达到 92.6-98.3% 有效信息可用率和 98.4-99.5% 无效信息拒收率,而 Restore 和 Reset 两种基线策略各自在一边归零。在 STALE Type II 上,TRACE 的 Overall 比最强对比策略高出 22.3(Qwen)、18.5(Gemini)、27.5(DeepSeek)分。
TRACE: Governing Memory Validity in Evolving Multi-Agent Systems
Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, and faithful to its source, and nonetheless be inadmissible for action: an itinerary saved before a pause still names the hotel the team has since replaced. We formalize this as temporal memory admission and present TRACE, a training-free layer that treats re-entry as an eligibility decision rather than a storage or retrieval operation, reconciling a departure checkpoint against absence-period updates, resolving explicit and implicit invalidation, and releasing a bounded Return View only when it covers the returning role's open obligations. We evaluate TRACE under three actor models on Memora, STALE Type II, and a derived ManBench-Return setting, each recast as return episodes: one agent departs, four teammates change the shared state, and the agent rejoins. What separates methods is not overall accuracy but whether one can retain valid memory and reject stale memory at once, and no single-policy baseline can: Restore (reinstate the departure checkpoint in full) admits stale state, Reset (start the return from an empty memory) discards valid state, each bottoming out at 0% on one of the two. TRACE is the only method high on both, reaching 92.6-98.3% valid-information availability with 98.4-99.5% invalid-information rejection on ManBench-Return, within 3.8 points of the best baseline's overall accuracy. On STALE Type II it improves Overall over the strongest comparison policy by 22.3 (Qwen), 18.5 (Gemini), and 27.5 (DeepSeek) points at roughly 2.3 times their tokens, while a write-time consolidation pipeline is more accurate still at 3.99 times TRACE's.