FlowState:把执行状态当作记忆,让 LLM 智能体更好处理长任务
FlowState: Execution State as Memory for Long-Horizon LLM Agents
一篇讲怎么给智能体做记忆的论文,思路是把执行状态当记忆存起来按需取,token 省了四成,成绩还涨了。
论文提出 FlowState,把智能体的执行状态作为可保留、可回访的记忆,用语义类型化的状态节点及其关系替代完整对话历史。其内部包含增量状态更新(ISU)和渐进式状态访问(PSA)两个机制,让智能体在推理过程中按需回看历史状态与证据。基于 DeepSeek-V4-Flash 的对比测试中,FlowState 在 MemoryArena 上平均成功率提高 4.55 个百分点,在 τ³-Bench 上平均通过率提高 13.95 个百分点,token 消耗分别减少 43.2% 和 40.6%。
FlowState: Execution State as Memory for Long-Horizon LLM Agents
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.