论文

论文提出情境条件化思考策略,让长期 LLM 智能体复用推理经验

Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents

精选理由

arXiv 上的新论文,教长期智能体把过往推理经验压成一个轻量策略,实验里把 DeepSeek 的推理 F1 从 0.789 拉到 0.868,做智能体记忆的可以看看。

arXiv 论文提出情境条件化思考记忆框架,把历史推理经验转成轻量策略,用于预测当前情境下该思考什么,细节推理仍交给 LLM。临时经验会跨多个独立片段周期性分析,提炼出重复的长程规律并内化为新的思考知识。实验中策略在时间规则泛化上达到 1.000 F1,将 DeepSeek 推理 F1 从 0.789 提升到 0.868,在 30,000 条历史情境下单查询在线处理时间从 0.3636 ms 降至 0.0382 ms。跨经验证据充分积累后,关系发现 F1 和未来思考准确率达到 1.000。

原文 · arXiv: DeepSeek

Learning Situation-Conditioned Thinking Policies for Long-Term LLM Agents

Long-running autonomous agents must reuse accumulated reasoning experience without allowing explicit historical memory and LLM context to grow indefinitely. However, existing memory mechanisms mainly retrieve, summarize, or compress past content and do not directly learn when particular kinds of thinking should be activated or discover new thinking knowledge from temporally dispersed experiences. This paper proposes a situation-conditioned thinking memory framework that transforms historical reasoning experience into a lightweight policy for predicting what should be thought about in the current situation, while leaving detailed reasoning to a large language model. Situations may represent temporal or spatiotemporal evolution rather than only current states. Temporary experiences are also periodically analyzed across multiple independent episodes to identify repeated long-range regularities, which are consolidated into new thinking knowledge and further internalized by the lightweight policy. Experiments show that the learned policy achieves 1.000 F1 on temporal-rule generalization, improves DeepSeek reasoning F1 from 0.789 to 0.868, reduces online processing time from 0.3636 ms to 0.0382 ms per query at 30,000 historical situations, and reaches 1.000 relation-discovery F1 and future-thinking accuracy after sufficient repeated cross-experience evidence.