Meta AI 发布 MIRA:拆分架构解决长程研究智能体的决策难题
Meta AI 出的论文,把研究智能体拆成元推理器加执行器,专门解决跑长任务时不知道下一步研究啥的问题,做 agent 的可以看看。
Meta AI 的 MIRA 论文针对长程研究智能体中"下一步研究什么"难以学习的问题,原因是这类决策在长轨迹中稀少且效果滞后。MIRA 将智能体拆成两部分:一个元推理器读取持久的研究记录并写下工作指令,一个全新执行器负责执行。决策只发生在工作指令边界,作者在这些节点训练 critic 预测剩余回报,进而得到同时评估进展和选择下一步的单一 actor-critic 模型 MIRA-AC。未训练的拆分架构已能提升定理证明和开放式架构研究,MIRA-AC 在全部四个 autoresearch 环境中提高了 gold scores。
Banger paper from Meta AI on research agents that decide what to investigate next.
(bookmark it)
If you run long-horizon research agents, choosing the next investigation is hard to learn, because those decisions are rare in long traces and their effects show up several steps later.
MIRA splits the agent into two.
An outer meta-reasoner reads a persistent research record and writes a work order for the next investigation.
A fresh executor carries out each work order.
Decisions only happen at work-order boundaries, so the authors train a critic at those points to forecast remaining return, then a single actor-critic (MIRA-AC) that both values partial progress and picks the next investigation.
Even without training, the split improves theorem proving and open-ended architecture research.
Trained on the model's own proxy signals, MIRA-AC improves gold scores in all four autoresearch environments.
Paper: https://t.co/gmcUdCVB3i
Chat with Paper: https://t.co/9wJLW7o8jV