Google 论文提出 AIM:研究智能体靠想法排序地图更快出好结果
Google 发的这篇论文很实用:给智能体加个想法排序地图加审计员,比 ScientistOne 快 3 倍出结果,做 agent 的可以抄作业。
Google 的新论文介绍 AIM,让研究智能体维护一份按主题排序的想法地图,并把每轮探索分配给强主题和未测试主题。AIM 配备审计员角色,丢弃被投机取巧的结果,并把想法重新标注到与代码匹配的类别,避免代码偏离想法时学到错误经验。在 2 个任务组上,AIM 分别比此前最优的 ScientistOne 高出 1.6 和 4.9 分,并且最多提前 3.1 倍就达到 ScientistOne 的最佳分数。论文指出,在可行路径多但好路径少的任务上收益最大。
New Google paper shows research agents get better results sooner when they keep a ranked map of ideas and check that code matches each idea, so build both in.
Most agents just keep editing code.
When code drifts from the idea it's scored as, the agent learns the wrong lesson.
AIM sorts ideas into ranked themes and splits each round between strong themes and untested ones. An auditor tosses gamed results and relabels ideas to match the code.
It beat the best prior agent, ScientistOne, by 1.6 and 4.9 points on 2 task groups, and matched ScientistOne's best score up to 3.1x sooner.
Expect the biggest gains on tasks with many possible approaches and few good ones.