Sakana AI发布递归自改进论文
Sakana AI的新论文让AI自己组织团队来改进自己,在多个研究基准上实现了1.2-1.6倍性能提升。
Sakana AI团队提出多代理自监督(MASS)方法,解决AI评估瓶颈问题。该方法通过27B参数模型在两个循环后,在MLR-Bench等基准上性能提升1.2-1.6倍。模型通过组织自身副本形成团队结构,学习协调技能并内化到权重中。
New paper from the team at Sakana AI: Recursive Self-Improvement through Multi-Agent Self-Supervision (MASS)
How do we keep improving AI when the tasks become too complex for humans to reliably evaluate?
For open-ended research, we eventually hit a "supervision bottleneck." If we use the model itself as the solver, optimizer, and evaluator, it gets trapped in a homogeneous loop, reinforcing its own blind spots rather than correcting them.
To break this loop, we drew inspiration from Marvin Minsky’s Society of Mind. What if you prompt a single model to organize copies of itself into a structured team, then distill what the team learned back into the base model?
MASS alternates between two loops:
1/ Workflow Optimization: the model iteratively evolves its own multi-agent workflows, figuring out how subagents should communicate and route information to each other.
2/ Workflow Internalization: fine-tune the shared model on the successful multi-agent trajectories so the coordination behavior gets baked into the weights.
After two cycles with a 27B open-weights model, performance improved 1.2-1.6x per output token on tough research benchmarks (MLR-Bench, ScienceAgentBench, DSBench).
Multi-agent traces were also much better training data than single-agent ones. The model learns faster from jointly learning high-level orchestration and bounded subagent execution. The results showing that the model can bootstrap its own coordination skills by organizing copies of itself into teams feels promising as a mechanism towards self-improvement.
Blog: https://t.co/rTEQ5mJxS8 Paper: https://t.co/kRvqLHxpOM
Incredible work by Hyunin Lee, Jinglue Xu, Jeffrey Seely, Yujin Tang, and our collaborators from UC Berkeley (Somayeh Sojoudi, Matei Zaharia, Donghyun Lee)!