Meta发布自我提升智能体研究
Banger paper from Meta Superintelligence Labs on self-improving agents. (bookmark it) It's hard to...
Meta研究了不同模型自我提升效率,发现学得快的模型会重用已有成果,而慢学习者会忽略已写内容。
Meta Superintelligence Labs发表关于自我提升智能体的研究论文。他们提出"智能体可塑性"概念,即在保持模型权重冻结的情况下,每美元投入在学习保留任务上的增益。在象棋、围棋和Hex游戏中,Claude Fable 5达到最高最终得分,而GPT-5.6 Sol每美元收益最高。在NetHack中,只有Claude Opus 5.5显著提升,约花费1073美元学习后提高了66个标准化点。
Banger paper from Meta Superintelligence Labs on self-improving agents. (bookmark it) It's hard to...
Banger paper from Meta Superintelligence Labs on self-improving agents. (bookmark it) It's hard to know exactly what drives self-improvement, since so many variables are at play (notes, skills, tool calls between runs, etc.). Meta researchers explore and discuss a way to measure whether self-improvement pays off. They call it agent plasticity. It is the gain on held-out tasks per dollar spent on learning, with model weights frozen and every run starting from a fresh context. They find that the model that performs the best is often a different model from the one that learns most efficiently. In chess, Go, and Hex, Claude Fable 5 reaches the highest final score, while GPT-5.6 Sol gains the most per dollar. In NetHack, only Claude Opus 5.5 improves significantly, by 66 normalized points for about $1,073 of learning. Another interesting finding is that slow learners often ignore artifacts they already wrote. Faster learners reuse their artifacts and still fail when an artifact is low quality. Paper: arxiv.org/abs/2610.08902 Chat with Paper: academy.dair.ai/papers/agent-p… 💬 30 🔄 7 ❤️ 95 👀 7310 📊 52 ⚡