一个叫Faraday的27B智能体,把论文复现变成强化学习任务,在测试集上赢了Claude Opus 4.8和GPT-5.5,还把编码agent当工具用,有点意思。
Faraday是一个27B参数的智能体,基于Replica框架将论文复现转化为可扩展的强化学习任务。它调用编码agent作为工具,在held-out研究复现上击败了Claude Opus 4.8和GPT-5.5。奖励信号来自自动生成的rubric judge,与人类评估一致性高。论文已发布在arxiv.org/abs/2608.13331。
A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication. Replica turns p...
A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication. Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified. The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality. Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric. The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses. Paper: arxiv.org/abs/2608.13331 Track more trending AI papers in our academy: academy.dair.ai 💬 5 🔄 8 ❤️ 30 👀 2133 📊 11 ⚡
- Decoder16:06原文