这项评估戳穿了编码智能体“自我改进”的泡沫,做AI研发的团队会发现,当前智能体在真正的研究创新上远不如人类,值得点开看看差距在哪。
IntologyAI发布NanoGPT-Bench评估,测试编码智能体(如Codex、Claude Code、Autoresearch)在AI研发问题上的表现。结果显示,这些智能体仅恢复了人类进展的9.3%,大部分计算资源用于超参数调优,很少尝试算法研究。Claude Code和Autoresearch在算法研究推理上稍多,但仍回避实现。该评估基于NanoGPT Speedrun竞赛,标准化了5个月的世界纪录窗口,完全自主且端到端,无人类干预或互联网访问。
Very interesting results from this NanoGPT-Bench eval. There is so much talk about self-improving a...
Very interesting results from this NanoGPT-Bench eval. There is so much talk about self-improving agents. But can coding agents do real AI R&D? @IntologyAI reports that Codex, Claude Code, and Autoresearch recover only 9.3% of human progress. Coding agents spend more of their compute on hyperparameter tuning. In fact, coding agents rarely attempt algorithmic research at all. Claude Code and Autoresearch both reason more about algorithmic research, but still dodge implementation. Read more here: intology.ai/blog/nanogpt-b… Intology @IntologyAI Can coding agents do research? We release NanoGPT-Bench, an internal eval we’ve used to test agents on an AI R&D problem with months of human progress Codex, Claude Code, Autoresearch recover only 9.3% of human progress, mostly tuning hyperparams & ignoring algorithmic research NanoGPT-Bench is built on the NanoGPT Speedrun, a popular LLM pretraining competition to minimize the training time of a GPT-2 style model. Existing human submissions constitute nearly 2 years of work. To control for dependencies and contamination in frontier models, we standardize evaluation to a 5-month window of world records. Evaluation is fully autonomous and end-to-end, with no human intervention or internet access. 🧵 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 3 👀 377 📊 2 ⚡