OpenMLE把自我改进变成可跑全栈,Frontis-MA1在4090上71.21%,代码模型都开源。
OpenMLE是一个开源的递归自我改进全栈系统,包含可验证任务环境、算子学习和长时程搜索。团队在此基础上训练了Frontis-MA1,一个35B参数的元进化智能体,围绕Draft、Improve、Debug、Crossover四个原子编程演化算子构建。在MLE-Bench Lite上、单张RTX 4090且12GB显存限制下,Medal Average从39.39%提升到60.61%,配合异步搜索和经验先验达到71.21%。该成绩超过GPT-5.5 with Codex,接近GPT-5.6 Sol和2.8T的Kimi K3。论文见arxiv.org/abs/2607.28568。
Great release for recursive self-improving in ML Engineering.
Great release for recursive self-improving in ML Engineering. DAIR.AI @dair_ai Very interesting paper on recursive self-improvement. The whole stack is released. Machine learning engineering gives recursive self-improvement a concrete, executable testbed. OpenMLE is an open full-stack system for that research, spanning verifiable task environments with execution feedback, operator learning, and long-horizon search. On top of it the team post-trains Frontis-MA1, a 35B meta-evolution agent aligned around four atomic program-evolution operators. Draft, Improve, Debug, Crossover. The same four operators are trained through execution-grounded SFT and RL, then composed into long-horizon search, so learning and evolution run in one loop. On MLE-Bench Lite under a 12-hour per-task budget on a single RTX 4090 capped at 12 GB VRAM, Medal Average climbs from 39.39% to 60.61% over the base model, reaching 71.21% with asynchronous search and benchmark-independent experience priors. That exceeds GPT-5.5 with Codex and approaches GPT-5.6 Sol and the 2.8T Kimi K3. With the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%. With the model fixed, swapping in the search framework raises it from 20% to 50%. Paper: arxiv.org/abs/2607.28568 Learn to build effective AI agents in our academy: academy.dair.ai 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 6 👀 1601 📊 3 ⚡