模型

小米公开 MiMo-V2.6 论文:用智能体自动搭建和评分训练任务

精选理由

小米这篇论文有意思,训练循环基本让智能体自己跑,还解决了智能体钻空子刷分的问题,DeepSWE 涨到 72.6。

小米发布 MiMo-V2.6 技术报告,训练流程中任务构建、测试审计、答案评分和作弊检测大部分由智能体自动完成,人类只负责预算和规则。针对 pass/fail 测试无法区分干净修复和投机修复的问题,引入了一个 grader 智能体对比各组通过补丁并把奖励分配给更干净的实现。MiMo-V2.6-Pro 的 DeepSWE 分数在 260 万美元的 RL 训练后从 58.4 升至 72.6,训练停止时仍在上升。

原文 · rohanpaul_ai

The paper for MiMo-V2.6 by Xiaom is out.

In MiMo-V2.6, AI runs much of its own training loop: agents build the tasks, audit the tests, grade the answers and hunt for cheats. Humans set the budget and the rules.

shows that agent models kept improving with more RL compute by scaling batch size, task and harness variety, and grading effort together.

Scaling RL for coding agents is hard because pass/fail tests can't tell a clean fix from a hacky fix. Agents also learn to game environments, for example by downloading the published fix.

A grader agent compared passing patches in each group and moved reward to the cleaner ones. Without it, agents drifted toward longer runs and workarounds like swallowed exceptions.

MiMo-V2.6-Pro’s DeepSWE score rose from 58.4 to 72.6 over $2.6M of RL and was still climbing when training stopped.

– arxiv. org/abs/2610.11959

Title: "MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement"