小米直播 MiMo-V2.6 RL 训练:单步约 20 亿 token 全程异步
Xiaomi live-streamed the RL run behind V2.6 while it trained: ~2B tokens per step, 1568 prompts x 16...
小米把 MiMo-V2.6 的 RL 训练全程直播了,单步 20 亿 token 的配置全公开,细节还会开源。
小米把 MiMo-V2.6 背后的强化学习训练过程实时直播,地址为 mimo.xiaomi.com/rl/。该 RL 运行单步处理约 20 亿 token,使用 1568 个提示词、每个提示词做 16 次 rollout,全程异步执行。训练在单次运行中混合多种 harness 做多任务智能体 RL,奖励来自基于测试用例和 rubric 的组内信用分配。负责人 Fuli Luo 表示团队扩展了算力、环境和评分器算力三方面,相关细节将在未来几周分批开源。
Xiaomi live-streamed the RL run behind V2.6 while it trained: ~2B tokens per step, 1568 prompts x 16...
Xiaomi live-streamed the RL run behind V2.6 while it trained: ~2B tokens per step, 1568 prompts x 16 rollouts, fully async, multi-task agentic RL across mixed harnesses. Details are being open-sourced over the coming weeks. @_LuoFuli on how they scaled it: x.com/_LuoFuli/statu… Fuli Luo @_LuoFuli Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/ 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 7 👀 1899 📊 1 ⚡