模型78°

小米 MiMo-V2.6 模型 RL 训练细节公开

This should be the standard for building open-source AI. $1M+ so far.

精选理由

小米开源了他们新模型 MiMo-V2.6 的训练细节,能让你了解如何用 RL 构建开源 AI,和之前的方法比,他们把计算、环境和评估都做了更大规模的扩展。

小米的 MiMo-V2.6 模型正在进行大规模强化学习训练,当前计算量达到每步约 200 亿 tokens,使用 1568 个提示 × 16 个滚动,完全异步执行。训练同时覆盖多任务代理 RL 和多个 harness,并采用基于测试用例和评分标准的奖励机制进行评估。

原文 · elvis

This should be the standard for building open-source AI. $1M+ so far.

This should be the standard for building open-source AI. $1M+ so far. Fuli Luo @_LuoFuli Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks. Streaming the run: mimo.xiaomi.com/rl/ 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 6 👀 1544 📊 2 ⚡