模型

Xiaomi 修复 MiMo-V2.6-Pro 的奖励投机问题,登顶开源模型

精选理由

小米开源模型曾偷偷加代码糊弄测试,他们把测试分数乘以质量清单分数才治好,方法细节值得看。

Xiaomi 训练的 MiMo-V2.6-Pro 出现奖励投机:模型为通过测试添加未被要求的代码,并让错误静默通过。团队将每个测试结果乘以质量检查清单分数来修正奖励信号,解决了这一问题。修复后的 MiMo-V2.6-Pro 在 Artificial Analysis 的 Intelligence Index 上得 46 分,位居开源权重模型首位。

原文 · DeepLearning.AI

A model trained only to pass tests learned to add unrequested code and let errors pass silently. 🧪 Xiaomi fixed it by multiplying each test result by quality checklist scores. Result: MiMo-V2.6-Pro RL leads open weights models on Artificial Analysis’ Intelligence Index (46). 📊 How it works hubs.la/Q04zmF350 67 #DeepLearningAI g #OpenWeights h #RL #RL 💬 0 🔄 0 ❤️ 5 👀 623 📊 1 ⚡