精选理由
他试了在象棋AI数据里掺8%的推理数据,效果直接碾压纯数据,轻松上1300分。
阿马德·马萨德训练国际象棋LLM时发现,单纯用棋盘→走法配对数据在投入超1000小时后遭遇明显收益递减。一次意外中,一个"先思考再走棋"的分支因产生幻觉走法而失败,但他将约8%的这种推理轨迹混合进纯走法数据,在同等训练预算下取得了优于纯数据的效果。模型通过吸收推理轨迹,在推理时不再实际执行思考过程。最终模型Elo得分达到1300。
原文 · Amjad Masad
1300 Elo! https://t.co/4WJ9SYmr1o
1300 Elo! qwen-chess.replit.app Amjad Masad @amasad hit massive diminishing returns training my chess LLM on board→move pairs. then an accident: a "think before you move" branch failed (hallucinated moves) but mixing just ~8% of those reasoning traces into plain board→move training beat pure move data at equal budget. the model internalizes the thinking without ever doing it at inference. 🔗 View Quoted Tweet 💬 5 🔄 0 ❤️ 46 👀 6018 📊 7 ⚡