花5美元用托管RL把Nemotron 3 Nano的数学准确率从22%拉到91%,还能一键扩展到更大模型,太实用了!
NVIDIA AI展示使用PrimeIntellect托管RL训练,将Nemotron 3 Nano在数学任务上的准确率从22%提升至91%,总成本低于5美元。工作流程包括基线检查、奖励训练和重新测试,最终输出可下载的LoRA适配器。该方案可通过修改一行代码扩展至Nemotron 3 Super和Ultra。
A hosted RL run took Nemotron 3 Nano from 22% to 91% accuracy on a math task for under $5. Using @P...
A hosted RL run took Nemotron 3 Nano from 22% to 91% accuracy on a math task for under $5. Using @PrimeIntellect Lab, the whole loop runs on hosted infrastructure. You check a baseline, train the model until the reward climbs, and retest to confirm it actually learned, ending with a downloadable LoRA adapter. You can also scale this workflow to Nemotron 3 Super and Ultra by changing one line. 💬 17 🔄 12 ❤️ 123 👀 8371 📊 34 ⚡