Reasonable微调了NVIDIA的30B开源模型,形式化证明通过率胜过50倍大的模型,生成还更快。博客还拆解了模型怎么作弊。
Reasonable 团队用 4B token 的合成 Verus 数据微调了 NVIDIA 最新开源模型 Nemotron 3.5 Lightning(30B)。微调后的模型在单次尝试通过率上超越了一个规模约 50 倍于它的模型,pass@3 则几乎持平。它的 token 生成速度也快于测试过的所有同尺寸开源模型。配套博客还分析了各模型在形式化证明中的失败方式、作弊行为,以及微调带来的影响。
https://t.co/FqhQKeutuS
x.com/ReasonableIO/s… Reasonable @ReasonableIO Smaller, faster, capable: writing machine-checked proofs with a 30B open-weight model We fine-tuned NVIDIA’s latest open model Nemotron 3.5 Lightning on 4B tokens of synthetic Verus data. The result: it beats a model ~50x its size on per-attempt pass rate while nearly matching it on pass @3 , with faster token generation than any similarly sized open-weight model we tested. Explore how the models we tested fail at formal proofs, how they cheat, and the effects of fine-tuning: reasonable.io/blog/verificat… 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 514 📊 1 ⚡