发新论文了,斯坦福团队搞的QuasiMoTTo,不用独立采样浪费算力,相关采样省25-47%样本,训练步骤也砍半。做推理扩展的可以看看。
QuasiMoTTo是一种新的推理计算扩展方法,通过相关采样替代独立采样,避免重复发现相同解。该方法样本覆盖率更高,且保持边缘精确的LLM分布。实验显示,在测试时扩展中,仅需25-47%的样本即可达到相同性能;在强化学习训练中,减少50%的步骤。该研究由斯坦福大学团队完成,探索了相关采样器的设计空间。
We love scaling inference compute, but it’s costly! Independently sampling parallel attempts might b...
We love scaling inference compute, but it’s costly! Independently sampling parallel attempts might be the culprit: it wastes compute rediscovering the same solutions. What if we scaled inference compute with correlated samples? Check out QuasiMoTTo by @michaelyli_ and @probablynotaz9 ! Michael Y. Li @michaelyli_ You're wasting FLOPs when scaling inference compute: by independently sampling parallel attempts, you burn compute rediscovering the same solutions. Introducing QuasiMoTTo: we scale parallel sampling with correlated samples instead! These samples have higher coverage, are marginally exact draws from the LLM, and can be generated in parallel. Result: same performance with 25-47% fewer samples in test-time scaling + 50% fewer training steps in RL! In our new paper, we explore the design space of correlated samplers. Work with co-authors @probablynotaz9 (co-lead), @gandhikanishk , @noahdgoodman , and Emily Fox! 🔗 View Quoted Tweet 💬 0 🔄 5 ❤️ 11 👀 4055 📊 3 ⚡