后训练中的锐化税研究
MIT研究揭示预训练LLM在测试时间预算充足时可能优于后训练模型,并推出锐化税指标和PTGS采样方法。
研究显示,配备轻量推理工具的预训练大语言模型可作为智能体,尽管通过率较低,但在足够测试时间预算下,其解决方案覆盖率常超越后训练模型。研究者提出锐化税作为诊断指标,量化后训练导致的测试可扩展性损失。团队还提出了后验温度调节组采样(PTGS),一种简单的贝叶斯采样器,可根据提示难度自适应调整采样温度。
Sharpening Tax in Post-Training
"Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget."
"we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training"
"we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty."