Meta 研究发现预训练模型加轻量脚手架在智能体任务上可超越 RL 后训练版本
Meta 这篇论文挺反直觉的:给基座模型配个轻量脚手架,多采样几次,居然比 RL 后训练版本更能解智能体任务,做训练的人可以看看。
Meta Superintelligence Labs 的论文提出 Sharpening Tax 概念:RL 后训练让模型每题趋于必对或必错,一致性上升但覆盖面下降。在 BFCL v4 多轮、ACEBench 和 WebShop 上,42 对基座与后训练模型对比显示,采样数足够大时基座模型常能解出后训练模型从未解开的任务。团队据此提出 PTGS 方法,在 RL 阶段按每条提示词的估计难度设置采样温度,既缩小 Sharpening Tax 又提升 pass@1。
Great paper for all you pre-training and post-training nerds. The surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents and can even surpass post-trained counterparts with a sufficient test-time budget. DAIR.AI @dair_ai Banger paper from Meta Superintelligence Labs. They find something super interesting and unexpected. (bookmark it) Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass @1 . At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop. This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down. The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts. Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass @1 . Paper: academy.dair.ai/papers/sharpen… 🔗 View Quoted Tweet 💬 11 🔄 2 ❤️ 28 👀 2326 📊 11 ⚡