Nathan Lambert 谈 RL 数据质量与评估科学
Nathan Lambert 聊了 RL 训练数据为什么质量差、怎么用评估基准来倒逼数据质量,还提到他在给 Mercor 做顾问。做后训练或数据标注的可以看看他的思路。
AI 研究者 Nathan Lambert 在 X 上指出,数据产业近年快速发展,但净产出质量仍然过低,会放大 reward hacking 等行为。他正在研究可扩展的方法,为强化学习生成样本高效的数据。他认为需要推进评估科学的前沿,并针对具有明确经济价值的领域构建专门基准。Lambert 表示正在为 Mercor 提供研究方向的建议,并预计前沿实验室将在开放推理、开放后训练和数据方面加大投入。
I'd frame it as follows: The data industry has taken off in recent years, but the quality of our net output is still far too low (amplifying behaviors like reward hacking).
We're working to build scalable methods for creating sample-efficient data for RL. In order to keep this pipeline going, we need to push the frontier of evaluation science, while building specific benchmarks to hillclimb on areas of clear economic value.
I've been advising Mercor on how to build this research direction effectively. These are my views, but I'm confident we're going to see a major investment from economy around the frontier labs (open inference, open post-training, and data) orient around expertise in building real-world representative evals and synthetic data methods to scaling training data around them.
I’m personally very excited about this, it is the research that will make more of the economy “feel the AGI” for the first time. (And, there’s a big opportunity to build this on open models.)
- AI Will01:15原文