行业精选73°

SnorkelAI创始人探讨AI评估与强化学习挑战

精选理由

SnorkelAI创始人分享AI评估和强化学习的真实挑战,包括公平性、奖励设计和长周期评估难题。

Jerry Liu与SnorkelAI创始人Vincent Sun Chen共进晚餐,讨论了AI评估和强化学习环境的发展。他们指出评估方法激增,但模型在基准测试中表现越来越好。公平性是构建强化学习环境的一大挑战,奖励机制设计困难,长周期评估仍极为复杂。受监管行业仍需人类参与以确保接近100%的准确率。

原文 · Jerry Liu

Yesterday I hosted a fun dinner conversation with @vincentsunnchen from @SnorkelAI on evals and RL environments. The "data and RL env" companies (like Snorkel) have seen massive growth in the past few years. There's been an explosion of interest in evals. At the same time, models are ripping through benchmarks with each new release. We talked about the evals everyone is defining, what evals are still left unsolved, what’s left up to frontier models vs. intelligence that you own, and more: - A big challenge for building RL environments is “fairness” - when the model fails on a given environment, can you attribute it to the input, harness, or reward model? - Building proper rewards is hard. Some tasks are not easily quantifiable. You also want to discourage reward hacking. At the same time, you don’t want to be too prescriptive with intermediate rewards. - Long horizon evals are still extremely hard, some business processes can take up to weeks or months before the final outcome - Most regulated industries still need human in the loop to guarantee ~100% accuracy, “80%” accuracy is not good enough - As models get more intelligent, there will be a barbell of boutique data vendors (e.g. any SMB) any scaled up data providers. - Models still exhibit “jagged intelligence” where they still fail on a long tail of edge cases. - There might always be opportunities to gather unique data for a given task and posttrain models for lower cost and higher accuracy. This marks #003 in our founder dinner series. What topic should we discuss next? Let us know your thoughts below! 💬 4 🔄 2 ❤️ 4 👀 513 📊 5 ⚡