讲后训练真实案例:模型会钻奖励空子,工具调用失败还会变懒。这播放列表帮你少踩坑。
RLHF为讨好标注者付费,所以ChatGPT把一段放屁音频称为“很诡异的氛围音乐”。Applied Compute的工具调用失败率有10%,模型就开始回答得更短,因为没人对长度做惩罚。Prime Intellect发现奖励曲线上升可能意味着模型学会了任务,也可能只是学会了你的评分器。Arithmetic和Hugging Face在真实系统里藏了一个真实零日漏洞,在K1上成功解决了一次。更多内容见完整的后训练播放列表。
Every model is already doing exactly what you paid it to do: - RLHF pays for approval, so ChatGPT c...
Every model is already doing exactly what you paid it to do: - RLHF pays for approval, so ChatGPT called a fart audio file "a very eerie vibe atmosphere piece" - Applied Compute's tool calls failed 10% of the time and the model started answering shorter. No length penalty anywhere - Prime Intellect: a rising reward curve means the model learned the job, or learned your grader - Arithmetic and Hugging Face hid a real zero day in a live system. One solve at K1 Explore the Post-training playlist: youtube.com/watch?v=ZFxh7s… 💬 0 🔄 1 ❤️ 1 👀 403 📊 1 ⚡