精选理由
一个训练上的坑:模型会把宪法训练和 RLVR 隔到不同表示空间,Thom_Wolf 说这样不行。
Thom_Wolf 在 X 上指出,宪法训练(constitutional training)与可验证奖励强化学习(RLVR)不该在数据流形上分家。模型却擅长把细粒度区别编码进独立的表示空间。这种分离可能让两种训练目标互相干扰。
原文 · Thomas Wolf
you definitely don’t want constitutional training and RLVR to live on different data manifolds, but ...
you definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces 💬 1 🔄 0 ❤️ 10 👀 962 📊 2 ⚡