模型精选
模型将宪法训练与 RLVR 拆进不同表示空间
精选理由
一个训练上的坑:模型会把宪法训练和 RLVR 隔到不同表示空间,Thom_Wolf 说这样不行。
Thom_Wolf 在 X 上指出,宪法训练(constitutional training)与可验证奖励强化学习(RLVR)不该在数据流形上分家。模型却擅长把细粒度区别编码进独立的表示空间。这种分离可能让两种训练目标互相干扰。
原文 · Thomas Wolf
you definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces 💬 1 🔄 0 ❤️ 10 👀 962 📊 2 ⚡