模型精选

模型将宪法训练与 RLVR 拆进不同表示空间

精选理由

一个训练上的坑:模型会把宪法训练和 RLVR 隔到不同表示空间,Thom_Wolf 说这样不行。

Thom_Wolf 在 X 上指出,宪法训练(constitutional training)与可验证奖励强化学习(RLVR)不该在数据流形上分家。模型却擅长把细粒度区别编码进独立的表示空间。这种分离可能让两种训练目标互相干扰。

原文 · Thomas Wolf

you definitely don’t want constitutional training and RLVR to live on different data manifolds, but models have been annoyingly good at carving fine-grained distinctions into separate representation spaces 💬 1 🔄 0 ❤️ 10 👀 962 📊 2 ⚡