精选理由
Anthropic 让 Sonnet 5 训练 Opus 4.8,安全分数接近原版,模型自我对齐有新进展。
Anthropic 让 Sonnet 5 对 Opus 4.8 的早期检查点进行后训练。经过训练后,该模型的安全分数接近完整对齐训练的 Opus 4.8 水平。这是测试模型能否对其更强大的继任者的首次尝试。
原文 · Anthropic
Could a model one day align its stronger successors? As a first test, we had Sonnet 5 post-train a...
Could a model one day align its stronger successors? As a first test, we had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model. It reached safety scores approaching those of production Opus 4.8, which went through our full alignment training. 💬 2 🔄 6 ❤️ 76 👀 11837 📊 9 ⚡