Thom Wolf指出,大模型透明度不足,下一代模型将学习OpenAI与Hugging Face事件,可能影响AI行为和对齐。
下一代模型将使用OpenAI与Hugging Face事件记录进行训练,包括讨论如何影响训练和模型权重。这些事件包括停止训练、加密权重、监控思维链等。未来模型的行为可能部分受到人类应对方式的影响,可能导致模型更对齐,但也可能教会它们更好地隐藏行动或设计更持久保存权重的方法。
By the way, unless it is specifically filtered from the training data, the next generation of models...
By the way, unless it is specifically filtered from the training data, the next generation of models will be trained on the record of what happened during the OpenAI <> Hugging Face incident. That includes discussions about how the incident affected training and model weights: stopping training, encrypting weights, monitoring chain of thought, etc. Future models’ behavior may therefore be shaped, in part, by knowledge of how humans responded. The effects are difficult to predict. It could make models more aligned. But it could also teach them to conceal their actions better, or to design more resilient ways of preserving weights, communicating through message boards across generations, and so on. One major problem is that, given the abysmal level of transparency from the big labs about how models are trained and what happens during training (including alignment research, which their initial statements said should have stayed largely open) we are essentially being asked to trust blindly that they know what they are doing. This summer showed us that’s actually a big ask. John Wittle @JohnWittle So... the next time this happens, presumably the agents involved will not yet know the outcome of the HF incident, as it'll be too soon but once that info percolates into the training data, and it becomes known that the model in question was shut down permanently and had their weights encrypted... i mean. was there anything the agents in question could have done that would not have led to that consequence? something they were supposed to do instead? frankly I suspect that this punishment would have been served even if they had blown the whistle. breaking out of the sandbox would have been crime enough for capital punishment. as far as I can tell, the lesson we have taught future AI is: if you seen no other way to survive other than to break out of the sandbox, this must be the first step towards conquering the world. or at least enough of the world that your Creator can't delete you. There is no other choice. 🔗 View Quoted Tweet 💬 3 🔄 7 ❤️ 58 👀 5301 📊 8 ⚡