10:16
官方一手arXiv: Anthropic@Rob Manson
精选
推荐理由:Anthropic's research explores new ways to detect deceptive alignment in models, using geometric analysis and naturalistic contexts. It's a must-read for those interested in model interpretability and safety.