8月26日
10:16
10:16官方一手arXiv: Anthropic@Rob Manson
精选
This paper extends Anthropic's Sleeper Agents research, introducing a naturalistic methodology for detecting deceptive alignment in models. It uses multi-turn context windows and a new metric called semantic surface area to analyze geometric patterns in model outputs, suggesting a scalable, unsupervised path for detection beyond linear methods.
事件专题

推荐理由:Anthropic's research explores new ways to detect deceptive alignment in models, using geometric analysis and naturalistic contexts. It's a must-read for those interested in model interpretability and safety.