全部动态

AI 相关资讯全量信息流 · 6327 条
8月26日
论文多源确认精选10:16
Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes

Anthropic's research explores new ways to detect deceptive alignment in models, using geometric analysis and naturalistic contexts. It's a must-read for those interested in model interpretability and safety.

事件专题
arXiv: Anthropic@Rob Manson8 个信源在谈原文
8月25日
论文11:12
Traceable Spectral Inference via Influence Functions

This paper presents a novel approach to data attribution and error proxying for machine learning models in space missions, offering a more efficient and accurate way to assess model performance. It's a must-read for those interested in explainable AI and its applications in scientific research.

arXiv cs.LG@Nikki Grens 等 4 人原文
Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding

This research offers a significant advancement in privacy-enhancing instance encoding, providing stronger theoretical guarantees and broader applicability. It's a must-read for those interested in data privacy and encoding techniques.

arXiv cs.LG@Seokjin Hwang 等 4 人原文

仅展示最近 2000 条内容,更早的内容请查阅 AI 日报存档