精选理由
Boris Cherny说多叠几层防御,间接提示注入能降到接近零,下周Claude Code的Auto Mode也默认开了,搞安全的可以看看。
Boris Cherny(@bcherny)在推文中称,通过叠加模型训练、输入探测和意图检查分类器等多层防御,可将间接提示注入在未见攻击上的成功率降至接近0。他表示一年前没预料到这一结果。Cherny同时宣布,Claude Code的Auto Mode将于下周起默认启用,细节见claude.com/blog/auto-mode。
原文 · AI Will
源:https://t.co/2SKaaqzZ4Y
源: x.com/bcherny/status… Boris Cherny @bcherny turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto mode is default in claude code as of next week claude.com/blog/auto-mode… 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 438 ⚡