别再靠感觉测智能体了,Phil Schmid告诉你用负面用例、500行限制和毫秒级正则来搞自动化评估,实用干货。
Phil Schmid在aiDotEngineer世博会上分享了为何依赖直觉检查智能体技能会导致生产故障,并提出了自动化评估方案。他建议添加负面测试用例以防止无关提示的关键词劫持。技能文件超过500行会降低模型推理质量,应保持精简。使用毫秒级正则断言在10-20个生产提示上验证结果。通过消融测试比较有/无技能时的表现,判断何时可退役技能。
At the @aiDotEngineer World's Fair, I gave a talk on why vibe-checking agent skills breaks in produc...
At the @aiDotEngineer World's Fair, I gave a talk on why vibe-checking agent skills breaks in production and how to build reliable automated evals before your users find the bugs. I talked about: 🎯 Why vague descriptions cause skill failures and how adding negative test cases prevents keyword hijacking on unrelated prompts ✂️ Why skill files over 500 lines degrade model reasoning ⚡ How to validate outcomes using simple, millisecond regex assertions across 10 to 20 production prompts. 🗑️ How to run ablation tests (with vs. without skills) to know when models have caught up and you can retire skills. Check out the full talk and resources below: Your browser does not support the video tag. 🔗 View on Twitter 💬 1 🔄 2 ❤️ 8 👀 1214 📊 2 ⚡