ParanoiaEval 基准发布:评测编程智能体的过度防御行为
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
编程智能体有时会做没必要的防御性改动,这个基准用 200 个任务对量化了这问题,测了 8 个模型,做智能体评测的可以看看。
研究团队推出 ParanoiaEval,这是首个统一评测编程智能体风险处理能力的基准。该基准基于软件工程风险管理中的避免-转移-缓解-接受四类处理框架,包含 200 个证据受控的仓库级任务对。实验覆盖 8 个代表性模型,结果显示 11.2%-58.7% 的运行存在不必要的风险处理,且更强的任务能力并不保证更合适的风险处理。研究还引入风险处理违规和证据响应度指标,并使用经人工校准的智能体裁判进行评估。
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents. Grounded in the well-established Avoidance-Transfer-Mitigation-Acceptance framework in software engineering risk management, ParanoiaEval operationalizes its 4 fundamental treatments for coding-agent settings and contains 200 evidence-controlled repository-level task pairs, each differing only in treatment-defining evidence. We further introduce dedicated metrics for risk-treatment violations and evidence responsiveness, using a human-calibrated agentic judge for reliable evaluation. Large-scale experiments on 8 representative models and a post-hoc human study reveal that (I) unnecessary risk treatment occurs in 11.2%-58.7% of runs despite explicit evidence, with substantial variation across agent configurations; (II) stronger task capability does not ensure more appropriate risk treatment, while treatment violations substantially harm developers' experience, establishing risk treatment as an independent capability dimension; and (III) agents exhibit systematic patterns consistent with established risk-management findings, suggesting that knowledge from human practice can guide the diagnosis and improvement of this capability.