NAVSIM v2.2评分有个大坑:数值后端一换,不看路的Ignore-All反而排名第一。论文给出复现步骤和审计协议,跑防御性驾驶评估前先测回放稳定性。
NAVSIM v2.2原始场景评分遭审计,数值后端不稳定使路线盲Ignore-All探测和演员盲探测在12,146-token navtest分割上排名高于人类回放与PDM-Closed。新安装按公共规范在32-token诊断集复现回放发散,450-token控制池仅替换求解器即消除发散并恢复盲后排序。研究表明参考条件原谅将共享参考失败传播为合规信用,作者提出分数基础与栈披露、盲探针及回放稳定性测试等审计协议。
When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit
Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do not. Re-simulation benchmarks may use reference-conditioned forgiveness, under which an agent receives credit when the logged human reference fails a compliance channel. When agent and reference share an unstable rollout transformation, this rule can propagate shared reference failures into broad compliance credit. We audit this risk in NAVSIM v2.2 original scene single-stage scoring. Under the affected documented-stack condition on the audited numerical backend, the route-blind Ignore-All probe and a route-aware actor-blind probe outrank human replay and PDM-Closed over the complete 12,146-token navtest split. A fresh installation following the public specification reproduces rollout divergence on a fixed 32-token diagnostic set. A same-source dependency stack control and an exact-input diagnostic isolate dependency-sensitive numerical behavior in the shared velocity refit. On a 450-token control pool, replacing only the solver eliminates rollout divergence and restores blind-last ordering while keeping forgiveness enabled. Thus, the numerical instability is the direct trigger. Reference-conditioned forgiveness propagates the resulting shared reference failures into compliance credit. We contribute an audit protocol requiring score basis and stack disclosure, blind probes, overwrite reporting, and rollout stability tests before using such scores for defensive driving claims.