这篇论文用消融实验戳穿了VLM做零样本控制的假象,多数模型其实没在看图像。还给出了一个能到0.973准确率的守护方案,想搞具身智能的值得看。
这项研究通过输入消融测试,检查视觉语言模型(VLM)在零样本控制中是否真正依赖视觉输入。在9个直接控制模型、6个结构化局部VLM和1个VLM-MPC层级中,分析了32,874次评分调用。结果显示多数直接控制模型表现负面,常数SLOW策略甚至优于脚本几何控制器。图像仅确定性正对照以0.090米MAE和精确镜像等变性估算前车间距,证明视觉刺激信息充分。后验对称一致性守护者选出的模型在272帧上达到0.954平衡准确率,弃权时提升至0.973。
Visual Grounding in Zero-Shot Vision-Language Control
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks. Across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy, we analyse 32,874 scored calls over two embodiments and three simulators. The direct-control results are largely negative: a constant-SLOW policy outperforms a scripted geometric controller, several models are image-invariant or nearly constant, and models that recognize longitudinal hazards still fail to transform LEFT and RIGHT under reflection. No local VLM meets the joint longitudinal and lateral grounding criteria. However, an image-only deterministic positive control estimates the lead gap with 0.090 m MAE and exact mirror equivariance, confirming the stimuli carry sufficient visual information; the failures are modular, not universal. A post-hoc, leakage-controlled symmetry-consensus guardian selects two models from 16 calibration frames and freezes a 2-of-4 hazard vote across original and reflected views. On 272 held-out frames it reaches 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); nested leave-one-episode-out recovers the same pair and threshold in all 12 folds. Abstaining on ties raises committed balanced accuracy to 0.973 at 0.824 coverage. With deterministic perception retaining lateral authority, offline modular replay achieves 0.934 action agreement and exact mirror equivariance. These results support current VLMs as bounded, selective hazard assistants, not monolithic zero-shot controllers.