论文

VLA模型在相机故障下的物理失效模式分析

Blackout vs. Freeze: Analyzing Physical Failure Modes of VLAs under Camera Faults

精选理由

MIT团队分析了VLA模型在视觉故障下的物理失效模式,发现黑屏和冻结会导致不同危险行为,并提出两种缓解方法。

研究分析了π0.5和GR00T模型在图像黑屏和冻结故障下的行为表现。黑屏和冻结即使任务成功率相似,也会产生不同的物理失效模式。本体感知部分补偿了机器人视觉信息的缺失,但无法在腕部视图物体信息被移除时充分恢复任务成功率。两种缓解方法(相机黑屏训练和故障视觉嵌入替换)在特定条件下提高了任务成功率,但可能增加非预期接触。

原文 · arXiv cs.AI

Blackout vs. Freeze: Analyzing Physical Failure Modes of VLAs under Camera Faults

Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how $π0.5$ and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even when task-success rates are similarly low: freezing causes more extreme joint behavior, whereas blackout after gripper closure can cause more object drops, most markedly without proprioception. Selective intervention studies reveal that proprioception (current robot state) partly compensates for the removed robot depictions and reduces non-target contact. However, it cannot sufficiently restore task success when wrist-view object information is removed, even when aided by the remaining scene view. We then evaluate two mitigation approaches: camera-blackout training and training-free replacement of faulty visual embeddings. Both improve task success in selected conditions, but can increase unintended contact or disturbance to surrounding objects. Real-robot trials further show that successful execution under camera faults can still involve unintended physical interactions. These findings motivate designing VLA policies that use the robot and object information still available under camera faults to limit hazardous motion.