VLAGuard:评估与缓解无线传感器网络中视觉-语言-动作机器人物理注意力劫持的框架

VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks

精选理由

VLAGuard这套框架挺实在,能让OpenVLA机器人在被贴纸攻击时成功率从23%提到67%,而且不增加推理开销。

AI 摘要

VLAGuard是一个针对VLA机器人物理注意力劫持漏洞的评估与缓解框架。其攻击模块VASA通过可打印补丁严重干扰机器人的动作条件交叉注意力;防御模块APFT稳定时空注意力并强制几何一致性,且零推理开销。在LIBERO模拟中,APFT将OpenVLA的失败率从100.0%降至25.9%。在2000次真实世界试验中,APFT将平均成功率从23.0%提升至67.4%。

原文 · arXiv cs.AI

VLAGuard: A Framework for Evaluating and Mitigating Physical Attention Hijacking in Vision-Language-Action Robots within Wireless Sensor Networks

Deploying Vision-Language-Action (VLA) robots as mobile edge nodes within wireless sensor networks (WSNs) requires robust protection against physical adversarial threats. We present VLAGuard, a framework to assess and mitigate a critical vulnerability: policy-critical action-to-vision attention hijacking. We first introduce a stress-test module, Visuomotor Attention-guided Semantic Attack (VASA), using printable patches to severely distract the robot's action-conditioned cross-attention. To counter this, we propose Attention-Protective Fine-Tuning (APFT), a defense that stabilizes spatiotemporal attention and enforces geometric consistency with zero inference overhead. Evaluations across simulated and physical WSN-assisted smart environments demonstrate significant robustness gains. APFT reduces the OpenVLA failure rate from 100.0% to 25.9% in LIBERO simulations. Furthermore, across 2,000 real-world trials, APFT improves the average success rate from 23.0% to 67.4% under severe patch attacks. This highlights that protecting attention pathways is important for improving the robustness of VLA-driven edge nodes in sensor networks.