这篇论文把热气流变成攻击红外遥感模型的武器,攻击成功率48.5%,还能让模型更自信地犯错,搞AI安全的一定要看。
研究者提出AirflowAttack,首个针对红外遥感视觉语言模型的对抗攻击方法,利用热气流湍流作为扰动先验。该方法通过轻量生成器合成输入无关的扰动,在五个CLIP骨干上平均零样本场景分类攻击成功率达48.5%,超过四个物理基线(27.7-37.0%)。对六个SOTA视觉语言模型,场景分类准确率相对降低38.2%,但某些模型反而更自信地将扰动误判为真实热证据。消融实验显示气流先验在提升物理合理性时未牺牲攻击成功率。该基准覆盖十一个模型和四项任务,暴露了红外遥感视觉语言模型的脆弱性。
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their adversarial robustness remains unexamined. We present AirflowAttack, to our knowledge the first adversarial attack for IR remote-sensing VLMs and the first to weaponize thermal-airflow turbulence as the perturbation prior. A lightweight generator synthesizes a single input-agnostic perturbation regularized toward physically plausible airflow patterns. Optimized on one surrogate CLIP model, it attains a mean zero-shot scene-classification attack success rate (ASR, the fraction of samples whose top-1 class changes) of 48.5% across five diverse CLIP backbones, far exceeding four IR-specific physical baselines (27.7--37.0%). Applied to six state-of-the-art VLMs, it cuts scene-classification accuracy by up to 38.2% relative, yet paradoxically makes some models more confident in their IR analysis, confabulating the perturbation as genuine thermal evidence such as temperature gradients and convection. Ablations show the airflow prior raises physical plausibility at no measurable cost to attack success. Together with a benchmark spanning eleven models and four tasks, these findings expose critical vulnerabilities in the rapidly expanding IR VLM ecosystem.