QR-STT:红外视觉语言模型的隐秘热触发攻击

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

精选理由

红外视觉模型能被QR码图案定向误导?论文提出QR-STT,黑盒无需训练,改分类还影响问答,AI安全党可以看看。

AI 摘要

论文提出QR-STT,一种无需训练的黑盒攻击框架,用于定向操控红外视觉语言模型的语义输出。该方法将QR码划分为冷、中、热三种温度模块,并通过无梯度优化搜索模块布局和渲染参数。实验在多个CLIP风格编码器上验证,QR-STT能稳定将图像-文本对齐导向攻击者选定的概念。针对分类任务优化的扰动还能迁移到图像描述和视觉问答,使生成文本出现目标一致的语义偏移。

原文 · arXiv cs.AI

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering. However, their robustness to structured thermal perturbations and the stability of cross-modal semantic alignment remain insufficiently studied. We propose QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free, black-box framework for targeted semantic steering of IR-VLMs. QR-STT preserves the functional regions of a QR pattern while optimizing its internal modules, each of which is assigned a cold, neutral, or hot thermal state. The framework jointly searches module topology and rendering parameters, including position, scale, rotation, intensity, blur, and roundness. A three-stage gradient-free procedure with greedy module-flip refinement efficiently handles the mixed discrete and continuous search space. The objective promotes alignment with an attacker-selected target, suppresses source-class evidence, and regularizes QR structure and visual similarity. Experiments on multiple CLIP-style encoders show that QR-STT consistently redirects image-text alignment toward chosen concepts while maintaining visual stealth. Perturbations optimized for classification also transfer to image captioning and VQA, causing target-consistent semantic drift in generated outputs. These results identify QR-structured thermal patterns as an interpretable attack surface for language-driven infrared perception and highlight the need for robustness evaluation against structured cross-task semantic attacks.