单提示词即可洗除水印:基础图像模型的新威胁

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

精选理由

这篇论文揭示了基础图像模型如何被用来洗除水印,比传统攻击方式更隐蔽有效。

AI 摘要

研究人员发现,攻击者可通过单个重建提示词,从OpenAI和Google的六种图像编辑模型中获取无法可靠解码水印的输出。研究评估了三种水印方案和1800个重建结果,发现OpenAI模型在有效载荷干扰方面表现最强,而Nano Banana 2显示DwtDct在高保真重建下仍然脆弱。提示词消融实验表明,无需特定的移除指令即可实现有效载荷干扰,这主要是由重建路径而非攻击性措辞引起的。

原文 · arXiv: OpenAI

One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single reconstruction prompt and obtain a visually faithful output from which the invisible watermark can no longer be decoded reliably. We formalize this failure mode as watermark laundering and evaluate it using a joint payload-fidelity profile that combines bit error rate (BER) with visual and semantic preservation. Across six OpenAI and Google image editing models, three representative watermarking schemes, and 1,800 reconstructed outputs, we identify two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction. Prompt ablations show that no single removal-oriented instruction is necessary for payload disruption, indicating that the effect is primarily induced by the reconstruction pathway rather than by explicit attack wording. Comparisons with conventional attacks further show that prompt-conditioned reconstruction constitutes a distinct operational attack interface. These findings motivate foundation-model reconstruction as a missing robustness condition in invisible watermark evaluation.