这篇论文告诉你,连后门攻击的触发器颜色都不能随便选。在CelebA发色任务上白trigger专克金发、黑trigger专克黑发,实验设计很扎实。
该论文研究联邦学习中的语义后门攻击,利用口罩、墨镜等自然视觉对象作为触发器,仅改变颜色。在四类CelebA发色分类任务上,白色触发器对攻击金发类别更有效(成功率显著更高),黑色触发器对攻击黑发类别更有效。实验采用标准投毒目标与SABLE增强目标(结合分类损失、触发目标损失、特征分离损失及正则化),发现即使语义、位置和投毒预算不变,颜色也能显著改变攻击成功率,该结论在鲁棒聚合下依然成立。
Color Matters: Trigger Color Affects Success in Federated Backdoor Attacks
Federated learning is vulnerable to backdoor attacks in which malicious clients inject poisoned updates while preserving benign-task performance. In this paper, we study a semantics-driven backdoor mechanism in which attackers use natural visual accessories as triggers and manipulate only the trigger color while keeping the attack pipeline fixed. Our framework considers semantic trigger objects such as masks and sunglasses, instantiated in black and white variants, and evaluates their effect in a controlled federated learning setting. Malicious clients construct poisoned samples by applying a trigger to source-class images and relabeling them to an attacker-chosen target class, while benign clients train only on clean data. We analyze this mechanism under both a standard poisoning objective and a stronger SABLE-based objective that combines clean classification loss, triggered target loss, feature-separation loss in the penultimate representation space, and regularization to keep malicious updates close to the global model. This design enables the attack to remain effective while reducing excessive update drift. Experiments on a four-class CelebA hair-color task show that trigger color significantly changes attack success rate even when trigger semantics, placement, and poisoning budget are unchanged. White triggers are more effective for attacks targeting the blond class, whereas black triggers perform better for attacks targeting the black class. The same trend persists under robust aggregation, showing that trigger color is a meaningful factor in the operation, persistence, and evaluation of semantic backdoor mechanisms in federated learning.