Real-world deepfake detection requires robustness against both unseen forgery methods and adversarial attacks. These two challenges are typically studied in isolation under closed-set settings, where deepfake detectors detect images produced by generators that synthesize the training set for the detectors. This leaves open the question of how deepfake models behave when adversarial perturbations occur alongside open-set scenarios, when test images are produced by generators not included in the training phase. In this paper, we introduce a modular evaluation framework for closed-set and open-set adversarial robustness of deepfake detectors, and propose Saliency-Guided Gradient Regularization (SG-GradReg), a defense strategy that restricts feature space perturbations to the most discriminative regions of the model using a soft spatial saliency mask and a confidence-driven adaptive scaling mechanism. Our method relies on first-order computations and maintains the original parameter count without the need for auxiliary networks. We evaluate six deepfake architectures (ResNet-50, Xception, EfficientNet-B0, ConvNeXt-Base, ViT-Small, ViT-Base) on the WILD dataset under seven adversarial threats ranging from white-box and adaptive attacks to transferable black-box perturbations (FGSM, PGD, A-PGD, AutoAttack, GNP, WJSMA, BSR). SG-GradReg achieves the lowest attack success rate on ResNet50, Xception, and ViT-Small, showing an average of 11× training speedup and 2.7× lower peak memory compared to state-of-the-art defenses. We conduct an ablation study over four adaptive scaling configurations to show that the best SG-GradReg schedule depends on the gradient stability of the backbone and can be chosen with minimal overhead using a small validation set.
A comprehensive evaluation framework and efficient regularization strategy for adversarial robustness of deepfake detectors under closed and open set scenarios / Alhalawani, N., Cirillo, L., Amerini, I.. - In: COMPUTER VISION AND IMAGE UNDERSTANDING. - ISSN 1077-3142. - 271:(2027). [10.1016/j.cviu.2026.104899]
A comprehensive evaluation framework and efficient regularization strategy for adversarial robustness of deepfake detectors under closed and open set scenarios
Lorenzo Cirillo
;Irene Amerini
2027
Abstract
Real-world deepfake detection requires robustness against both unseen forgery methods and adversarial attacks. These two challenges are typically studied in isolation under closed-set settings, where deepfake detectors detect images produced by generators that synthesize the training set for the detectors. This leaves open the question of how deepfake models behave when adversarial perturbations occur alongside open-set scenarios, when test images are produced by generators not included in the training phase. In this paper, we introduce a modular evaluation framework for closed-set and open-set adversarial robustness of deepfake detectors, and propose Saliency-Guided Gradient Regularization (SG-GradReg), a defense strategy that restricts feature space perturbations to the most discriminative regions of the model using a soft spatial saliency mask and a confidence-driven adaptive scaling mechanism. Our method relies on first-order computations and maintains the original parameter count without the need for auxiliary networks. We evaluate six deepfake architectures (ResNet-50, Xception, EfficientNet-B0, ConvNeXt-Base, ViT-Small, ViT-Base) on the WILD dataset under seven adversarial threats ranging from white-box and adaptive attacks to transferable black-box perturbations (FGSM, PGD, A-PGD, AutoAttack, GNP, WJSMA, BSR). SG-GradReg achieves the lowest attack success rate on ResNet50, Xception, and ViT-Small, showing an average of 11× training speedup and 2.7× lower peak memory compared to state-of-the-art defenses. We conduct an ablation study over four adaptive scaling configurations to show that the best SG-GradReg schedule depends on the gradient stability of the backbone and can be chosen with minimal overhead using a small validation set.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


