Real-world deepfake detection requires robustness against both unseen forgery methods and adversarial attacks. These two challenges are typically studied in isolation under closed-set settings, where deepfake detectors detect images produced by generators that synthesize the training set for the detectors. This leaves open the question of how deepfake models behave when adversarial perturbations occur alongside open-set scenarios, when test images are produced by generators not included in the training phase. In this paper, we introduce a modular evaluation framework for closed-set and open-set adversarial robustness of deepfake detectors, and propose Saliency-Guided Gradient Regularization (SG-GradReg), a defense strategy that restricts feature space perturbations to the most discriminative regions of the model using a soft spatial saliency mask and a confidence-driven adaptive scaling mechanism. Our method relies on first-order computations and maintains the original parameter count without the need for auxiliary networks. We evaluate six deepfake architectures (ResNet-50, Xception, EfficientNet-B0, ConvNeXt-Base, ViT-Small, ViT-Base) on the WILD dataset under seven adversarial threats ranging from white-box and adaptive attacks to transferable black-box perturbations (FGSM, PGD, A-PGD, AutoAttack, GNP, WJSMA, BSR). SG-GradReg achieves the lowest attack success rate on ResNet50, Xception, and ViT-Small, showing an average of 11× training speedup and 2.7× lower peak memory compared to state-of-the-art defenses. We conduct an ablation study over four adaptive scaling configurations to show that the best SG-GradReg schedule depends on the gradient stability of the backbone and can be chosen with minimal overhead using a small validation set.

A comprehensive evaluation framework and efficient regularization strategy for adversarial robustness of deepfake detectors under closed and open set scenarios / Alhalawani, N., Cirillo, L., Amerini, I.. - In: COMPUTER VISION AND IMAGE UNDERSTANDING. - ISSN 1077-3142. - 271:(2027). [10.1016/j.cviu.2026.104899]

A comprehensive evaluation framework and efficient regularization strategy for adversarial robustness of deepfake detectors under closed and open set scenarios

Lorenzo Cirillo
;
Irene Amerini
2027

Abstract

Real-world deepfake detection requires robustness against both unseen forgery methods and adversarial attacks. These two challenges are typically studied in isolation under closed-set settings, where deepfake detectors detect images produced by generators that synthesize the training set for the detectors. This leaves open the question of how deepfake models behave when adversarial perturbations occur alongside open-set scenarios, when test images are produced by generators not included in the training phase. In this paper, we introduce a modular evaluation framework for closed-set and open-set adversarial robustness of deepfake detectors, and propose Saliency-Guided Gradient Regularization (SG-GradReg), a defense strategy that restricts feature space perturbations to the most discriminative regions of the model using a soft spatial saliency mask and a confidence-driven adaptive scaling mechanism. Our method relies on first-order computations and maintains the original parameter count without the need for auxiliary networks. We evaluate six deepfake architectures (ResNet-50, Xception, EfficientNet-B0, ConvNeXt-Base, ViT-Small, ViT-Base) on the WILD dataset under seven adversarial threats ranging from white-box and adaptive attacks to transferable black-box perturbations (FGSM, PGD, A-PGD, AutoAttack, GNP, WJSMA, BSR). SG-GradReg achieves the lowest attack success rate on ResNet50, Xception, and ViT-Small, showing an average of 11× training speedup and 2.7× lower peak memory compared to state-of-the-art defenses. We conduct an ablation study over four adaptive scaling configurations to show that the best SG-GradReg schedule depends on the gradient stability of the backbone and can be chosen with minimal overhead using a small validation set.
2027
deepfake detection; adversarial attacks; open-set robustness; efficient robustness
01 Pubblicazione su rivista::01a Articolo in rivista
A comprehensive evaluation framework and efficient regularization strategy for adversarial robustness of deepfake detectors under closed and open set scenarios / Alhalawani, N., Cirillo, L., Amerini, I.. - In: COMPUTER VISION AND IMAGE UNDERSTANDING. - ISSN 1077-3142. - 271:(2027). [10.1016/j.cviu.2026.104899]
File allegati a questo prodotto
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1774530
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 0
social impact