Deep learning models have shown competitive performance on many computer vision tasks, including image-based malware detection and classification. However, evaluating adversarial robustness is critical for their application in real-world scenarios. In this paper, we propose a framework to analyze the adversarial robustness and attack transferability in image-based models for malware detection and classification. Specifically, we adopt a broad spectrum of image-domain attacks, ranging from Fast Gradient Sign Method (FGSM) to AutoAttack, to generate adversarial samples and evaluate the model performance drop and the attack success rate. Moreover, we investigate whether adversarial manipulations crafted in the binary domain remain effective after conversion to an image representation. We implement our framework on state-of-the-art models, providing a comprehensive evaluation of their adversarial robustness. Furthermore, we conduct generalization studies to analyze the capabilities of the adopted models under distribution shift. Experimental results reveal high model susceptibility to adversarial attacks, showing average attack success rate values of 64.6% and 98.8% for FGSM and AutoAttack, respectively, in the image domain. Furthermore, manipulations crafted in the binary domain remain effective upon transfer to the image domain, leading to accuracy drops up to 42%. To the best of our knowledge, this is the first work to compare a wide range of adversarial attacks on image-based malware models, analyzing the attack transferability from the binary to the image domain.
A comprehensive study of cross-domain adversarial robustness and attack transferability in image-based malware detection and classification / Daidone, G., Cirillo, L., Querzoni, L., Amerini, I.. - In: IMAGE AND VISION COMPUTING. - ISSN 0262-8856. - 176:(2026). [10.1016/j.imavis.2026.106232]
A comprehensive study of cross-domain adversarial robustness and attack transferability in image-based malware detection and classification
Giuseppe Daidone
Primo
;Lorenzo CirilloSecondo
;Leonardo QuerzoniPenultimo
;Irene AmeriniUltimo
2026
Abstract
Deep learning models have shown competitive performance on many computer vision tasks, including image-based malware detection and classification. However, evaluating adversarial robustness is critical for their application in real-world scenarios. In this paper, we propose a framework to analyze the adversarial robustness and attack transferability in image-based models for malware detection and classification. Specifically, we adopt a broad spectrum of image-domain attacks, ranging from Fast Gradient Sign Method (FGSM) to AutoAttack, to generate adversarial samples and evaluate the model performance drop and the attack success rate. Moreover, we investigate whether adversarial manipulations crafted in the binary domain remain effective after conversion to an image representation. We implement our framework on state-of-the-art models, providing a comprehensive evaluation of their adversarial robustness. Furthermore, we conduct generalization studies to analyze the capabilities of the adopted models under distribution shift. Experimental results reveal high model susceptibility to adversarial attacks, showing average attack success rate values of 64.6% and 98.8% for FGSM and AutoAttack, respectively, in the image domain. Furthermore, manipulations crafted in the binary domain remain effective upon transfer to the image domain, leading to accuracy drops up to 42%. To the best of our knowledge, this is the first work to compare a wide range of adversarial attacks on image-based malware models, analyzing the attack transferability from the binary to the image domain.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


