Anomaly detection in industrial visual inspection is challenged by the scarcity of defective samples and the computational demands of modern architectures. While recent methods have achieved strong performance through synthetic anomaly generation or pretrained feature extractors, these approaches introduce domain bias and may fail to capture the true distribution of real-world defects. Vision Transformer based methods offer powerful global context modeling but suffer from slow convergence and high data requirements that limit practical deployment. In this paper, we propose Caps-ViT, a reconstruction-based architecture by training solely on normal, real-world industrial imagery. Our approach eliminates dependence on artificial defect synthesis and external pretraining through three key contributions: (1) the integration of a DigitCaps layer for hierarchical feature encoding in the latent space, (2) a decaying noise schedule for the Gaussian Mixture Density Network that balances regularization with reconstruction precision across training stages, and (3) the adoption of Level Weighted SSIM loss to accelerate color and texture convergence. We evaluate Caps-ViT on five benchmark datasets comprising 39 product categories: MVTec AD, BTAD, KSDD, VisA, and the recently introduced MVTec AD 2. Training exclusively on normal authentic images, our approach surpasses the synthetic-free transformer-based VT-ADL, in 27 out of 31 categories on established benchmarks. On the challenging MVTec AD 2 dataset, Caps-ViT achieves 60.4% AU-PRO0.30, surpassing the embedding-based method EfficientAD (58.7%) by +1.7%. Latent space analysis reveals that performance correlates strongly with the geometric separability of normal and anomalous representations, providing interpretable insights into model behavior across different anomaly types.

Caps-ViT: A synthetic-free capsule vision Transformer for reconstruction based anomaly detection and localization / Onorato, G., Tarantelli, K., Guida, A., Schiavella, C., Amerini, I.. - (2026). [10.2139/ssrn.6301415]

Caps-ViT: A synthetic-free capsule vision Transformer for reconstruction based anomaly detection and localization

Gabriele Onorato;Kristjan Tarantelli;Alberto Guida;Claudio Schiavella
;
Irene Amerini
2026

Abstract

Anomaly detection in industrial visual inspection is challenged by the scarcity of defective samples and the computational demands of modern architectures. While recent methods have achieved strong performance through synthetic anomaly generation or pretrained feature extractors, these approaches introduce domain bias and may fail to capture the true distribution of real-world defects. Vision Transformer based methods offer powerful global context modeling but suffer from slow convergence and high data requirements that limit practical deployment. In this paper, we propose Caps-ViT, a reconstruction-based architecture by training solely on normal, real-world industrial imagery. Our approach eliminates dependence on artificial defect synthesis and external pretraining through three key contributions: (1) the integration of a DigitCaps layer for hierarchical feature encoding in the latent space, (2) a decaying noise schedule for the Gaussian Mixture Density Network that balances regularization with reconstruction precision across training stages, and (3) the adoption of Level Weighted SSIM loss to accelerate color and texture convergence. We evaluate Caps-ViT on five benchmark datasets comprising 39 product categories: MVTec AD, BTAD, KSDD, VisA, and the recently introduced MVTec AD 2. Training exclusively on normal authentic images, our approach surpasses the synthetic-free transformer-based VT-ADL, in 27 out of 31 categories on established benchmarks. On the challenging MVTec AD 2 dataset, Caps-ViT achieves 60.4% AU-PRO0.30, surpassing the embedding-based method EfficientAD (58.7%) by +1.7%. Latent space analysis reveals that performance correlates strongly with the geometric separability of normal and anomalous representations, providing interpretable insights into model behavior across different anomaly types.
2026
File allegati a questo prodotto
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1774432
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact