Linking a ground-view image with a satellite one representing the same location, known as the ground-to-aerial image matching task, poses a significant challenge in the field of computer vision given the drastic variation in perspective, image quality, and field of view. Recent models mainly focus on extracting invariant features from ground and aerial perspectives to reduce their domain gap, often relying on opaque feature representations that make their decisions difficult to interpret. In this work, we introduce Attribution-Regularized Ground-to-Aerial (Arg2a), a regularization technique designed to enhance the explainability of deep models for ground-to-aerial retrieval. Our method encourages the network to focus on spatially and semantically consistent input regions, resulting in more aligned and interpretable patterns in both ground and aerial views. Quantitative analyses validate our method, showing improvements in top-k retrieval accuracy, insertion, deletion, and attribution fidelity compared to the baselines on a subset of the CVUSA dataset. Moreover, the experimental results show the Arg2a technique improves the robustness of the regularized model against pixel-level perturbations, providing clearer insights into perturbed scenarios. Qualitatively, the generated explainability maps focus on semantically interpretable structures (e.g., trees, streets, buildings), and remain consistent across both ground and aerial views, providing an improved semantic alignment and geometric patterns disambiguation.
Attribution-Guided Regularization: an Explainable Perspective for Ground-to-Aerial Image Matching Models / Halitchi, A., Pro, F., Cirillo, L., Amerini, I.. - 3107:(2027), pp. 46-60. (2nd International Workshop on Explainable AI in Space, EASi 2026, Held in Conjunction with the 35th International Joint Conference on Artificial Intelligence, IJCAI-ECAI 2026 deu ) [10.1007/978-3-032-36806-5_4].
Attribution-Guided Regularization: an Explainable Perspective for Ground-to-Aerial Image Matching Models
Andrei Halitchi;Francesco Pro;Lorenzo Cirillo
;Irene Amerini
2027
Abstract
Linking a ground-view image with a satellite one representing the same location, known as the ground-to-aerial image matching task, poses a significant challenge in the field of computer vision given the drastic variation in perspective, image quality, and field of view. Recent models mainly focus on extracting invariant features from ground and aerial perspectives to reduce their domain gap, often relying on opaque feature representations that make their decisions difficult to interpret. In this work, we introduce Attribution-Regularized Ground-to-Aerial (Arg2a), a regularization technique designed to enhance the explainability of deep models for ground-to-aerial retrieval. Our method encourages the network to focus on spatially and semantically consistent input regions, resulting in more aligned and interpretable patterns in both ground and aerial views. Quantitative analyses validate our method, showing improvements in top-k retrieval accuracy, insertion, deletion, and attribution fidelity compared to the baselines on a subset of the CVUSA dataset. Moreover, the experimental results show the Arg2a technique improves the robustness of the regularized model against pixel-level perturbations, providing clearer insights into perturbed scenarios. Qualitatively, the generated explainability maps focus on semantically interpretable structures (e.g., trees, streets, buildings), and remain consistent across both ground and aerial views, providing an improved semantic alignment and geometric patterns disambiguation.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


