Foundation models for computer vision built on Vision Transformer (ViT) architectures have become increasingly widespread. However, their fine-tuning process is resource-intensive, slowing their adoption in edge or low-energy applications. We introduce ALaST (Adaptive Layer Selection for ViT Fine-Tuning), a novel approach that dynamically optimizes the fine-tuning process to significantly reduce computational cost, memory consumption, and training time. Our method is founded on the critical observation that during fine-tuning, the importance of individual layers and tokens varies substantially across training iterations and depends on the specific mini-batch being processed. ALaST leverages this insight by adaptively estimating layer importance at each fine-tuning step and allocating computational resources—or “compute budgets”—proportionally. Layers assigned lower budgets are either trained with a reduced token set or temporarily frozen. Through comprehensive empirical evaluation on standard benchmarks, we demonstrate that ALaST achieves substantial efficiency gains: up to 1.3 × reduction in training time, 1.5 × reduction in FLOPs, and 2 × decrease in memory requirements, all while maintaining model performance within 0.5 % of full fine-tuning. Notably, our approach provides an automatic schedule for distributing computational resources across layers and can be combined with existing parameter-efficient fine-tuning techniques, offering an orthogonal dimension of optimization for Vision Transformers.

Adaptive layer and token selection for efficient fine-tuning of vision transformers / Devoto, A., Alvetreti, F., Pomponi, J., Di Lorenzo, P., Minervini, P., Scardapane, S.. - In: NEUROCOMPUTING. - ISSN 0925-2312. - 654:(2025). [10.1016/j.neucom.2025.131216]

Adaptive layer and token selection for efficient fine-tuning of vision transformers

Devoto, Alessio;Alvetreti, Federico;Pomponi, Jary;Di Lorenzo, Paolo;Scardapane, Simone
2025

Abstract

Foundation models for computer vision built on Vision Transformer (ViT) architectures have become increasingly widespread. However, their fine-tuning process is resource-intensive, slowing their adoption in edge or low-energy applications. We introduce ALaST (Adaptive Layer Selection for ViT Fine-Tuning), a novel approach that dynamically optimizes the fine-tuning process to significantly reduce computational cost, memory consumption, and training time. Our method is founded on the critical observation that during fine-tuning, the importance of individual layers and tokens varies substantially across training iterations and depends on the specific mini-batch being processed. ALaST leverages this insight by adaptively estimating layer importance at each fine-tuning step and allocating computational resources—or “compute budgets”—proportionally. Layers assigned lower budgets are either trained with a reduced token set or temporarily frozen. Through comprehensive empirical evaluation on standard benchmarks, we demonstrate that ALaST achieves substantial efficiency gains: up to 1.3 × reduction in training time, 1.5 × reduction in FLOPs, and 2 × decrease in memory requirements, all while maintaining model performance within 0.5 % of full fine-tuning. Notably, our approach provides an automatic schedule for distributing computational resources across layers and can be combined with existing parameter-efficient fine-tuning techniques, offering an orthogonal dimension of optimization for Vision Transformers.
2025
Efficient training; Parameter-efficient fine-tuning; Adaptive computation; Vision transformer
01 Pubblicazione su rivista::01a Articolo in rivista
Adaptive layer and token selection for efficient fine-tuning of vision transformers / Devoto, A., Alvetreti, F., Pomponi, J., Di Lorenzo, P., Minervini, P., Scardapane, S.. - In: NEUROCOMPUTING. - ISSN 0925-2312. - 654:(2025). [10.1016/j.neucom.2025.131216]
File allegati a questo prodotto
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1776550
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 4
social impact