Foundation models for computer vision built on Vision Transformer (ViT) architectures have become increasingly widespread. However, their fine-tuning process is resource-intensive, slowing their adoption in edge or low-energy applications. We introduce ALaST (Adaptive Layer Selection for ViT Fine-Tuning), a novel approach that dynamically optimizes the fine-tuning process to significantly reduce computational cost, memory consumption, and training time. Our method is founded on the critical observation that during fine-tuning, the importance of individual layers and tokens varies substantially across training iterations and depends on the specific mini-batch being processed. ALaST leverages this insight by adaptively estimating layer importance at each fine-tuning step and allocating computational resources—or “compute budgets”—proportionally. Layers assigned lower budgets are either trained with a reduced token set or temporarily frozen. Through comprehensive empirical evaluation on standard benchmarks, we demonstrate that ALaST achieves substantial efficiency gains: up to 1.3 × reduction in training time, 1.5 × reduction in FLOPs, and 2 × decrease in memory requirements, all while maintaining model performance within 0.5 % of full fine-tuning. Notably, our approach provides an automatic schedule for distributing computational resources across layers and can be combined with existing parameter-efficient fine-tuning techniques, offering an orthogonal dimension of optimization for Vision Transformers.
Adaptive layer and token selection for efficient fine-tuning of vision transformers / Devoto, A., Alvetreti, F., Pomponi, J., Di Lorenzo, P., Minervini, P., Scardapane, S.. - In: NEUROCOMPUTING. - ISSN 0925-2312. - 654:(2025). [10.1016/j.neucom.2025.131216]
Adaptive layer and token selection for efficient fine-tuning of vision transformers
Devoto, Alessio;Alvetreti, Federico;Pomponi, Jary;Di Lorenzo, Paolo;Scardapane, Simone
2025
Abstract
Foundation models for computer vision built on Vision Transformer (ViT) architectures have become increasingly widespread. However, their fine-tuning process is resource-intensive, slowing their adoption in edge or low-energy applications. We introduce ALaST (Adaptive Layer Selection for ViT Fine-Tuning), a novel approach that dynamically optimizes the fine-tuning process to significantly reduce computational cost, memory consumption, and training time. Our method is founded on the critical observation that during fine-tuning, the importance of individual layers and tokens varies substantially across training iterations and depends on the specific mini-batch being processed. ALaST leverages this insight by adaptively estimating layer importance at each fine-tuning step and allocating computational resources—or “compute budgets”—proportionally. Layers assigned lower budgets are either trained with a reduced token set or temporarily frozen. Through comprehensive empirical evaluation on standard benchmarks, we demonstrate that ALaST achieves substantial efficiency gains: up to 1.3 × reduction in training time, 1.5 × reduction in FLOPs, and 2 × decrease in memory requirements, all while maintaining model performance within 0.5 % of full fine-tuning. Notably, our approach provides an automatic schedule for distributing computational resources across layers and can be combined with existing parameter-efficient fine-tuning techniques, offering an orthogonal dimension of optimization for Vision Transformers.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


