The integration of Artificial Intelligence (AI) into information systems has led to increasingly complex models, ranging from deep Neural Networks (NN) with millions of parameters to Large Language Models (LLMs) with billions of parameters. While cloud-based AI-as-a-Service platforms offer scalability and centralized management, they are inadequate for cyber-physical systems and edge environments where connectivity, latency, security, and resource constraints pose significant challenges. This paper presents SlimAI4Edge, a service-oriented cloud-edge framework that addresses the deployment of these models on resource-constrained devices through structural downsizing. The framework offloads the computationally intensive model reduction phase to cloud resources, maintaining local execution on edge devices. The system relies on two reduction methods, the first one is ImproveNet for deep NN and the second one is SiRE (Saliency-informed REduction) for transformer-based models. Through experiments on edge devices (Raspberry Pi 4, NVIDIA Jetson Orin, Intel NUC), we analyze the differences between the two approaches and the results obtained, pointing out how more important it is to select architectures appropriate to the device rather than delegating everything to universal solutions based on NNs/LLMs, especially when efficiency and real-time performance are required.
SlimAI4Edge: A Cloud-Edge Framework for Downsizing AI Models as-a-Service / Puglisi, A., Monti, F., Leotta, F., Napoli, C., Mecella, M.. - 16558:(2026), pp. 341-357. (38th International Conference on Advanced Information Systems Engineering, CAiSE 2026 Verona; Italy ) [10.1007/978-3-032-28110-4_19].
SlimAI4Edge: A Cloud-Edge Framework for Downsizing AI Models as-a-Service
Puglisi A.;Monti F.
;Leotta F.;Napoli C.;Mecella M.
2026
Abstract
The integration of Artificial Intelligence (AI) into information systems has led to increasingly complex models, ranging from deep Neural Networks (NN) with millions of parameters to Large Language Models (LLMs) with billions of parameters. While cloud-based AI-as-a-Service platforms offer scalability and centralized management, they are inadequate for cyber-physical systems and edge environments where connectivity, latency, security, and resource constraints pose significant challenges. This paper presents SlimAI4Edge, a service-oriented cloud-edge framework that addresses the deployment of these models on resource-constrained devices through structural downsizing. The framework offloads the computationally intensive model reduction phase to cloud resources, maintaining local execution on edge devices. The system relies on two reduction methods, the first one is ImproveNet for deep NN and the second one is SiRE (Saliency-informed REduction) for transformer-based models. Through experiments on edge devices (Raspberry Pi 4, NVIDIA Jetson Orin, Intel NUC), we analyze the differences between the two approaches and the results obtained, pointing out how more important it is to select architectures appropriate to the device rather than delegating everything to universal solutions based on NNs/LLMs, especially when efficiency and real-time performance are required.| File | Dimensione | Formato | |
|---|---|---|---|
|
Puglisi_SlimAI4Edge_2026.pdf
solo gestori archivio
Tipologia:
Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza:
Tutti i diritti riservati (All rights reserved)
Dimensione
20.95 MB
Formato
Adobe PDF
|
20.95 MB | Adobe PDF | Contatta l'autore |
|
Puglisi_SlimAI4Edge_2026.pdf
solo gestori archivio
Tipologia:
Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza:
Tutti i diritti riservati (All rights reserved)
Dimensione
9.81 MB
Formato
Adobe PDF
|
9.81 MB | Adobe PDF | Contatta l'autore |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


