The integration of Artificial Intelligence (AI) into information systems has led to increasingly complex models, ranging from deep Neural Networks (NN) with millions of parameters to Large Language Models (LLMs) with billions of parameters. While cloud-based AI-as-a-Service platforms offer scalability and centralized management, they are inadequate for cyber-physical systems and edge environments where connectivity, latency, security, and resource constraints pose significant challenges. This paper presents SlimAI4Edge, a service-oriented cloud-edge framework that addresses the deployment of these models on resource-constrained devices through structural downsizing. The framework offloads the computationally intensive model reduction phase to cloud resources, maintaining local execution on edge devices. The system relies on two reduction methods, the first one is ImproveNet for deep NN and the second one is SiRE (Saliency-informed REduction) for transformer-based models. Through experiments on edge devices (Raspberry Pi 4, NVIDIA Jetson Orin, Intel NUC), we analyze the differences between the two approaches and the results obtained, pointing out how more important it is to select architectures appropriate to the device rather than delegating everything to universal solutions based on NNs/LLMs, especially when efficiency and real-time performance are required.
SlimAI4Edge: A Cloud-Edge Framework for Downsizing AI Models as-a-Service / Puglisi, A., Monti, F., Leotta, F., Napoli, C., Mecella, M.. - 16558:(2026), pp. 341-357. (38th International Conference on Advanced Information Systems Engineering, CAiSE 2026 ita ) [10.1007/978-3-032-28110-4_19].
SlimAI4Edge: A Cloud-Edge Framework for Downsizing AI Models as-a-Service
Puglisi A.;Monti F.;Leotta F.;Napoli C.;Mecella M.
2026
Abstract
The integration of Artificial Intelligence (AI) into information systems has led to increasingly complex models, ranging from deep Neural Networks (NN) with millions of parameters to Large Language Models (LLMs) with billions of parameters. While cloud-based AI-as-a-Service platforms offer scalability and centralized management, they are inadequate for cyber-physical systems and edge environments where connectivity, latency, security, and resource constraints pose significant challenges. This paper presents SlimAI4Edge, a service-oriented cloud-edge framework that addresses the deployment of these models on resource-constrained devices through structural downsizing. The framework offloads the computationally intensive model reduction phase to cloud resources, maintaining local execution on edge devices. The system relies on two reduction methods, the first one is ImproveNet for deep NN and the second one is SiRE (Saliency-informed REduction) for transformer-based models. Through experiments on edge devices (Raspberry Pi 4, NVIDIA Jetson Orin, Intel NUC), we analyze the differences between the two approaches and the results obtained, pointing out how more important it is to select architectures appropriate to the device rather than delegating everything to universal solutions based on NNs/LLMs, especially when efficiency and real-time performance are required.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


