The use of artificial intelligence (AI) and machine learning (ML) models to forecast urban energy consumption is becoming more widespread, but these models often rely on large, clean and well-distributed datasets. In reality, particularly at a local level, the available data are often limited, inconsistent or incomplete. This study systematically examines the robustness and reliability of AI predictive models in urban small data conditions using a real energy dataset from a neighborhood monitored over 24 months. The analysis compares several ML models trained on progressively shorter historical windows (6, 12, 18 and 24 months) and assesses performance degradation through controlled data quality stress tests, including missing values and noise. Results show that ensemble-based models achieve high accuracy when at least 18–24 months of data are available (normalized R2 up to 0.87), while performance declines markedly below 12 months. Gradient boosting demonstrates the highest robustness under severe data constraints, maintaining normalized R2 values above 0.70 with 12 months of data. Regularized linear models perform competitively in longer, well-structured time series but degrade under extreme data scarcity. An ultra-conservative data augmentation strategy yields limited but consistent improvements (≈1–2%) in short-horizon scenarios.

Assessment of the reliability of AI models in predicting urban energy consumption under conditions of small or incomplete data / Piras, G., Muzi, F., Ziran, Z.. - In: APPLIED SCIENCES. - ISSN 2076-3417. - 16:3(2026). [10.3390/app16031457]

Assessment of the reliability of AI models in predicting urban energy consumption under conditions of small or incomplete data

Piras G.
Primo
;
Muzi F.
Secondo
;
Ziran Z.
Ultimo
2026

Abstract

The use of artificial intelligence (AI) and machine learning (ML) models to forecast urban energy consumption is becoming more widespread, but these models often rely on large, clean and well-distributed datasets. In reality, particularly at a local level, the available data are often limited, inconsistent or incomplete. This study systematically examines the robustness and reliability of AI predictive models in urban small data conditions using a real energy dataset from a neighborhood monitored over 24 months. The analysis compares several ML models trained on progressively shorter historical windows (6, 12, 18 and 24 months) and assesses performance degradation through controlled data quality stress tests, including missing values and noise. Results show that ensemble-based models achieve high accuracy when at least 18–24 months of data are available (normalized R2 up to 0.87), while performance declines markedly below 12 months. Gradient boosting demonstrates the highest robustness under severe data constraints, maintaining normalized R2 values above 0.70 with 12 months of data. Regularized linear models perform competitively in longer, well-structured time series but degrade under extreme data scarcity. An ultra-conservative data augmentation strategy yields limited but consistent improvements (≈1–2%) in short-horizon scenarios.
2026
artificial intelligence (AI); building energy consumption; data augmentation; ensemble learning; machine learning models; predictive model; small data; urban energy forecasting
01 Pubblicazione su rivista::01a Articolo in rivista
Assessment of the reliability of AI models in predicting urban energy consumption under conditions of small or incomplete data / Piras, G., Muzi, F., Ziran, Z.. - In: APPLIED SCIENCES. - ISSN 2076-3417. - 16:3(2026). [10.3390/app16031457]
File allegati a questo prodotto
File Dimensione Formato  
Piras_Assessment_2026.pdf

accesso aperto

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 2.61 MB
Formato Adobe PDF
2.61 MB Adobe PDF

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1771757
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 4
  • ???jsp.display-item.citation.isi??? 2
social impact