Achieving safe and autonomous glycemic regulation for Type-1 Diabetes care is an urgent challenge. Although Reinforcement Learning (RL) emerged as a promising paradigm, practical deployment is hindered by the risk of uncontrolled hyperglycemia or hypoglycemia. This work adapts two safe deep RL approaches in the context of automated insulin delivery. The first consists of a Lagrangian constrained Markov decision process that solves a primal–dual scheme with adaptive multipliers, thereby delivering constraint satisfaction in expectation; the second adopts a Barrier–Lyapunov Actor–Critic framework that embeds discrete-time control-barrier conditions and Lyapunov decrease into the learning updates, ensuring stepwise feasibility and promoting stability by design. Simulations under randomized meal timing and size, benchmarked against a standard clinical practice protocol and an unconstrained DRL baseline, indicate improved time-in-range with reduced hypoglycemic events.

Safe Deep Reinforcement Learning Control of Type 1 Diabetes / Baldisseri, F., Lops, G., Atanasious, M.M.H., Menegatti, D., Becchetti, V., Delli Priscoli, F., Mascolo, S., Wrona, A.. - (2026), pp. 3603-3608. (2026 European Control Conference (ECC) Rekjavik ).

Safe Deep Reinforcement Learning Control of Type 1 Diabetes

Federico BALDISSERI
;
Mohab M. H. ATANASIOUS;Danilo MENEGATTI;Valentina BECCHETTI;Francesco DELLI PRISCOLI;Saverio MASCOLO;Andrea WRONA
2026

Abstract

Achieving safe and autonomous glycemic regulation for Type-1 Diabetes care is an urgent challenge. Although Reinforcement Learning (RL) emerged as a promising paradigm, practical deployment is hindered by the risk of uncontrolled hyperglycemia or hypoglycemia. This work adapts two safe deep RL approaches in the context of automated insulin delivery. The first consists of a Lagrangian constrained Markov decision process that solves a primal–dual scheme with adaptive multipliers, thereby delivering constraint satisfaction in expectation; the second adopts a Barrier–Lyapunov Actor–Critic framework that embeds discrete-time control-barrier conditions and Lyapunov decrease into the learning updates, ensuring stepwise feasibility and promoting stability by design. Simulations under randomized meal timing and size, benchmarked against a standard clinical practice protocol and an unconstrained DRL baseline, indicate improved time-in-range with reduced hypoglycemic events.
2026
2026 European Control Conference (ECC)
Adaptive control systems; Constrained optimization; Deep learning; Deep reinforcement learning; Discrete time control systems; Lagrange multipliers; Learning algorithms; Markov processes; Reinforcement learning
04 Pubblicazione in atti di convegno::04b Atto di convegno in volume
Safe Deep Reinforcement Learning Control of Type 1 Diabetes / Baldisseri, F., Lops, G., Atanasious, M.M.H., Menegatti, D., Becchetti, V., Delli Priscoli, F., Mascolo, S., Wrona, A.. - (2026), pp. 3603-3608. (2026 European Control Conference (ECC) Rekjavik ).
File allegati a questo prodotto
File Dimensione Formato  
Baldisseri_preprint_Safe-Deep-Reinforcement_2026.pdf

accesso aperto

Tipologia: Documento in Pre-print (manoscritto inviato all'editore, precedente alla peer review)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 483.53 kB
Formato Adobe PDF
483.53 kB Adobe PDF
Baldisseri_Safe-Deep-Reinforcement_2026.pdf

solo gestori archivio

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 3.51 MB
Formato Adobe PDF
3.51 MB Adobe PDF   Contatta l'autore

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1753378
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact