Purpose: To evaluate the inter-rater reliability of DISE interpretation using the VOTE classification, assess variability in therapeutic decision-making, and explore the potential role of artificial intelligence (AI) in DISE evaluation. Methods: Twenty DISE recordings were retrospectively and independently evaluated by 12 raters (7 junior and 5 senior) and by a general-purpose AI model without task-specific training. Inter-rater agreement among human raters for obstruction grading and pattern classification was assessed using percentage agreement, Fleiss’ kappa, and weighted kappa. Agreement with a reference standard (unanimous expert consensus) was analyzed. Treatment proposals were evaluated using exact match ratio and Jaccard similarity. Results: At baseline, agreement among human raters ranged from 43.6% to 68.0% for obstruction grading (weighted κ, 0.255–0.425) and from 60.8% to 79.3% for collapse pattern (Fleiss’ κ, 0.287–0.370). Senior raters showed higher baseline agreement than junior raters across all anatomical levels for collapse pattern and at three of four levels for obstruction grading. Agreement with the expert consensus ranged from 70.0% to 90.0% for senior raters and from 35.0% to 85.0% for AI in baseline grading, and from 75.0% to 95.0% and 50.0% to 80.0%, respectively, in baseline pattern classification. For therapeutic recommendations, exact agreement was 70.0% for senior raters, 55.0% for junior raters, and 50.0% for AI; mean Jaccard similarity was 0.867, 0.833, and 0.742, respectively. Conclusions: DISE interpretation remains complex and operator-dependent. Pattern classification appears more reproducible than grading. While a non-task-specific AI model captured clinically relevant patterns, it did not match expert performance in complex decision-making. AI may serve as a supportive tool, particularly for less experienced clinicians.

Inter-rater reliability in drug-induced sleep endoscopy: experience-driven interpretation, decision-making, and the emerging role of artificial intelligence / Leone, F., Omenetti, F., Bianchi, A., Iannella, G., Collettini, A., Nicoletti, S., Distefano, M.E., Germini Moysey, Y., Rosciano, E., Imbrogno, G., Preti, A., Vultaggio, F., Minnella, F., Zanzi, R., Salamanca, F., Ambrogi, F., Mozzanica, F.. - In: SLEEP & BREATHING. - ISSN 1522-1709. - 30:5(2026), pp. 1-8. [10.1007/s11325-026-03807-8]

Inter-rater reliability in drug-induced sleep endoscopy: experience-driven interpretation, decision-making, and the emerging role of artificial intelligence

Iannella, Giannicola
Investigation
;
Collettini, Andrea
Investigation
;
Nicoletti, Saverio
Investigation
;
2026

Abstract

Purpose: To evaluate the inter-rater reliability of DISE interpretation using the VOTE classification, assess variability in therapeutic decision-making, and explore the potential role of artificial intelligence (AI) in DISE evaluation. Methods: Twenty DISE recordings were retrospectively and independently evaluated by 12 raters (7 junior and 5 senior) and by a general-purpose AI model without task-specific training. Inter-rater agreement among human raters for obstruction grading and pattern classification was assessed using percentage agreement, Fleiss’ kappa, and weighted kappa. Agreement with a reference standard (unanimous expert consensus) was analyzed. Treatment proposals were evaluated using exact match ratio and Jaccard similarity. Results: At baseline, agreement among human raters ranged from 43.6% to 68.0% for obstruction grading (weighted κ, 0.255–0.425) and from 60.8% to 79.3% for collapse pattern (Fleiss’ κ, 0.287–0.370). Senior raters showed higher baseline agreement than junior raters across all anatomical levels for collapse pattern and at three of four levels for obstruction grading. Agreement with the expert consensus ranged from 70.0% to 90.0% for senior raters and from 35.0% to 85.0% for AI in baseline grading, and from 75.0% to 95.0% and 50.0% to 80.0%, respectively, in baseline pattern classification. For therapeutic recommendations, exact agreement was 70.0% for senior raters, 55.0% for junior raters, and 50.0% for AI; mean Jaccard similarity was 0.867, 0.833, and 0.742, respectively. Conclusions: DISE interpretation remains complex and operator-dependent. Pattern classification appears more reproducible than grading. While a non-task-specific AI model captured clinically relevant patterns, it did not match expert performance in complex decision-making. AI may serve as a supportive tool, particularly for less experienced clinicians.
2026
artificial intelligence; drug-induced sleep endoscopy; inter-rater reliability; obstructive sleep apnea; VOTE classification
01 Pubblicazione su rivista::01a Articolo in rivista
Inter-rater reliability in drug-induced sleep endoscopy: experience-driven interpretation, decision-making, and the emerging role of artificial intelligence / Leone, F., Omenetti, F., Bianchi, A., Iannella, G., Collettini, A., Nicoletti, S., Distefano, M.E., Germini Moysey, Y., Rosciano, E., Imbrogno, G., Preti, A., Vultaggio, F., Minnella, F., Zanzi, R., Salamanca, F., Ambrogi, F., Mozzanica, F.. - In: SLEEP & BREATHING. - ISSN 1522-1709. - 30:5(2026), pp. 1-8. [10.1007/s11325-026-03807-8]
File allegati a questo prodotto
File Dimensione Formato  
Leone_Inter-rater reliability_2026.pdf

solo gestori archivio

Tipologia: Versione editoriale (versione pubblicata con il layout dell'editore)
Licenza: Tutti i diritti riservati (All rights reserved)
Dimensione 752.79 kB
Formato Adobe PDF
752.79 kB Adobe PDF   Contatta l'autore

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1776008
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
social impact