The identification of compact, stable, and biologically meaningful gene signatures is a central challenge in high-dimensional biomedical data analysis. We propose an explainable mixed-integer linear programming (MILP) framework for robust gene signature selection that integrates predictive relevance, feature stability, empirical robustness, biological evidence, and model sparsity within a unified optimization approach. Candidate genes are scored using feature-attribution information derived from ensemble machine-learning models, cross-validation stability, an empirical robustness measure, and prior biological knowledge. These criteria are incorporated into a multi-objective optimization model solved through an ε-constraint strategy, enabling the systematic exploration of trade-offs rather than the selection of a single optimal signature. The framework is evaluated on TCGA-BRCA gene-expression data for tumor-versus-normal classification. Results show that the Pareto-efficient solution set is structured around a stable core of genes shared across solutions, while additional genes account for alternative trade-offs among robustness, biological relevance, predictive utility, and signature size. The proposed approach provides a transparent and decision-oriented methodology for gene signature selection, yielding compact and interpretable solutions while explicitly accounting for stability and robustness.

Explainable and optimization-driven Machine Learning for biomedical data analysis / Ottaviani, E., Santoni, D., Felici, G.. - (2027).

Explainable and optimization-driven Machine Learning for biomedical data analysis

Eleonora Ottaviani
Primo
;
Giovanni Felici
Ultimo
Supervision
2027

Abstract

The identification of compact, stable, and biologically meaningful gene signatures is a central challenge in high-dimensional biomedical data analysis. We propose an explainable mixed-integer linear programming (MILP) framework for robust gene signature selection that integrates predictive relevance, feature stability, empirical robustness, biological evidence, and model sparsity within a unified optimization approach. Candidate genes are scored using feature-attribution information derived from ensemble machine-learning models, cross-validation stability, an empirical robustness measure, and prior biological knowledge. These criteria are incorporated into a multi-objective optimization model solved through an ε-constraint strategy, enabling the systematic exploration of trade-offs rather than the selection of a single optimal signature. The framework is evaluated on TCGA-BRCA gene-expression data for tumor-versus-normal classification. Results show that the Pareto-efficient solution set is structured around a stable core of genes shared across solutions, while additional genes account for alternative trade-offs among robustness, biological relevance, predictive utility, and signature size. The proposed approach provides a transparent and decision-oriented methodology for gene signature selection, yielding compact and interpretable solutions while explicitly accounting for stability and robustness.
2027
File allegati a questo prodotto
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11573/1777803
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact