The identification of compact, stable, and biologically meaningful gene signatures is a central challenge in high-dimensional biomedical data analysis. We propose an explainable mixed-integer linear programming (MILP) framework for robust gene signature selection that integrates predictive relevance, feature stability, empirical robustness, biological evidence, and model sparsity within a unified optimization approach. Candidate genes are scored using feature-attribution information derived from ensemble machine-learning models, cross-validation stability, an empirical robustness measure, and prior biological knowledge. These criteria are incorporated into a multi-objective optimization model solved through an ε-constraint strategy, enabling the systematic exploration of trade-offs rather than the selection of a single optimal signature. The framework is evaluated on TCGA-BRCA gene-expression data for tumor-versus-normal classification. Results show that the Pareto-efficient solution set is structured around a stable core of genes shared across solutions, while additional genes account for alternative trade-offs among robustness, biological relevance, predictive utility, and signature size. The proposed approach provides a transparent and decision-oriented methodology for gene signature selection, yielding compact and interpretable solutions while explicitly accounting for stability and robustness.
Explainable and optimization-driven Machine Learning for biomedical data analysis / Ottaviani, E., Santoni, D., Felici, G.. - (2027).
Explainable and optimization-driven Machine Learning for biomedical data analysis
Eleonora Ottaviani
Primo
;Giovanni FeliciUltimo
Supervision
2027
Abstract
The identification of compact, stable, and biologically meaningful gene signatures is a central challenge in high-dimensional biomedical data analysis. We propose an explainable mixed-integer linear programming (MILP) framework for robust gene signature selection that integrates predictive relevance, feature stability, empirical robustness, biological evidence, and model sparsity within a unified optimization approach. Candidate genes are scored using feature-attribution information derived from ensemble machine-learning models, cross-validation stability, an empirical robustness measure, and prior biological knowledge. These criteria are incorporated into a multi-objective optimization model solved through an ε-constraint strategy, enabling the systematic exploration of trade-offs rather than the selection of a single optimal signature. The framework is evaluated on TCGA-BRCA gene-expression data for tumor-versus-normal classification. Results show that the Pareto-efficient solution set is structured around a stable core of genes shared across solutions, while additional genes account for alternative trade-offs among robustness, biological relevance, predictive utility, and signature size. The proposed approach provides a transparent and decision-oriented methodology for gene signature selection, yielding compact and interpretable solutions while explicitly accounting for stability and robustness.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


