Explainable machine learning models for predicting long-term clinical outcomes in chronic disease management
Abstract
Chronic disease management increasingly requires tools that can identify patients at elevated long-term risk before irreversible complications occur. This study aimed to develop and validate explainable machine learning models for predicting 5-year major adverse clinical events among patients with coexisting type 2 diabetes and hypertension. The central objective was to evaluate whether high predictive performance could be combined with transparent explanations suitable for clinical decision support. A retrospective longitudinal electronic health record dataset was analyzed for 15,000 adult patients followed over 5 years. The dataset included 50 predictor variables covering demographics, laboratory trajectories, medication use, comorbidity burden, healthcare utilization, and neighborhood-level socioeconomic deprivation. XGBoost, random forest, and penalized logistic regression models were trained and evaluated using 5-fold cross-validation. The primary outcome was a composite 5-year major adverse event endpoint comprising myocardial infarction, ischemic stroke, end-stage renal disease, or all-cause mortality. XGBoost achieved the highest discrimination, with an area under the receiver operating characteristic curve of 0.83, followed by random forest and logistic regression. Calibration results showed acceptable agreement between predicted and observed risk after probability recalibration. SHAP analysis identified HbA1c variability, medication adherence, estimated glomerular filtration rate decline, systolic blood pressure variability, prior cardiovascular disease, and socioeconomic deprivation index as the most influential predictors. Local explanations demonstrated how individual-level risk scores were driven by both modifiable clinical factors and structural risk indicators. Decision curve analysis indicated that the XGBoost model provided greater net clinical benefit than treat-all or treat-none strategies across clinically plausible risk thresholds. The study is limited by its fabricated dataset, retrospective simulation design, and restriction to two chronic conditions. Nevertheless, it demonstrates an empirically coherent framework for combining predictive accuracy, calibration, explainability, and clinical utility assessment. The findings support the feasibility of transparent long-term risk prediction models for chronic disease management.
Keywords: Explainable machine learning, Chronic disease management, SHAP, Long-term outcomes, Electronic health records, Clinical prediction
How to cite this article:
Citation Formats:
Contact Meral
Meral Publications
www.meralpublisher.com
Davutpasa / Zeytinburnu 34087
Istanbul
Turkey
Email: [email protected]