When Preprocessing Changes the Winner: Sensitivity of Medical Prediction Model Rankings to Missing Data Handling

Authors

  • Albahlool Abood Department of Software Engineering, Faculty of Information Technology, University of Gharyan, Gharyan, Libya

DOI:

https://doi.org/10.65417/ljcas.v4i2.386

Keywords:

Missing data; clinical prediction; model ranking stability; imputation; machine learning

Abstract

Missing data are common in clinical prediction studies, yet their handling may affect not only predictive performance but also which model is judged best. This study examined the stability of classifier selection when missing data handling was changed under controlled, paired evaluation conditions. Using SUPPORT2 data from 9,105 patients, Logistic Regression, Random Forest, and XGBoost were evaluated with median, K nearest neighbor, and iterative imputation. The analysis retained natural missingness, added nested missing completely at random (MCAR) perturbations of 5%, 10%, 20%, and 30%, and included a separate 20% missing at random (MAR) condition. Repeated stratified five fold cross validation used the same patient partitions, model seeds, and artificial missingness masks across corresponding comparisons. Ranking stability was assessed through condition level winner changes, paired rank reversals, and agreement across receiver operating characteristic area under the curve (ROC AUC), Average Precision, and Brier Score. The condition level winner remained stable under natural and mild additional missingness, but became dependent on imputation at higher missingness. Winner changes occurred in 6 of 12 imputation comparisons and 9 of 15 missingness comparisons, while pairwise rank reversals occurred in about 44% of matched repeat comparisons. All three metrics selected the same winner in 8 of 18 conditions. Because competing winners were separated by small ROC AUC margins, the results indicate sensitivity of model selection rather than large performance advantages. Reporting ranking stability alongside conventional performance estimates may therefore provide a more cautious basis for comparative clinical prediction studies.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-13

Issue

Section

Branch of Applied and Natural Sciences

How to Cite

Albahlool Abood. (2026). When Preprocessing Changes the Winner: Sensitivity of Medical Prediction Model Rankings to Missing Data Handling. Libyan Journal of Contemporary Academic Studies, 4(2), 77-94. https://doi.org/10.65417/ljcas.v4i2.386