Efficient Hospital Readmission Prediction with Reduced Feature Sets Under Class Imbalance  
Author

Kehan Gao

 

Co-Author(s)

Fatma Pakdil;  Steve Muchiri; Garrett Dancik; Ece Pakdil; H. Nail Aydin

 

Abstract Hospital readmissions remain a major challenge for healthcare systems due to their impact on patient outcomes and healthcare costs. Accurate prediction of high-risk patients can support targeted interventions and reduce avoidable readmissions. This study investigates hospital readmission prediction using machine learning models under multiple feature selection configurations. Three classifiers (CatBoost, Logistic Regression, and Random Forest) were evaluated using four feature sets: Full, No-LOS (excluding length of stay), Top-10, and Top-19. SHAP-based feature importance was used to identify reduced and
interpretable feature subsets. Performance was assessed using 5-fold cross-validation with precision, recall, F1-score, ROC-AUC, and PR-AUC.
Experimental results show that CatBoost consistently achieved the best overall performance. The Top-19 feature set preserved nearly all predictive power of the Full model while reducing feature dimensionality by approximately 50%. In contrast, aggressive feature reduction (Top-10) and removing length of stay
(LOS) resulted in statistically significant performance degradation. Paired t-tests further confirmed that CatBoost significantly outperformed Logistic Regression and Random Forest. These findings demonstrate that interpretable feature reduction combined with gradient boosting provides an efficient and effective approach for readmission prediction.

 

Keywords hospital readmission prediction, machine learning, CatBoost, feature selection, SHAP, class imbalance, healthcare analytics, predictive modeling
   
    Article #:  RQD2026-21
 

Proceedings of 31st ISSAT International Conference on Reliability & Quality in Design
August 5-7, 2026