Data subdivision approach enhances machine learning-based mortality prediction in pediatric ICU patients

Wenqian Chen, Benjamin Lee, Zexi Zang, Junfeng Li, Lingna Huang, Hang Xing, Yanli Ren

Abstract

To evaluate machine learning–based models for predicting all-cause mortality in pediatric ICU patients using comprehensive biochemical panels, with a focus on addressing missing data and class imbalance.

Introduction

Children admitted to pediatric intensive care unit (PICU) with critical illnesses are at high risk of clinical deterioration and subsequently increased risk of mortality [1]. Therefore, PICUs are designed to provide specialized critical care and continuous monitoring for severely ill children. Significant evidence has suggested that biochemical and clinical warning signs can help predict the risk of mortality in PICU [2,3].

Materials and methods

This study utilized a publicly available PICU dataset, consisting of 12,881 patient records collected between 2010 and 2019 at The Children’s Hospital of Zhejiang University School of Medicine.

Results

Of the collected data, the mean age of the overall cohort was 3.2 ± 3.8 years, with no significant difference observed between the survivor (3.2 ± 3.8 years) and non-survivor (3.0 ± 3.9 years) groups (p = 0.419).

Discussion

Mortality risk prediction in pediatric critical care remains a critical challenge [15]. Increased focus on ML-based risk prediction models in PICUs can be attributed to limitations of existing risk prediction tools. These generic severity scores were designed for population-level mortality assessment rather than guiding care for individual patients.

Conclusion

Our findings indicate that stacking model significantly mitigates the inherent class imbalance in the dataset, possibly by ensuring adequate representation of minority classes (mortality cases) during model training while preserving the diversity of majority-class samples, and improves model stability as a clinical prediction model.

Citation: Chen W, Lee B, Zang Z, Li J, Huang L, Xing H, et al. (2026) Data subdivision approach enhances machine learning-based mortality prediction in pediatric ICU patients. PLoS One 21(6): e0349772. https://doi.org/10.1371/journal.pone.0349772

Editor: Asadullah Shaikh, Najran University College of Computer Science and Information Systems, SAUDI ARABIA

Received: August 15, 2025; Accepted: May 5, 2026; Published: June 16, 2026
Copyright: © 2026 Chen et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Data Availability: The data used in this study are sourced from the publicly accessible Pediatric Intensive Care (PIC) database. Detailed information about this database can be found on the official website (http://pic.nbscn.org). All relevant data are within the paper and its Supporting information files.

Funding: This work was supported by the Joint Funds for the innovation of science and Technology, Fujian province (Grant number:2023Y9358) and the Fundamental Research Funds for the Central University (SCU2025D005). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.

Abbreviations: AUC, Area Under the Curve; AUROC, Area Under the Receiver Operating Characteristic Curve; AUPRC, Area Under the Precision-Recall Curve; ALT, Alanine Aminotransferase; AST, Aspartate Aminotransferase; EHR, Electronic Health Record; ICU, Intensive Care Unit; PICU, Pediatric Intensive Care Unit; INR, International Normalized Ratio; Lasso, Least Absolute Shrinkage and Selection Operator; ML, Machine Learning; PELOD, Pediatric Logistic Organ Dysfunction; PIM, Pediatric Index of Mortality; PRISM, Pediatric Risk of Mortality; PT, Prothrombin Time; PTT, Partial Thromboplastin Time; ROC, Receiver Operating Characteristic; SMOTE, Synthetic Minority Oversampling Technique; TT, Thrombin Time.