Md. Mahfuz Uddin,Md. Binyamin,Md. Muhtasim Munif Fahim,Md. Rezaul Karim
Abstract
Early and accurate prediction of cardiovascular disease (CVD) is fundamental for reducing morbidity and mortality. Machine learning (ML) algorithms provide a data-driven, actionable foundation to strengthen clinical decision-making and enable more precise risk stratification. The purpose of this study is to identify the most significant risk factors for CVD and to compare the predictive performance of eight machine learning algorithms.
Introduction
Cardiovascular disease (CVD) is a broad term that refers to a group of disorders affecting the heart and blood vessels, and it is the common cause of heart failure worldwide [1]. The term “CVDs” refers to a collection of illnesses that affect the heart and blood vessels, such as rheumatic heart disease, coronary heart disease, and cerebrovascular disease. An estimated 17.9 million people worldwide died from cardiovascular diseases (CVDs), making them the top cause of death worldwide.
Materials and methods
Data source and study variables
The dataset used in this study includes 68205 observations for 17 variables and is obtained from the publicly accessible Kaggle database. The main goal is to use a variety of patient parameters to predict whether cardiovascular disease will be present [12].
Results
Descriptive statistics
The descriptive statistics provided in Table 2 represent key insights into the dataset. Age (years) ranges from 39 to 64, with a mean of 52.89 years and a coefficient of variation (CV) of 12.75%, indicating moderate variability. BMI shows a moderate dispersion (CV = 15.94%) and an average of 26.94. With a mean of 126.30 mmHg, systolic blood pressure (mmHg) shows greater skewness (0.74), suggesting the presence of extreme values
Discussion
This study uses the publicly accessible Kaggle dataset to perform a thorough comparative analysis of eight machine learning algorithms for cardiovascular disease (CVD) prediction. This work offers several significant scientific and practical advancements, including identifying LightGBM as the top-performing model for cardiovascular disease (CVD) prediction in 5-fold cross-validation. To identify reliable risk factors for CVD, a comparison of four advanced feature selection techniques (Boruta, RRF, RFE, and LASSO) reveals the consistency and stability of predictor importance.
Conclusion
This study uses a publicly available dataset to compare eight machine learning algorithms for predicting cardiovascular disease. The results show that ensemble boosting techniques outperform conventional techniques in achieving balanced, reliable classification. Among the evaluated models, LightGBM demonstrates modestly better, more balanced predictive performance, achieving comparatively higher F1-score and AUC values across multiple cross-validation schemes.
Acknowledgments
The authors are thankful to the Kaggle website
Citation: Uddin MM, Binyamin M, Fahim MMM, Karim MR (2026) Comparative analysis of machine learning techniques for cardiovascular disease prediction. PLoS One 21(8): e0356170. https://doi.org/10.1371/journal.pone.0356170
Editor: Chong Liu, The University of Aukland, NEW ZEALAND
Received: March 18, 2026; Accepted: July 29, 2026; Published: August 18, 2026
Copyright: © 2026 Uddin et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The dataset used in this study is freely available on the Kaggle website: https://www.kaggle.com/datasets/colewelkins/cardiovascular-disease.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.