Predicting survival rates of patients with cardiovascular diseases using ensemble techniques
* Author to whom correspondence should be addressed.
Sigma Journal of Engineering and Natural Sciences 2026, Vol. 44, Issue 1, pp. 404-422; doi.org/10.14744/sigma.2026.1989
Abstract
Keywords: Cardiovascular Disease; Kaplan-Meier Estimator; Machine Learning; Survival Rate
Introduction
Heart and blood vessel problems like heart attacks, strokes, and heart failure are known as cardiovascular diseases (CVDs). Around 17 million people die from these diseases every year worldwide [1]. In India, the number of deaths due to CVDs has been increasing for the first time
in 30 years. Heart failure occurs when cardiovascular system cannot circulate blood throughout the human body. Excessive hypertension, diabetic complications, or other cardiac problems frequently bring it on. Doctors classify heart failure into two types based on how much blood the heart pumps out with each beat [2]. One type, heart failure with reduced ejection fraction (HFrEF) [3], happens when
*Corresponding author. *E-mail address: kararebharati@gmail.com This paper was recommended for publication in revised form by Editor-in-Chief Ahmet Selim Dalkilic Published by Yıldız Technical University Press, İstanbul, Turkey © Author. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
the heart pumps less than 45% of the blood. The other type, heart failure with preserved ejection fraction (HFpEF), occurs when the heart contracts well but doesn’t relax properly to fill with blood. Most heart attacks can be prevented by using population-wide strategies to address lifestyle risk factors such as smoking, unhealthy diets and obesity, and alcohol consumption [4]. The machine learning is an effective strategy to early prediction and proper medication of patient with heart failure or who are at risk due to the combination of several risk factors such as BP, diabetics, and other medical history [5]. This potential brings optimism for the future of cardiovascular disease management. Machine learning (ML) might surpass current modeling methods to accurately predict high blood pressure within specific racial groups and elucidate critical factors contributing to high blood pressure development across diverse races [6]. The majority of heart failure cases can be attributed to structural or physiological issues in the heart, leading to increased intracardiac pressure or reduced cardiovascular output based on the individual’s state of rest or stress [7]. Consequently, heart failure is associated with a diminished quality of life and decreased engagement in physical and mental activities. Approximately 1-2% of the general population and 10% of older people in developed nations are affected by heart failure, with its prevalence expected to rise alongside an aging population. In hospital discharge, patients with heart failure (HF) experience a high 56.6% readmission rate. Addressing high frequency promptly is crucial to prevent future severe complications, with a current urgent focus on minimizing readmissions. Cardiovascular diseases like coronary artery disease (CAD), atrial fibrillation (AF), and vascular conditions remain the primary global cause of death [8]. The rising incidence of cardiovascular diseases is a pressing issue that requires immediate attention and innovative solutions. If lifestyle conditions rise and the amount of stress increases, the incidence of cardiovascular diseases is alarmingly increasing. To address this problem, an ensemble approach and survival rate prediction of patients have developed as a possible treatment. Motivation Based on the current studies [9]-[10] cardiovascular disease (CVD) is projected to cause the deaths of approximately 23 million individuals by 2030. There are many causes such as heart disease, irregular heartbeat, and heart attack, which are three different forms of cardiovascular disease [11]-[12]. Age, gender, BMI, height, waist circumference, and results from blood tests that check cholesterol levels, liver health, and kidney function are some of the factors used to analysed cardiovascular disease [13]–[14]. Several health issues could arise from the complex interactions across the risk factors. Traditional statistically effective methods are unable to investigate the complex relationship across risk-associated factors due to the large number of components present [15]–[16].
Over the last few decades, several researchers utilized the artificial intelligence (AI) techniques to investigate the new clinical data that help the physicians to analyze the signs and effects of many diseases and the prediction of survival of patients. The continuous efforts to collect all health examination records and consistent clinical data [17]–[18] focus to standardizing of clinical data before investigating previously unknown risk factors. A number of possible risk factors shows the associations regarding the development of diseases that indicate the basic causes of the disorders. Furthermore, a significant amount of medical data must be analyzed in order to create accurate prediction models for disease occurrences [19]–[20]. The risk assessment of CVD frameworks increasingly utilizes AI and large amounts of clinical data aggregation. Problem Statement Cardiovascular diseases (CVD) are the leading cause of death globally, with an estimated 17.9 million deaths each year. Early prediction of survival rates for patients with CVDs is crucial for providing timely and effective interventions to improve patient outcomes. Machine learning (ML) algorithms have shown promise in predicting survival rates based on patient data, including demographics, medical history, and clinical measurements. This study designs the ensemble models to detect and classify CVD and also predict the survival rates of patients with cardiovascular diseases. The models trained on a CVD dataset contain the risk factors of CVD. To measure the performance of the models using several evaluation parameters and also predict the survival rate of patients over a specified period. The proposed model helps to improved patient care that enable the healthcare experts to identify high-risk patients and early diagnosis to prevent adverse effect. Limitation of Existing System • The existing model worked on limited set of features, such as age, gender, and a few clinical elements. • The quality and quantity of dataset is used to train the models could vary performance significantly that shows the biased or inaccurate predictions. • Some existing models might have needed to be more complex, making them difficult to interpret and implement in clinical settings. This need for interpretability is a pressing issue in healthcare machine learning. We must address this complexity to avoid overfitting, where the model performed well on training data but failed to generalize to new data. Machine learning models worked effectively on present clinical procedure in order to be useful in the field of medicine. The medical professionals’ resistance and accessibility issues result that many existing methods were not designed with the combination in perspective. Contribution Following is a list of this investigation’s primary contribution as well as uniqueness:
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
A technique for increasing the accuracy of cardiovascular disease prediction has been developed by Mohan et al. [28], employing machine learning algorithms for recognizing critical features. Regarding heart disease prediction, the suggested hybrid RF and linear model obtained the 87.9% accuracy. Au et al. [29] suggested the hybrid model to detect the CVD based on the logistic regression and obtained the accuracy score of 88.00%. Investigators have suggested a hybrid approach for forecasting cardiovascular disease [30]. The framework was employed by three ML techniques: DT, RF, and a combination. At 87.8%, the hybrid technique had the most excellent accuracy score. As reported [31], author suggested the several ML models used to detect CVD and also suggested the ML-based ensemble model to predict and improve the efficiency of model. We use information on coronary artery disease for our investigation. So, this might be our final investigation, with about 710,000 people and ten features in the data set. In addition, we employ multiple deep learning and machine learning techniques to determine which is most effective in identifying coronary artery disease. Last two decade there is an advancements and large use of ML based models to predict the CVD and there are still several gaps that are unaddressed. The existing models depends on limited feature sets that fail to capture the complex features among risk factors such as lifestyle, biochemical markers, and physiological features. The usefulness of ensemble approaches in real-world scenarios is limited by the lack of thorough evaluations in several publications. The majority of approaches do not offer apparent findings for medical decision-making, making and ML models difficult to interpret in healthcare settings. Additionally, the datasets utilized in earlier studies are either limited or region-specific, and also limits their applicability to larger populations. Practical interactions between ML and survival analysis techniques, including Kaplan-Meier estimators, are not fully investigated. Addressing these gaps could lead to more robust, interpretable, and clinically actionable models for predicting survival rates in cardiovascular patients.
To extract seven unique features from the CVD dataset. Following the feature extraction, we meticulously normalized the data and divided the CVD dataset into training and testing sets using a 70:30 split, a crucial step in our thorough methodology. This aids in creating an ensemble model that includes a base ML classifier and a DT as a meta-classifier. Utilizes the Kaplan-Meier estimator to predict the survival rate of cardiovascular diseases for continuous variables in the dataset. Ultimately, we obtained results through various performance parameter analyses and predicted the survival rate of patients.
Paper organization Section 2 presents the previous work in cardiovascular disease prediction based on ML and DL techniques. Section 3 presents the proposed methodology; Section 4 discusses the result analysis of the ensemble model. Section 4 addressed the conclusion and future scope of the study.
40. On the other hand, the maximum age varies between
genders, with males reaching up to 95, while females have a maximum age of 90. These age specifications provide insights into the age range covered in the research, indicating that individuals below 40 are omitted, and the maximum ages differ slightly for males and females in the studied cohort. The impact of patient aging on the probability of survival is depicted in Figure 5. The 50–70 age category has
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
a significantly greater survival probability than the population. Every age category still carries some chance of not suffering a cardiovascular occurrence; the risk is most significant in the range of 60-65 age. Beyond the age of 80, the chances of survival sharply decline. The patterns suggest that age plays a significant role in survival outcomes after a heart failure event. Figure 6 shows the distribution of various parameters and their survival rate. It is observed that survival outcomes in the studied population are based on gender. For the male population, 44.1% (132 individuals) have survived heart failure, while 20.7% (62 individuals) unfortunately did not survive. In the female population, 23.7% (71 individuals) survived the heart failure event, and 11.4% (34) did not.
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Figure 6. Distribution of parameters and their survival rate.
These percentages highlight the gender-specific variations in survival rates after experiencing heart failure, indicating a higher survival rate for males compared to females in the studied cohort. Around the 35.00% develop hypertension that excessive blood pressure. With this subgroup, 22% have sudden cardiac arrest events, and 13% unfortunately did not make it. Turning to the 65% of the population without hypertension, 45.8% have successfully survived heart failure, underscoring a higher survival rate compared to the hypertensive group, where 19% succumbed to the condition. Around 42% of individuals are identified as having diabetes, while approximately 58% do not have diabetes. 28.4% of people with diabetes have survived a cardiovascular attack, but unfortunately, 13.4% did not survive. In contrast, among those without diabetes, a higher percentage, precisely 39.5%, have successfully survived after a heart attack, even though 18.7% have unfortunately given
in to the medical condition. 32% of individuals have smoking habits, while around 68% do not smoke. Among those who smoke, 22.1% of people have prevented cardiac arrest, and unfortunately, 10% did not survive. On the other hand, among individuals without smoking habits, a higher percentage, precisely 45.8%, have successfully survived after cardiac arrest, while 22.1% have unfortunately given an approach to the illness. 43.1% of individuals exhibit symptoms of anemia, while around 56.9% do not show any signs of anemia. Among those with anemia, 27.8% survived after cardiac arrest, while 15.4%, unfortunately, failed to survive. Conversely, among individuals without anemia, a higher percentage, precisely 40.1%, have successfully survived after cardiac arrest, while 16.7% have unfortunately succumbed to the condition. 46.5% of the entire population has lower blood sodium levels, while 53.5% have levels within the acceptable range. Among those with low
sodium levels, 26.8% after cardiac arrest, while 19.7% did not survive. In contrast, some patients whose sodium levels are below the appropriate limit have a higher percentage, precisely 41.1%, after cardiac arrest, with a smaller number, 12.4%, succumbing to the condition. These findings suggest a potential correlation between blood sodium levels and heart failure outcomes, with individuals within the acceptable range demonstrating a higher survival rate than those with lower sodium levels in the studied population. Figure 7 shows information on various factors and laboratory test results related to heart failure outcomes in the
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
studied population. First, it notes that people with cardiac failure who failed to survive generally have higher levels of CPK (creatine phosphokinase) enzyme. The violin histogram indicates outliers with high CPK levels in survival and death events. Additionally, it mentions that individuals who passed away with cardiac arrest often exhibited less than typical values for the ejection proportion, indicating inadequate pumping of blood from the heart. This shows outliers that observed in patients who survived heart failure. Serum creatinine levels 72.9% of cases reported elevated levels. The 43.8% are survived, and 29.1% succumbed to heart
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Figure 7. Distribution of parameters and their effect on survival rate. failure. Cases with serum levels in the normal range show a higher survival rate of 24.1%, contrasting with a lower percentage of 3.01% succumbing to the condition. However, 96 instances have passed away from cardiovascular disease. Of those instances, 59 people had sodium values below the normal range. Table 1 presents the most correlated values associated with death in the dataset. The age shows the positive correlation of 0.253729 that indicate modest relationship across increase age and the likelihood of death. On the other hand, ejection fraction shows a negative correlation of -0.268603, suggesting that a lower ejection fraction, representing the proportion of blood that the coronary artery
pumps continuously per contraction, is associated with a higher likelihood of death. Serum creatinine exhibits a positive correlation of 0.294278, indicating that higher serum
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
creatinine levels, a marker of kidney function, are connected to a higher chance of passing away. Build Baseline Classifiers The CVD detection using ML based classifiers such as LR, RF, XG Boost, and NB serve as initial models to measure the performance of baseline models. The RF is a simple linear model measure the probabilities. The RF is a group of decision trees that collect correlations between the features. The XG Boost algorithm known for its accuracy and efficiency; and NB for feature independence. The classifiers provide a starting point for evaluating more sophisticated models, allow us to compare their performance against these more straightforward approaches, and measure the effectiveness of their ML classifiers in predicting cardiovascular diseases. Ensemble Model Combine the predictions of multiple individual classifiers to improve the overall prediction accuracy. We designed ensemble classifiers for CVD prediction using ML baseline classifiers such as LR, RF, XG Boost, and NB. Ensemble classifiers are often more robust and accurate than individual classifiers that effectively used the strengths of different models and mitigate their weaknesses, ultimately improving the performance of CVD prediction models [37]. The suggested ensemble models
develops the outcomes using the weighted majority to combine the predicted results of many ML models. The most favorable results are shown following the fine-tuning of each categorization model. Equation 1 show to obtained the maximum votes. (1) Normalize and specify the loss function that provide an objective function to measure the performance of the ensemble model (2) Equation 2 shows the variables measured using the plus operator and derived from the supplied inputs. The normalization factor (Theta′) and the loss function used to calculate the model generalization. Figure 9 shows the suggested ML-based ensemble model for predicting CVD. This model uses a voting approach by allowing each model to “vote” on the final prediction. Each model independently makes its prediction, and the most common prediction among the models is selected as the final. By combining these models, the ensemble can achieve higher accuracy and robustness in predicting CVD than any single model.
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Pseudo Code: Ensemble Model Input: Training Dataset CVD = {(p1, q1), (p2, q2),...(pn, qn)} Baseline Model M = (LR, RF, XGB, NB) Meta Classifier DT
Table 2. Hyperparameter settings of ML models in ensemble Model
Output: Learn Ensemble Model EM Start Step-1: Learn the Baseline classifiers M on CVD for i = 1 to n do
Bi = Mi(CVD) end for Step-2: Construct new CVD Dataset for prediction CVD' for j=1 to n' do for i = 1 to n do Used Bi to classify training parameters pj
xij = Bi(pj) end for CVD = (xj, qj), where xj = {xij, x2j,.... xnj} end for Step-3: Learn Meta Classifier DT EM = DT(CVD) Return EM END
Hyperparameter settings Hyperparameter optimization plays a crucial role in improving the performance of ensemble models for cardiovascular disease detection by fine-tuning the parameters of individual ML models included in the ensemble. Table 3 shows the hyperparameter setting of ML models used to build the ensemble model
Results And Discussion
The following section examines the baseline and proposed ensemble classifier performance and briefly overviews the outcomes. The main goal of this investigation is to investigate the effectiveness of suggested ML algorithms for identifying cardiovascular illness. We used the CVD dataset in the tests conducted for this study. We divided the CVD dataset into 70% training and 30% testing. We
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Table 3. Comparative result analysis of baseline model with ensemble model Classifiers
16. GB RAM, and 4 GeForce RTX graphics cards. Several
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
percentage of actual positive instances the model accurately recognized. A higher recall indicates better performance in identifying positive cases. The ensemble model obtained the recall scores of training and testing of 86.90% and 82.53% that shows the superior ability to identify positive cases.
Table 2 presents a comparative analysis of baseline models, including Logistic Regression, Random Forests, XG Boost, Naive Bayes, and an Ensemble Model, based on their performance metrics. The finding shows the ensemble model performed well ad compared to individual models and obtained the accuracy of 82.00%, precision of 85.00%, Recall of 80.00%, and F1-Score of 82.00%. The suggested model is more effective due to ability to combine the strengths of different ML classifiers. Table 4 shows different machine learning models’ training and test recall scores for cardiovascular disease prediction. Recall, sometimes called sensitivity, indicates the
Survival Prediction Using Kaplan Meier This study uses the Kaplan-Meier estimator to predict the survival rate of cardiovascular diseases for continuous variables, such as Age, Creatinine Phosphokinase, Ejection Fraction, Platelets, Serum Creatinine, Serum Sodium, and Time. The first step is to discretize each continuous variable into intervals. The proportion of observations that survive beyond each time interval is calculated for each combination of intervals. This is done similarly to the standard Kaplan-Meier estimator but using the intervals for the continuous variables. Figure 11 is a powerful tool that visually represents the estimated survival probabilities against each interval of parameters, allowing us to quickly grasp the survival function over the range of the continuous variables. Another name for the Kaplan-Meier estimate is “product limit estimate.” This mesure the possibility of an event occurred at a specific moment [38,39]. Combine this sequential probability with all previously estimated possibilities to obtain the highest possible estimation.
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
(7) The chance of surviving risk for every period is determined by dividing the total number of vulnerable people by the total number of patients who survive. Individuals who have passed away, stopped participating, or moved away are not included in the denominator and are not regarded as “under threat. p Values and Statistical Significance of Parameters The summary’s p-values shows risk exhibit considerable significance, with p-values below 0.0005 that shows statistical significance at a confidence level of 99.9995% or higher. The attributes demonstrate a strong correlation with the occurrence of the death event. Conversely, Smoking, Sex, Platelets, and Diabetes yield notably high p-values, making it uncertain whether they hold statistical significance. Consequently, their impact on the hazard rate may be disregarded in the analysis. This observation is supported by their substantial standard errors and the consequent broad confidence intervals [40]. Anemia is a binary variable represented by 1 or 0, indicating whether the subject is anemic. The coefficient of 0.481 is interpreted as:
The hazard ratio for anemia is 1.618. This indicate that the patient has anemia, the risk of death increases by 61.8%.
High BP represented by 1 and 0 that shows that subject has high blood pressure (hypertension) or not. The coefficient of 0.406 is interpreted as:
The HR for high BP is 1.5. This shows that the patient has hypertension. The risk of death increases by 50%. Hazard ratio (HR) The HR shows the impact of a covariate on the hazard rate. An HR of 1 suggests no effect that indicate the covariate does not influence the hazard rate. An HR more significant than 1 increases the hazard rate as the covariate value increased that indicate the higher risk of the event occurring [40]. The HR less than 1 shows to decrease in the hazard rate as the covariate value increases that suggest lower risk of the event. Figure 12 represents the coefficients (i.e., log hazard ratios) for predicting the survival rate of cardiovascular diseases based on various factors. Each factor, such as anemia, high blood pressure, serum creatinine, diabetes, smoking, age, creatinine phosphokinase, platelets, serum sodium, ejection fraction, and sex, is listed along the vertical axis. The horizontal bars extending from the central axis represent the magnitude and direction of the coefficients. A bar extending to the right indicates a positive effect on the hazard ratio, meaning an increased risk of cardiovascular disease survival, while a bar extending to the left indicates a negative impact, implying a decreased risk.
Figure 12. Graphical Visualization of the Coefficients (i.e. log hazard ratios).
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
The impact of covariates on survival outcomes has been understood through the HR which measures the change in hazard rate with a one-unit change in the covariate. For Ejection Fraction, a one-unit change results in a 5.2% increase in survival time, as indicated by an HR of 0.95. Conversely, a one-unit rise in Serum Creatinine causes a 28.1% decrease in survival time, given its HR of 1.392 [41]. Age shows a 3.9% decrease in survival time per unit increase, with an HR of 1.046. However, Creatinine Phosphokinase and Platelets have HRs of 1, suggesting no effect on the probability of the death occurrences. Figures 12 and 13 provide the graphical layout of coefficients, log hazard ratios, and hazard ratios, respectively, encompassing their sizes and standardized errors. Anemia,
High Blood Pressure, Serum Creatinine, Age, and the amount of Ejection are all within the 95% confidence interval of influencing the death incident. Figure 14 shows that with age, the survival probability decreases for any complication arising from a heart failure condition. Increasing age has a significant impact on the likelihood of surviving. Figure 15 shows that the volume of blood pumped out of the heart increases with increasing ejection fraction percentage. Consequently, the probability of survival also increases during any phase of heart failure. Increasing EF levels has a beneficial significant impact on the likelihood of survival. Figure 16 shows that with increasing serum creatinine levels in the blood, the survival probability decreases for
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Figure 16. Partial effect of varied serum creatinine vs non-anaemia.
any complication arising out of a heart failure condition. Increasing creatinine levels have a significant impact on the likelihood of survival. Figure 17 shows that anaemic patients are more likely to encounter a hazard due to heart failure condition—survival Probabilities over 280 days. Table 5 analyzes the survival probabilities of the patient in the test cohort over 280 days. The analysis assumes that the subjects have just entered the study without considering how long they have lived. It is observed that initially, each subject in the test cohort has high survival chances, hovering around the 98-99% mark. For Patient -298 and Patient -179, the survival probabilities remain consistent throughout the period, at 88% and 85.5%, respectively, by the 280th day. For Patient-42 and Patient -193, the survival probabilities hover around 61% and 34%, respectively, at the end of 280 days. However, for patient -5, the chances of survival show a decreasing trend. By day 15, the survival chance is
approximately 75%; by day 38, it hovers around the 50% mark; by the end of 180 days, it falls below 10%. Table 6 shows the comparative perfromance analysis of the proposed ensemble models and existing model to detect and classify the CVD. The proposed ensemble model obtained the accuracy of 82.53% that performed well as compared to existing individual ML models that obtained the accuracy score of DT has 81.23%, K-means of 78.00%, ANN of 82.10%, and Stacking Models of 82.35%. The XGBoost obtained the 82.44%, and SVM achieved the highest accuracy at 83.12%. This comparison demonstrates that the ensemble approach offers a balanced tradeoff between simplicity and performance, positioning it as a competitive alternative among state-of-the-art techniques for this task. The final results suggest whether the various CVD parameters are used to forecast a patient’s survival rate suffering heart disease. They also indicate that forecasts based
Table 5. Survival probabilities of patients last 10 days Days
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Table 6. Comparative analysis of proposed method with existing methods Author
Methods
simply on both variables may be more reliable than those based on the entire dataset. This is especially required for healthcare facilities environments: physicians might still be capable of forecasting the survival of patients using Age, smoking, serum creatinine amounts, and ejection fraction alone, even if the individual technological medical care contained a lot of not present clinical information and test examination from laboratories. However, invetigation that need for further study to design the strong models to daignosis the CVD ealier and implementable in healthcare settings. Further investigation offered several interesting results that were not discovered in the researchers’ earlier dataset analysis [41]. Ahmad et al. found that anemia, high blood pressure, ejection fraction, age, and serum creatinine a sign of kidney disease were actually the most common characteristics. The characteristics are important since are highly predictive of patient survival rates.
serum creatinine, and ejection fraction by bridging the gap between ML and medical survival analysis. This method offers useful applications for early treatment in cardiovascular care while also advancing predictive modeling. In order to enhance preventative and therapeutic approaches, future studies might investigate the connections among risk variables and CVD prognosis.
Conclusion
This study validated the importance of relevant feature extraction with machine learning by demonstrating that conventional statistical analysis identified cardiovascular disease factors as the most significant features. Furthermore, the proposed method showed that machine learning was applied successfully to the binary categorization of electronic healthcare of individuals with coronary artery disease disorders related to the heart system. This study demonstrates the effectiveness of machine learning techniques in predicting cardiovascular diseases and forecasting survival rates of patients. The ensemble model performed well as compared to baseline models that indicate its superiority in disease prediction. This study also identifies age, serum creatinine, and ejection fraction as significant factors affecting the death event, while smoking, sex, platelets, and diabetes may not hold statistical significance. The existing studies that focus on individual models or limited feature sets. The proposed model combined the diverse risk factors and shows superior accuracy compared to existing models. This work is unusual because it provides interpretable insights into important risk factors including age,
Acknowledgment
We acknowledge the support received from Yeshwantrao Chavan College of Engineering, Nagpur, Maharashtra, India. We are deeply grateful to the management and Principal Dr. U. P. Waghe for their invaluable support and encouragement.
Data Availability Statement
The authors confirm that the data that supports the findings of this study are available within the article. Raw data that support the finding of this study are available from the corresponding author, upon reasonable request.
Conflict Of Interest
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Sigma J Eng Nat Sci, Vol. 44, No. 1, pp. 404−422, February, 2026
Ethics
There are no ethical issues with the publication of this manuscript.
Statement On The Use Of Artificial Intelligence
Artificial intelligence was not used in the preparation of the article.
Share and Cite
KARARE, B.; PANDE, A.A.; WAGHALE, P.; PAUL, A.T.; TRIPATHI, C.; DAMAHE, L. Predicting survival rates of patients with cardiovascular diseases using ensemble techniques. Sigma Journal of Engineering and Natural Sciences 2026, Vol. 44, pp. 404-422. https://doi.org/10.14744/sigma.2026.1989

