Peningkatan Akurasi Klasifikasi Risiko Serangan Jantung Menggunakan Feature Selection dan Random Forest dengan Pendekatan Cross Validation

  • Bima Aditya (Corresponding Author) Universitas Teknokrat Indonesia
  • Erliyan Redi Susanto Universitas Teknokrat Indonesia
Keywords: Random Forest, SMOTE, Feature Selection, Cross-Validation, Klasifikasi Risiko Jantung

Abstract

Penyakit kardiovaskular, khususnya serangan jantung, menyumbang beban morbiditas dan mortalitas tertinggi di Indonesia, namun pengembangan model prediktif berbasis machine learning untuk skrining dini menghadapi tantangan fundamental berupa ketidakseimbangan kelas (class imbalance), dimensionalitas fitur tinggi, serta risiko overfitting pada evaluasi konvensional. Penelitian ini bertujuan membangun pipeline klasifikasi yang sistematis dengan menangani ketidakseimbangan data, mereduksi redundansi fitur, dan mengevaluasi generalisasi model secara robust. Dataset Heart Attack Prediction Indonesia terdiri dari 158.355 sampel dengan 27 fitur prediktor. Tahapan metodologi meliputi prapemrosesan data, penyeimbangan kelas menggunakan Synthetic Minority Over-sampling Technique (SMOTE) kustom dengan k-nearest neighbors = 5, seleksi fitur dua-tahap berbasis Mutual Information dan Random Forest Feature Importance dengan SelectFromModel, serta pelatihan model Random Forest yang dioptimasi melalui penyetelan hyperparameter. Evaluasi performa dilakukan dengan Stratified 5-Fold Cross-Validation menggunakan metrik accuracy, precision, recall,F1-score, dan ROC-AUC. Hasil menunjukkan bahwa SMOTE berhasil menyeimbangkan distribusi kelas menjadi 94.854 sampel per kelas (rasio 1:1), sementara seleksi fitur mereduksi 27 fitur menjadi 13 fitur terpilih yang klinis relevan. Evaluasi pada data full resampled menghasilkan performa mendekati sempurna (ROC-AUC 0,9997), namun Stratified 5-Fold Cross-Validation mengungkapkan estimasi yang lebih realistis dengan rata-rata accuracy 0,7758, precision 0,7943, recall 0,7444, F1-score 0,7685, dan ROC-AUC 0,8721. Temuan ini mengidentifikasi generalization gap sebesar ~21,5 poin persentase yang mengindikasikan adanya data leakage struktural akibat implementasi SMOTE sebelum validasi silang. Kontribusi utama penelitian ini terletak pada demonstrasi pipeline end-to-end yang mengintegrasikan teknik resampling,seleksi fitur berbasis domain, dan evaluasi robust, sekaligus memberikan argumentasi kritis mengenai interpretasi metrik pada data sintetis dan perlunya estimasi performa berbasis cross-validation sebagai acuan klinis yang lebih valid.

Downloads

Download data is not yet available.

References

W. S. P. Harmadha et al., “Explaining the increase of incidence and mortality from cardiovascular disease in Indonesia: A global burden of disease study analysis (2000–2019),” PLoS One, vol. 18, no. 12, hal. e0294128, Des 2023, doi: 10.1371/journal.pone.0294128.

T. Liu, A. J. Krentz, Z. Huo, dan V. Ćurčin, “Opportunities and Challenges of Cardiovascular Disease Risk Prediction for Primary Prevention Using Machine Learning and Electronic Health Records: A Systematic Review,” Rev. Cardiovasc. Med., vol. 26, no. 4, Apr 2025, doi: 10.31083/RCM37443.

G. C. M. Siontis, I. Tzoulaki, K. C. Siontis, dan J. P. A. Ioannidis, “Comparisons of established risk prediction models for cardiovascular disease: systematic review,” BMJ, vol. 344, no. may24 1, hal. e3318–e3318, Mei 2012, doi: 10.1136/bmj.e3318.

C. F. Chen, Se-Hyun Yang, B. Falsafi, dan A. Moshovos, “Accurate and Complexity-Effective Spatial Pattern Prediction,” in 10th International Symposium on High Performance Computer Architecture (HPCA’04), IEEE, hal. 276–276. doi: 10.1109/HPCA.2004.10010.

G. Ramaswami, T. Susnjak, dan A. Mathrani, “On Developing Generic Models for Predicting Student Outcomes in Educational Data Mining,” Big Data Cogn. Comput., vol. 6, no. 1, hal. 6, Jan 2022, doi: 10.3390/bdcc6010006.

A. Özçift, “Random forests ensemble classifier trained with data resampling strategy to improve cardiac arrhythmia diagnosis,” Comput. Biol. Med., vol. 41, no. 5, hal. 265–271, Mei 2011, doi: 10.1016/j.compbiomed.2011.03.001.

J. Ramesh, R. Aburukba, dan A. Sagahyroon, “A remote healthcare monitoring framework for diabetes prediction using machine learning,” Healthc. Technol. Lett., vol. 8, no. 3, hal. 45–57, Jun 2021, doi: 10.1049/htl2.12010.

P. Probst, M. N. Wright, dan A. Boulesteix, “Hyperparameters and tuning strategies for random forest,” WIREs Data Min. Knowl. Discov., vol. 9, no. 3, Mei 2019, doi: 10.1002/widm.1301.

V. Srinivasan et al., “Optimizing pipelines for power and performance,” in 35th Annual IEEE/ACM International Symposium on Microarchitecture, 2002. (MICRO-35). Proceedings., IEEE Comput. Soc, hal. 333–344. doi: 10.1109/MICRO.2002.1176261.

J. V. Alonso dan L. Escot, “Robust Cross-Validation of Predictive Models Used in Credit Default Risk,” Appl. Sci., vol. 15, no. 10, hal. 5495, Mei 2025, doi: 10.3390/app15105495.

V. C. Pezoulas et al., “A computational pipeline for data augmentation towards the improvement of disease classification and risk stratification models: A case study in two clinical domains,” Comput. Biol. Med., vol. 134, hal. 104520, Jul 2021, doi: 10.1016/j.compbiomed.2021.104520.

A. H. Syed, T. Khan, dan N. Alromema, “A Hybrid Feature Selection Approach to Screen a Novel Set of Blood Biomarkers for Early COVID-19 Mortality Prediction,” Diagnostics, vol. 12, no. 7, hal. 1604, Jun 2022, doi: 10.3390/diagnostics12071604.

I. Aruleba dan Y. Sun, “An Improved Ensemble Method With Data Resampling for Credit Risk Prediction,” IEEE Access, vol. 13, hal. 71275–71287, 2025, doi: 10.1109/ACCESS.2025.3563432.

M. Amelia dan D. Fitriyani, “Classification of Thyroid Disease Risk Using the XGBoost Method,” J. Intell. Syst. Technol. Informatics, vol. 1, no. 3, hal. 93–97, Nov 2025, doi: 10.64878/jistics.v1i3.26.

A. S. Deokar dan M. A. Pradhan, “Optimizing Cardiovascular Disease Detection Using Ranking-Based Feature Selection Machine Learning Models,” Eng. Technol. Appl. Sci. Res., vol. 15, no. 5, hal. 28172–28178, Okt 2025, doi: 10.48084/etasr.11923.

C. T. Nakas, T. A. Alonzo, dan C. T. Yiannoutsos, “Accuracy and cut‐off point selection in three‐class classification problems using a generalization of the Youden index,” Stat. Med., vol. 29, no. 28, hal. 2946–2955, Des 2010, doi: 10.1002/sim.4044.

S. A. Thamrin, D. S. Arsyad, H. Kuswanto, A. Lawi, dan S. Nasir, “Predicting Obesity in Adults Using Machine Learning Techniques: An Analysis of Indonesian Basic Health Research 2018,” Front. Nutr., vol. 8, Jun 2021, doi: 10.3389/fnut.2021.669155.

I. M. Alkhawaldeh, I. Albalkhi, dan A. J. Naswhan, “Challenges and limitations of synthetic minority oversampling techniques in machine learning,” World J. Methodol., vol. 13, no. 5, hal. 373–378, Des 2023, doi: 10.5662/wjm.v13.i5.373.

P. Cunningham dan S. J. Delany, “k-Nearest Neighbour Classifiers - A Tutorial,” ACM Comput. Surv., vol. 54, no. 6, hal. 1–25, Jul 2022, doi: 10.1145/3459665.

A. Sufyan, S. Mohsin, K. Hameed, E. Az- Zo’bi, dan M. Tashtoush, “Discretization of the Inverse Rayleigh-G Family: Theoretical Properties, Machine Learning-Based Parameter Estimation, and Practical Applications,” Stat. Optim. Inf. Comput., vol. 14, no. 3, hal. 1174–1197, Jul 2025, doi: 10.19139/soic-2310-5070-2618.

D. Chutia, D. K. Bhattacharyya, J. Sarma, dan P. N. L. Raju, “An effective ensemble classification framework using random forests and a correlation based feature selection technique,” Trans. GIS, vol. 21, no. 6, hal. 1165–1178, Des 2017, doi: 10.1111/tgis.12268.

D. Rengasamy et al., “Feature importance in machine learning models: A fuzzy information fusion approach,” Neurocomputing, vol. 511, hal. 163–174, Okt 2022, doi: 10.1016/j.neucom.2022.09.053.

D. Wilimitis dan C. G. Walsh, “Practical Considerations and Applied Examples of Cross-Validation for Model Development and Evaluation in Health Care: Tutorial,” JMIR AI, vol. 2, hal. e49023, Des 2023, doi: 10.2196/49023.

S. kamal Utsho, “The Decline of Synthetic Oversampling in Large-Scale Imbalanced Learning:A Post-SMOTE Empirical and Theoretical Study (2020–2025),” 2 Desember 2025. doi: 10.21203/rs.3.rs-8236211/v1.

S. Gholampour, “Impact of Nature of Medical Data on Machine and Deep Learning for Imbalanced Datasets: Clinical Validity of SMOTE Is Questionable,” Mach. Learn. Knowl. Extr., vol. 6, no. 2, hal. 827–841, Apr 2024, doi: 10.3390/make6020039.

D. P. Mitchell, “Generating antialiased images at low sampling densities,” in Proceedings of the 14th annual conference on Computer graphics and interactive techniques, New York, NY, USA: ACM, Agu 1987, hal. 65–72. doi: 10.1145/37401.37410.

I. Dey dan V. Pratap, “A Comparative Study of SMOTE, Borderline-SMOTE, and ADASYN Oversampling Techniques using Different Classifiers,” in 2023 3rd International Conference on Smart Data Intelligence (ICSMDI), IEEE, Mar 2023, hal. 294–302. doi: 10.1109/ICSMDI57622.2023.00060.

P.-Y. Wong et al., “Effects of feature selection methods in estimating SO2 concentration variations using machine learning and stacking ensemble approach,” Environ. Technol. Innov., vol. 37, hal. 103996, Feb 2025, doi: 10.1016/j.eti.2024.103996.

P. Dubey, P. Dubey, dan P. N. Bokoro, “Advancing CVD Risk Prediction with Transformer Architectures and Statistical Risk Factor Filtering,” Technologies, vol. 13, no. 5, hal. 201, Mei 2025, doi: 10.3390/technologies13050201.

A. Louati dan H. Louati, “Predictive Monitoring of Wage-Band Classification in GOSI Data with Leakage Control and Out-of-Time Validation,” Forecasting, vol. 8, no. 2, hal. 27, Mar 2026, doi: 10.3390/forecast8020027.

Published
2026-09-21
How to Cite
Aditya, B., & Susanto, E. R. (2026). Peningkatan Akurasi Klasifikasi Risiko Serangan Jantung Menggunakan Feature Selection dan Random Forest dengan Pendekatan Cross Validation. Journal of Artificial Intelligence and Technology Information (JAITI), 4(3), 427-437. https://doi.org/10.58602/jaiti.v4i3.289