Perbandingan Algoritma Naive Bayes dan K-Nearest Neighbor untuk Prediksi Risiko Penyakit Stroke dengan Teknik Oversampling SMOTE
DOI:
https://doi.org/10.63643/jodens.v6i1.389Kata Kunci:
Early Warning System , K-Nearest Neighbor, Naive Bayes , SMOTE, StrokeAbstrak
Stroke is a leading cause of death and permanent disability in Indonesia, where prevalence rose from 8.3 per 1,000 population in 2013 to 10.9 per 1,000 in 2018, creating an urgent need for accurate early detection. This study develops and compares stroke risk prediction models based on Naive Bayes (NB) and K-Nearest Neighbor (KNN), each integrated with the Synthetic Minority Over-sampling Technique (SMOTE) to handle class imbalance. The dataset consists of 512 medical records from a hospital in Batam collected between 2020 and 2023, covering 12 demographic, clinical, and lifestyle predictors. SMOTE was applied only to the training data of an 80 to 20 split, yielding a balanced distribution of 258 samples per class. The KNN parameter K was tuned over eight odd values and 7 was found optimal. Evaluation on 103 test records shows that KNN consistently outperforms Naive Bayes, with accuracy of 91.3 against 81.6 percent, F1-Score of 89.2 against 79.9 percent, and AUC of 0.934 against 0.862. SMOTE raised recall of the Stroke class by 18.8 percentage points. Blood glucose, age, and systolic blood pressure were the three strongest predictors. The model is proposed as an Early Warning System module for hospital information systems.
Referensi
] WHO, "The top 10 causes of death," World Health Organization, 2023. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/the-top-10-causes-of-death
] Kemenkes RI, "Laporan Nasional RISKESDAS 2018," Badan Penelitian dan Pengembangan Kesehatan, Jakarta, 2018.
] E. Dritsas dan M. Trigka, "Stroke risk prediction with machine learning techniques," Sensors, vol. 22, no. 13, p. 4670, 2022. DOI: 10.3390/s22134670.
] N. L. Fitriyani, M. Syafrudin, G. Alfian, dan J. Rhee, "HDPM: An effective heart disease prediction model for a clinical decision support system," IEEE Access, vol. 8, pp. 133034–133050, 2020. DOI: 10.1109/ACCESS.2020.3010511.
] G. Sailasya dan G. L. A. Kumari, "Analyzing the performance of stroke prediction using ML classification algorithms," Int. J. Adv. Comput. Sci. Appl., vol. 12, no. 6, pp. 539–545, 2021.
] C. M. Bhatt, P. Patel, T. Ghetia, dan P. L. Mazzeo, "Effective heart disease prediction using machine learning techniques," Algorithms, vol. 16, no. 2, p. 88, 2023. DOI: 10.3390/a16020088.
] V. L. Feigin et al., "World Stroke Organization (WSO): Global stroke fact sheet 2022," Int. J. Stroke, vol. 17, no. 1, pp. 18–29, 2022. DOI: 10.1177/17474930211065917.
] GBD 2019 Stroke Collaborators, "Global, regional, and national burden of stroke and its risk factors, 1990–2019: A systematic analysis for the Global Burden of Disease Study 2019," Lancet Neurol., vol. 20, no. 10, pp. 795–820, 2021. DOI: 10.1016/S1474-4422(21)00252-0.
] G. Battineni, G. G. Sagaro, N. Chintalapudi, dan F. Amenta, "Applications of machine learning predictive models in the chronic disease diagnosis," J. Pers. Med., vol. 10, no. 2, p. 21, 2020. DOI: 10.3390/jpm10020021.
] L. J. Muhammad et al., "Predictive supervised machine learning models for diabetes mellitus," SN Comput. Sci., vol. 1, no. 5, p. 240, 2020. DOI: 10.1007/s42979-020-00250-8.
] P. Cunningham dan S. J. Delany, "k-Nearest neighbour classifiers – A tutorial," ACM Comput. Surv., vol. 54, no. 6, pp. 1–25, 2021. DOI: 10.1145/3459665.
] J. Gou, B. Yu, S. J. Maybank, dan D. Tao, "A generalized mean distance-based k-nearest neighbor classifier," IEEE Trans. Syst. Man Cybern. Syst., vol. 52, no. 5, pp. 2863–2878, 2022. DOI: 10.1109/TSMC.2020.3035829.
] N. V. Chawla, K. W. Bowyer, L. O. Hall, dan W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002. DOI: 10.1613/jair.953.
] T. Wongvorachan, S. He, dan O. Bulut, "A comparison of undersampling, oversampling, and SMOTE methods for dealing with imbalanced classification in educational data mining," Information, vol. 14, no. 1, p. 54, 2023. DOI: 10.3390/info14010054.
] N. L. Fitriyani, M. Syafrudin, G. Alfian, dan J. Rhee, "Development of diabetes prediction model using machine learning techniques and oversampling," J. Pers. Med., vol. 12, no. 9, p. 1497, 2022. DOI: 10.3390/jpm12091497.
] R. R. Sarra, A. M. Dinar, M. A. Mohammed, dan M. K. A. Ghani, "Enhanced prediction of heart disease using multi-modal fusion and explainability of ensemble machine learning," Diagnostics, vol. 12, no. 6, p. 1474, 2022. DOI: 10.3390/diagnostics12061474.
] D. Chicco dan G. Jurman, "Machine learning can predict survival of patients with heart failure from serum creatinine and ejection fraction alone," BMC Med. Inform. Decis. Mak., vol. 20, no. 1, p. 16, 2020. DOI: 10.1186/s12911-020-1023-5.
] O. Rainio, J. Teuho, dan R. Klén, "Evaluation metrics and statistical tests for machine learning," Sci. Rep., vol. 14, p. 6086, 2024. DOI: 10.1038/s41598-024-56706-x.
Unduhan
Diterbitkan
Cara Mengutip
Terbitan
Bagian
Lisensi
Hak Cipta (c) 2026 hilda herasmus, Agus Suryadi

Artikel ini berlisensi Creative Commons Attribution 4.0 International License.









