A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction
DOI:
https://doi.org/10.33003/fjs-2026-1014-5650Keywords:
Class, Diabetes, Machine Learning, Prediction, Random ForestAbstract
Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise, creating an urgent need for scalable, low-cost tools for early risk identification. This study evaluates the effectiveness of five machine learning algorithms (Random Forest, XGBoost, Support Vector Machine [SVM], CatBoost, and TabNet) for predicting diabetes risk from routinely available clinical and lifestyle variables. Using the Pima Indians Diabetes dataset, a preprocessing pipeline was applied that included median imputation of physiologically implausible zero values, standardized (Z-score) feature scaling, Boruta-based feature selection, and class-imbalance handling through class-weight adjustment and the Synthetic Minority Oversampling Technique (SMOTE). Models were trained on an 80/20 train-test split and assessed using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (ROC-AUC). Random Forest achieved the strongest overall performance (accuracy 0.753; F1-score 0.689; ROC-AUC 0.810), followed by CatBoost, SVM, and XGBoost, whereas TabNet performed worst with very low recall for the diabetic class. The best-performing model (Random Forest) was deployed in a lightweight Flask web application that returns a probability-based diabetes risk assessment, categorising each prediction as low, moderate, or high risk together with a tailored recommendation. The findings confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings. Key limitations include dataset homogeneity, residual class imbalance, and limited feature coverage.
References
Ahmed, A., Aziz, S., Abd-alrazaq, A., Farooq, F., Househ, M., & Sheikh, J. (2023). The effectiveness of wearable devices using artificial intelligence for blood glucose level forecasting or prediction: Systematic review. Journal of Medical Internet Research, 25, e40259. https://doi.org/10.2196/40259
Al-Sideiri, A., Cob, B. C., & Drus, S. B. M. (2019). Machine learning algorithms for diabetes prediction. https://doi.org/10.1145/3388218.3388231
Almutairi, E. S., & Abbott, M. F. (2023). Machine learning methods for diabetes prevalence classification in Saudi Arabia. Modelling, 4(1), 37-55. https://doi.org/10.3390/modelling4010004
Arik, S. O., & Pfister, T. (2021). TabNet: Attentive interpretable tabular learning. Proceedings of the AAAI Conference on Artificial Intelligence, 35(8), 6679-6687. https://doi.org/10.1609/aaai.v35i8.16826
Ayoade, O. B., Shahrestani, S., & Ruan, C. (2025). Machine learning and deep learning approaches for predicting diabetes progression: A comparative analysis. Electronics, 14(13), 2583. https://doi.org/10.3390/electronics14132583
Bavkar, V. C., & Shinde, A. A. (2021). Machine learning algorithms for diabetes prediction and neural network method for blood glucose measurement. Indian Journal of Science and Technology, 14(10), 869-880. https://doi.org/10.17485/IJST/v14i10.2187
Bi, Q., Goodman, K. E., Kaminsky, J., & Lessler, J. (2019). What is machine learning? A primer for the epidemiologist. American Journal of Epidemiology, 188(12), 2222-2239. https://doi.org/10.1093/aje/kwz189
Bothra, R. (2021). Diabetes prediction using machine learning algorithms. International Journal of Engineering Applied Sciences and Technology, 6(5).
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32. https://doi.org/10.1023/A:1010933404324
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321-357. https://doi.org/10.1613/jair.953
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). https://doi.org/10.1145/2939672.2939785
Chowdhury, M. M., Ayon, R. S., & Hossain, M. S. (2023). An investigation of machine learning algorithms and data augmentation techniques for diabetes diagnosis using class imbalanced BRFSS dataset. Healthcare Analytics, 5, 100297. https://doi.org/10.1016/j.health.2023.100297
Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273-297. https://doi.org/10.1007/BF00994018
Ganie, S. M., Pramanik, P. K. D., Malik, M. B., Mallik, S., & Qin, H. (2023). An ensemble learning approach for diabetes prediction using boosting techniques. Frontiers in Genetics, 14, 1252159. https://doi.org/10.3389/fgene.2023.1252159
Gholampour, S. (2024). Impact of nature of medical data on machine and deep learning for imbalanced datasets: Clinical validity of SMOTE is questionable. Machine Learning and Knowledge Extraction, 6(2), 827-841. https://doi.org/10.3390/make6020039
Jithendra, V., Sai Mohit, R. M., Kusuma, S., Madhusudhan, M., & Jagadeesh, B. (2023). Diabetes prediction using machine learning techniques. Journal of Artificial Intelligence and Capsule Networks, 5(2), 190-206. https://doi.org/10.36548/jaicn.2023.2.008
Jordan, M. I., & Mitchell, T. M. (2015). Machine learning: Trends, perspectives, and prospects. Science, 349(6245), 255-260. https://doi.org/10.1126/science.aaa8415
Khanam, J. J., & Foo, S. Y. (2021). A comparison of machine learning algorithms for diabetes prediction. ICT Express, 7(4), 432-439. https://doi.org/10.1016/j.icte.2021.02.004
Kiran, M., Xie, Y., Anjum, N., Ball, G., Pierscionek, B., & Russell, D. (2025). Machine learning and artificial intelligence in type 2 diabetes prediction: A comprehensive 33-year bibliometric and literature analysis. Frontiers in Digital Health, 7, 1557467. https://doi.org/10.3389/fdgth.2025.1557467
Kivrak, M., Avci, U., Uzun, H., & Ardic, C. (2024). The impact of the SMOTE method on machine learning and ensemble learning performance results in addressing class imbalance in data used for predicting total testosterone deficiency in type 2 diabetes patients. Diagnostics, 14(23), 2634. https://doi.org/10.3390/diagnostics14232634
Kumar, S., Jaiswal, M., & Tiwari, H. (2023). Diabetes prediction using optimization techniques with machine learning algorithms. International Journal of Electronic Healthcare, 13(2), 158. https://doi.org/10.1504/IJEH.2023.130515
Kursa, M. B., & Rudnicki, W. R. (2010). Feature selection with the Boruta package. Journal of Statistical Software, 36(11), 1-13. https://doi.org/10.18637/jss.v036.i11
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30 (pp. 4765-4774).
Mahesh, B. (2020). Machine learning algorithms: A review. International Journal of Science and Research (IJSR), 9(1), 381-386. https://doi.org/10.21275/ART20203995
Mani Nageshwar, & Jayapradha, S. (2022). Diabetes prediction using machine learning algorithms. In International Conference on Advanced Computing and Communication Systems (ICACCS) (pp. 46-51). https://doi.org/10.1109/ICACCS54159.2022.9785073
Mujumdar, A., & Vaidehi, V. (2019). Diabetes prediction using machine learning algorithms. Procedia Computer Science, 165, 292-299. https://doi.org/10.1016/j.procs.2020.01.047
Okikiola, F. M., Adewale, O. S., & Obe, O. O. (2023). A diabetes prediction classifier model using Naive Bayes algorithm. FUDMA Journal of Sciences, 7(1), 253–260. https://doi.org/10.33003/fjs-2023-0701-1301
Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., & Gulin, A. (2018). CatBoost: Unbiased boosting with categorical features. In Advances in Neural Information Processing Systems 31 (pp. 6638-6648).
Rakshith, R., Varun, R., & Jagadeesh, M. G. (2022). A comparative analysis of diabetes prediction models using machine learning algorithms (pp. 261-265).
Rufai, M. A., Abdullahi, M. B., Abisoye, O. A., & Ojerinde, O. A. (2024). Utilizing a fusion of machine learning techniques for diabetes mellitus subtypes classification and identification. FUDMA Journal of Sciences, 8(3), 331–343. https://doi.org/10.33003/fjs-2024-0803-2510
Sarker, I. H. (2021). Machine learning: Algorithms, real-world applications and research directions. SN Computer Science, 2(3), 160. https://doi.org/10.1007/s42979-021-00592-x
Smith, J. W., Everhart, J. E., Dickson, W. C., Knowler, W. C., & Johannes, R. S. (1988). Using the ADAP learning algorithm to forecast the onset of diabetes mellitus. In Proceedings of the Symposium on Computer Applications and Medical Care (pp. 261-265). IEEE Computer Society Press.
Tan, Y., Chen, H., Zhang, J., Tang, R., & Liu, P. (2022). Early risk prediction of diabetes based on GA-stacking. Applied Sciences, 12(2), 632. https://doi.org/10.3390/app12020632
Tasin, I., Nabil, T. U., Islam, S., & Khan, R. (2023). Diabetes prediction using machine learning and explainable AI techniques. Healthcare Technology Letters, 10(1-2), 1-10. https://doi.org/10.1049/htl2.12039
Theerthagiri, P., Ruby, A. U., & Vidya, J. (2022). Diagnosis and classification of diabetes using machine learning algorithms. SN Computer Science, 4(1).
Wardhani, K. D. K., & Akbar, M. (2022). Diabetes risk prediction using extreme gradient boosting (XGBoost). Jurnal Online Informatika, 7(2), 244-250. https://doi.org/10.15575/join.v7i2.970
World Health Organization. (2024). Diabetes [Fact sheet]. https://www.who.int/news-room/fact-sheets/detail/diabetes
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Tosin Comfort Olayinka

This work is licensed under a Creative Commons Attribution 4.0 International License.