Leveraging machine learning for personalized type 2 diabetes prediction: A comparative analysis
2026 (English)In: Array, E-ISSN 2590-0056, Vol. 31, article id 101031Article in journal (Refereed) Published
Abstract [en]
This study presents a comparative machine learning (ML) benchmarking framework for the early prediction and risk stratification of Type 2 diabetes (T2D) in adult patients. The work is positioned as a comparative feasibility study rather than as a deployable clinical tool. Using the Pima Indian Diabetes Dataset (768 records of female subjects aged 21 years and older, 9 clinical features), we comparatively evaluate nine ML classifiers Support Vector Machine (SVM), CatBoost, Voting Classifier, K-Nearest Neighbors (KNN), XGBoost, Random Forest (RF), Logistic Regression (LR), Gradient Boosting (GB), and Decision Tree (DT) across five train-test splits (5%-50%) and four k-fold cross-validation settings (k = 2, 4, 6, 8). Model evaluation reports multiple complementary metrics (accuracy, precision, recall, F1-score, ROC-AUC, and confusion-matrix-based error counts) together with cross-validation standard deviations interpreted as approximate uncertainty bounds, providing a more reliable picture than accuracy alone. Among the evaluated models, Random Forest provided the most consistent overall performance, with a maximum accuracy of 79.3% and the largest ROC-AUC (0.83); SVM remained stable in the 72.5%-77.2% range; and the Voting Classifier produced balanced but non-superior aggregate performance. We explicitly acknowledge several scope limitations: (i) the dataset is single-source and demographically homogeneous (adult Pima Indian females), so the findings are not directly generalizable to children, adolescents, men, or other ethnic groups; (ii) the analysis is internal cross-validation only, without external or prospective multi-center validation; (iii) explainability techniques (SHAP, LIME), calibration, decision-curve analysis, and advanced class-imbalance handling are identified as future research directions and are not part of the current scope. Within these boundaries, the study contributes a reproducible, multi-metric methodological template for comparative T2D risk-prediction studies, and explicitly identifies which methodological gaps the subsequent research should address.
Place, publisher, year, edition, pages
Elsevier, 2026. Vol. 31, article id 101031
Keywords [en]
Diabetes prediction, Machine learning, Healthcare analytics, Predictive modeling, Risk stratification, Data-driven healthcare, Algorithmic approaches
National Category
Computer Sciences Endocrinology and Diabetes
Identifiers
URN: urn:nbn:se:uu:diva-594787DOI: 10.1016/j.array.2026.101031ISI: 001820487800001Scopus ID: 2-s2.0-105043644802OAI: oai:DiVA.org:uu-594787DiVA, id: diva2:2090258
Funder
Swedish Foundation for Strategic Research, FUS21-0067Swedish Foundation for Strategic Research, CH10-00032026-08-062026-08-062026-08-06Bibliographically approved