MARATTO

article

A Comparative Analysis of Random Forest and Decision Tree Classifiers for Predicting Type 2 Diabetes using K-Fold Cross-Validation

20242 citationsIbn Tofail University

Abstract

This study investigates the prediction accuracy of type 2 diabetes using the Pima-Indians-Diabetes dataset, employing a novel classifier that integrates Random Forest and K-fold cross-validation. Additionally, the findings are compared with predictions obtained from a Stochastic Gradient Descent (SGD) classifier. The dataset, consisting of 768 samples, was used for both training and testing. Using a significance level of α=0.05 and power=0.80 in the G*Power analysis, the results consistently approached an accuracy of 80%. The Decision Tree classifier achieved an accuracy of 86.245%, while the novel Random Forest classifier outperformed it with an accuracy of 96.179%. An Independent Samples T-Test yielded a p-value of 0.01 (p<0.05), indicating a statistically significant relationship between the performance of the Random Forest classifier, K-fold cross-validation, and the Decision Tree classifier. Overall, the Random Forest classifier demonstrated superior predictive accuracy, highlighting its effectiveness in type 2 diabetes prediction.

Research topics

  • Artificial Intelligence in Healthcare

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/isaect64333.2024.10799515

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.