article · International Journal of Environmental Research and Public Health
Bioinformatics approaches were combined with text mining and machine learning to uncover key genes and pathways associated with diabetes mellitus. By scanning over forty thousand abstracts, the work identified thousands of diabetes-linked genes, noting several frequently cited candidates tied to processes like glycogen regulation, adipogenesis, and macrophage differentiation. Subsequent gene expression analysis across three datasets comprising forty-four diabetic patients and fifty-seven healthy controls isolated one hundred and thirty-five differentially expressed genes, which were enriched in biological pathways such as aerobic respiration and immune signalling. Machine learning models, including decision trees and random forests, were trained on these findings to predict disease status, achieving predictive accuracies between sixty-three and eighty-eight percent. The computational screening pinpointed thirty-nine potential biomarkers, highlighting the HLA-DQB1 gene in particular as a prospective candidate for early detection.
Diabetes remains a widespread and complex disease whose underlying molecular causes are not completely understood. By coupling automated literature analysis with patient gene expression data and machine learning, this computational approach pinpoints specific genetic targets and biological processes. These discoveries help clarify how diabetes develops, supporting the development of earlier diagnostic tools and targeted interventions.
The findings could enable the development of early-stage diagnostic tools and biomarker panels for clinicians and diagnostic laboratories. Machine learning models trained on these genetic targets, especially HLA-DQB1, demonstrate predictive accuracy up to eighty-eight percent in distinguishing diabetic from non-diabetic profiles. This represents early-stage, computational research that requires further clinical validation before it can be integrated into commercial diagnostic products.
AI-generated from the published abstract. Always read the original work before citing.
The molecular basis of diabetes mellitus is yet to be fully elucidated. We aimed to identify the most frequently reported and differential expressed genes (DEGs) in diabetes by using bioinformatics approaches. Text mining was used to screen 40,225 article abstracts from diabetes literature. These studies highlighted 5939 diabetes-related genes spread across 22 human chromosomes, with 112 genes mentioned in more than 50 studies. Among these genes, HNF4A, PPARA, VEGFA, TCF7L2, HLA-DRB1, PPARG, NOS3, KCNJ11, PRKAA2, and HNF1A were mentioned in more than 200 articles. These genes are correlated with the regulation of glycogen and polysaccharide, adipogenesis, AGE/RAGE, and macrophage differentiation. Three datasets (44 patients and 57 controls) were subjected to gene expression analysis. The analysis revealed 135 significant DEGs, of which CEACAM6, ENPP4, HDAC5, HPCAL1, PARVG, STYXL1, VPS28, ZBTB33, ZFP37 and CCDC58 were the top 10 DEGs. These genes were enriched in aerobic respiration, T-cell antigen receptor pathway, tricarboxylic acid metabolic process, vitamin D receptor pathway, toll-like receptor signaling, and endoplasmic reticulum (ER) unfolded protein response. The results of text mining and gene expression analyses used as attribute values for machine learning (ML) analysis. The decision tree, extra-tree regressor and random forest algorithms were used in ML analysis to identify unique markers that could be used as diabetes diagnosis tools. These algorithms produced prediction models with accuracy ranges from 0.6364 to 0.88 and overall confidence interval (CI) of 95%. There were 39 biomarkers that could distinguish diabetic and non-diabetic patients, 12 of which were repeated multiple times. The majority of these genes are associated with stress response, signalling regulation, locomotion, cell motility, growth, and muscle adaptation. Machine learning algorithms highlighted the use of the HLA-DQB1 gene as a biomarker for diabetes early detection. Our data mining and gene expression analysis have provided useful information about potential biomarkers in diabetes.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3390/ijerph192113890
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.