article · Biostatistics & Epidemiology
Tuberculosis remains a critical global health problem, with recurrent cases posing serious threats to patient outcomes and transmission control within two years of completing treatment. To support earlier identification of individuals at high risk of recurrence, machine learning models were developed using patient data from the Guelmim region of Morocco. After applying a random forest feature selection method, seven distinct algorithms were tested, including logistic regression, random forest, support vector machines, and gradient boosting techniques. The primary predictors identified were weight, age, and household size. Performance varied across models, with extreme gradient boosting delivering the highest overall discriminative performance, light gradient boosting machine achieving the strongest sensitivity, and random forest yielding the highest precision and F1-score.
Recurrent tuberculosis often arises within two years of initial treatment completion, hindering disease control and worsening patient prognosis. By pinpointing key risk factors such as age, weight, and household size, predictive tools can help healthcare providers spot patients at high risk of recurrence early, enabling timely interventions that improve long-term outcomes and limit community transmission.
This research represents early-stage algorithmic development that could eventually inform decision-support software for healthcare providers and public health programmes managing tuberculosis cases. However, with the highest discriminative performance reaching an area under the curve of 63.4 percent and a top precision of 19 percent, the models remain far from clinical deployment and require significant refinement and validation before real-world use.
AI-generated from the published abstract. Always read the original work before citing.
Despite sustained global efforts, tuberculosis (TB) remains a major public health challenge. Recurrent TB, frequently occurring within two years after treatment completion, poses serious concerns for disease control and patient prognosis. Early identification of individuals at high risk of recurrence is therefore critical to improving outcomes and reducing disease transmission. This study aimed to develop a machine learning-based predictive model for recurrent TB in the Guelmim region of Morocco. A random forest-based method was applied for feature selection, followed by the development of seven machine learning models, including logistic regression, random forest, support vector machine, k-nearest neighbours, extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), and gradient boosting machine (GBM). Models were trained on 80% of the data using five-fold cross-validation and evaluated on an independent 20% test set, with performance assessed using several metrics, particularly AUROC and PR-AUC. Key predictors of recurrence included weight (12.54%), age (12.07%), and household size (8.57%). Among the seven machine learning models evaluated, XGBoost achieved the highest discriminative performance (AUROC = 63.4%), while LightGBM demonstrated the highest sensitivity (88.6%), and Random Forest yielded the best precision (19%), F1-score (21.6%), and PR-AUC (12.9%).
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1080/24709360.2026.2715243
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.