MARATTO

article · BMC Medical Informatics and Decision Making

Explainable machine learning identifies diagnostic patterns for paediatric respiratory diseases in a high-dimensional underrepresented South African dataset

Abstract

This study presents a clinically-derived, real-world paediatric pulmonology dataset across underrepresented areas in South Africa and benchmarks machine learning models for the diagnostic classification of three prevalent respiratory conditions: asthma, bronchiectasis, and bronchopulmonary dysplasia. The dataset comprises of 2,176 patient records with approximately 300 clinical features and 95 diagnostic labels, reflecting the multi-label nature of routine specialist care. The records used were from the period 2010 to 2024. Using a stratified 70–30 train-test split with controlled cross-validation and hyperparameter tuning, we evaluated logistic regression, linear support vector machines, decision trees, random forests, gradient-boosted trees, LightGBM, XGBoost, CatBoost, a multilayer perceptron, and an ensemble CatBoost neural network model. Across all diseases, modern boosted tree models and the ensemble achieved consistently strong performance, with AUC values up to 0.975 for asthma, 0.947 for bronchiectasis, and 0.961 for BPD, and high precision–recall performance under class imbalance. Probabilistic scores and calibration analyses further highlighted differences in probability reliability between models, even when discrimination was similar. This is relevant for clinical screening and decision-support where reliable risk estimates are necessary. We further evaluated alternative decision thresholds and used decision-curve analysis to assess clinical usefulness, showing that the operating point can be aligned with an intended clinical objective and that the strongest models provided positive net benefit over default strategies. Lastly, model interpretability using SHAP-based explanations identified internally consistent predictor patterns for each disease target. These results demonstrate that a robust and internally consistent classification signal is present in a real-world, multi-label paediatric respiratory dataset and that explainable machine learning evaluation can support transparent comparison of modelling approaches in complex clinical environments.

Research topics

  • Chronic Obstructive Pulmonary Disease (COPD) Research
  • Respiratory and Cough-Related Research
  • Phonocardiography and Auscultation Techniques

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1186/s12911-026-03730-8

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.