dataset · Zenodo (CERN European Organization for Nuclear Research)
Hybrid deep-learning systems for classifying chest radiographs often combine handcrafted visual features with neural networks, but testing whether these complex additions actually improve performance requires isolating each design choice. Using an audited dataset of 20,342 images screened for patient leakage and duplicates, this study evaluated models diagnosing normal scans, COVID-19, pneumonia, and tuberculosis. Purely handcrafted models lagged behind convolutional neural network architectures by roughly five percentage points. Crucially, adding a handcrafted feature branch or gating mechanism to deep-learning models offered no statistically meaningful gain at full sample sizes, contributing only about 0.01 accuracy points. Handcrafted features demonstrated modest advantages only when training data was severely restricted, losing measurable benefit once training sets exceeded a quarter of their full size. Diagnostic errors concentrated heavily in tuberculosis cases, where 6.3 percent were incorrectly identified as normal.
Medical software developers often assume that combining traditional image-processing techniques with deep learning produces superior diagnostic tools. These findings show that such hybrid complexity offers no measurable performance advantage once sufficient training data is available. Recognising this ceiling prevents researchers and engineers from expending time and computational resources on unnecessary feature engineering that fails to improve patient diagnosis.
This research informs developers of clinical decision-support software and automated chest radiography screening tools. Positioned at an early research stage, the findings suggest developers should avoid building hybrid handcrafted feature pipelines when sufficient image data is accessible, thereby simplifying model deployment and maintenance. Real-world diagnostic use remains distant because residual error in tuberculosis detection remains high and external validation on an independent cohort failed its control.
AI-generated from the published abstract. Always read the original work before citing.
Background/Objectives: Attribution fails when two design choices move together. Hybrid deep-learning classifiers for normal, COVID-19, pneumonia and tuberculosis on chest radiographs usually change what is fused and how at once, so a gain is attributable to neither, and this study crosses them. Methods: Three pooled public sources passed exact-byte, integrity and perceptual screens; screening removed 142 near-duplicates. Patient grouping kept 703 of the 4,069 an image-level split would have produced from sharing a subject with training. Seven architectures were trained under one protocol with five-fold cross-validation on 20,342 images; four form the 2-by-2 design. Results: The handcrafted-only model trailed the six convolutional configurations, all near 97%, by 4.5 to 5.1 points, the only one of eighteen comparisons to survive correction. Neither crossed factor did: the handcrafted branch was worth +0.010 accuracy points (95% confidence interval −0.36 to +0.39), the gate −0.004 (−0.54 to +0.53). A label rule reading two directory levels had put that branch at +0.81 in one fold. The trained gate spans 0.12 to 0.88 across images, and a hundredfold higher fine-tuning rate moved its parameters further than the schedule reported here without a distinguishable change in accuracy. The null held for ranking, convergence speed and saliency: 130 of 140 zonal contrasts reversed sign. Residual error concentrates on tuberculosis: 6.3% read as normal, which is not a detection rate. A probe naming the source reaches 88.8% against 47.1%, yet transfer between collections gives 96.79% and 93.34% against 72.89% and 52.68%. Conclusions: Neither component is bought at full data; both are paid for. The handcrafted effect falls monotonically from +1.302 to +0.010 as the training fold grows; its interval excludes zero at a quarter of the fold and not at a half, which locates the ceiling. An attempt on a second cohort failed its pre-registered control, so external validation remains necessary.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.22261531
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.