review · Diagnostic and Prognostic Research
BACKGROUND: Deep learning (DL)-assisted low-dose computed tomography (LDCT) may improve lung cancer screening, but the available evidence is heterogeneous and patient-level diagnostic accuracy remains uncertain. METHODS: MEDLINE, Embase, and Web of Science were searched from January 2010 to December 2025 for studies evaluating DL-assisted LDCT in lung cancer screening or screening-relevant populations. Two reviewers independently screened studies, extracted data, and assessed risk of bias using QUADAS-2. AI-specific reporting completeness was assessed descriptively using items adapted from CLAIM and STARD-AI. Studies with complete or reconstructible patient-level 2 × 2 data at a defined threshold were pooled using bivariate random-effects and hierarchical summary receiver operating characteristic models. Studies without sufficient threshold-specific data were synthesised narratively. RESULTS: Eleven studies met the inclusion criteria. Five studies, comprising 2,220 participants, 232 lung cancer cases, and 1,988 non-cases, provided usable threshold-specific 2 × 2 data and were included in the meta-analysis. Six studies were synthesised narratively because patient-level TP, FP, FN, and TN values were unavailable or not reconstructible. Pooled sensitivity was 85.4% (95% CI, 79.1-90.0), and pooled specificity was 83.5% (95% CI, 73.5-90.2). The positive likelihood ratio was 5.17, the negative likelihood ratio was 0.18, and the diagnostic odds ratio was 29.53. No study was at low risk of bias across all QUADAS-2 domains, and AI-specific reporting gaps were common. Subgroup, meta-regression, and sensitivity analyses were not feasible because only five studies were quantitatively eligible. No statistical evidence of small-study effects was detected, although this assessment was inconclusive because only five studies were pooled. CONCLUSIONS: DL-assisted LDCT shows promising but preliminary diagnostic accuracy for lung cancer screening. However, the small, heterogeneous, and methodologically limited evidence base does not support autonomous clinical use. DL is currently best considered a decision-support tool within radiologist-led screening pathways, pending prospective external validation and workflow-based evaluation.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1186/s41512-026-00235-w
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.