article
Artificial intelligence shows strong potential for early detection of neurodegenerative diseases (NDDs) using medical imaging and multimodal data. However, many models suffer from limited generalizability on independent external cohorts. This systematic review and meta-analysis (PRISMA 2020) of 152 deep-learning studies ($2020-2025$) quantifies the performance gap: a consistent drop of $\mathbf{8} \boldsymbol{-} \mathbf{1 5}$ percentage points in AUC occurs between internal and external validation. Patientselection bias and reliance on curated cohorts (ADNI, PPMI) are the main contributors (QUADAS-2). Multimodal fusion achieves the highest pooled AUC (0.96) but remains fragile in real-world settings. The study calls for prioritizing reproducibility, external validation and clinical robustness over architectural novelty.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset68627.2026.11538848
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.