article
Deep learning models for autism detection using resting-state fMRI often report overly optimistic results, largely due to limited cross-site validation. In many cases, averaged performance metrics obscure important differences between sites. To address this, we evaluated cross-site generalization using a leave-one-site-out strategy on 884 participants (408 autism, 476 controls) from 17 ABIDE sites. Feature selection and normalization were performed strictly within training folds, and model training included validation splits, early stopping, and bootstrap confidence intervals. The model achieved a mean AUC of 0.665 (95% CI: 0.615$0.713; \mathrm{p}\lt 0.0001$), consistent with prior rigorous studies. However, performance varied substantially across sites (AUC range: $0.411-0.808; \sigma=0.102$). Only $29 \%$ of sites exceeded an AUC of 0.70, while $47 \%$ remained below 0.65, with no significant link to sample size $(\mathrm{r}=-0.19, \mathrm{p}=0.46)$. These findings show that averaged metrics can be misleading, as they hide strong site-level variability. Nearly half of the sites fall below clinically relevant thresholds, highlighting the need for site-specific calibration or harmonization before clinical use.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset68627.2026.11538564
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.