A review of deep-learning methods that rebuild accelerated brain MRI scans has found that current testing often cannot answer the question clinicians care about most: whether a reconstructed image preserves what matters for diagnosis. Across 263 primary studies, only 18, or 6.8%, reported both a conventional image-fidelity score and a reader assessment on the same data. The authors conclude that this evaluation pattern cannot certify diagnostic safety.
The paper is an arXiv v1 preprint dated 28 Aug 2026. It examines reported evaluation practices for deep-learning and accelerated brain MRI reconstruction. The review followed PRISMA 2020, searched seven databases without a date limit, and combined the findings in a narrative synthesis rather than a pooled clinical analysis.
The missing test
Image-fidelity metrics reduce a reconstruction to a numerical comparison with a reference scan. That can be useful for engineering, but the review found little direct evidence showing how such scores track radiologists' judgements. No included study reported the relevant pixel-wise metric-to-reader rank correlation, and none evaluated a model observer, a computational stand-in for a human reader.
The review also illustrates why a strong whole-image score may overlook a small but important abnormality. In its example, a 100 mm3 acute lacunar infarct was assigned a local error ratio of 100. Under the stated error model, the resulting change in the global PSNR score was only 0.027 to 0.036 decibels. PSNR, or peak signal-to-noise ratio, is a numerical measure of image similarity, so a change of that size could be difficult to distinguish from ordinary score-reporting precision. The calculation is an author-derived bound, not a measured rate of missed lesions.
The review did find that reader agreement can vary with the task and conditions. In the fastMRI 2020 reader phase, concordance among six radiologists was 0.457 in the four-times acceleration track and 0.386 in the eight-times track, rising to 0.781 in the transfer track. These figures come from one named primary study and are not pooled estimates, but they underline that reader-based evaluation is itself dependent on the setting.
More warnings, fewer readers
Reader assessment did not expand in step with the review's publication eras. Its recorded share fell from 32.4% in studies published from 1995 to 2019 to 18.2% in 2020 to 2022, then recovered to 20.7% in 2023 to 2026. The gap was especially visible in method families that reported concerns about hallucinated or unstable content: generative studies recorded those concerns in 38.8% of cases but reader assessment in 11.3%, while self-supervised studies recorded hallucination or instability in 47.1% and no reader assessment.
The corpus also showed a split between computational reproducibility and clinical verification. Only 13 of the 263 studies, or 4.9%, both released code and reported reader assessment. A post-hoc comparison found a venue association with an odds ratio of 0.43 and a p-value of 0.013, but the authors describe this analysis as exploratory and do not treat it as evidence that publication venue causes the difference.
The benchmark problem was equally stark for urgent disease. No named dataset in the corpus covered acute stroke or haemorrhage. In practical terms, the review's recorded benchmark set could not directly test the failure mode in which a reconstruction alters or removes a time-sensitive lesion. It does not show that no such resource exists anywhere.
What the review can and cannot settle
The review included 263 primary studies after 1,068 reports were assessed in full. Its prevalence figures are lower-bound summaries of what was recorded, not estimates of how often deployed systems erase clinically important lesions. The extraction labels had no independent reliability estimate; the only quoted check found four design-field errors among 61 records, or 6.6%.
The authors call for five elements in safety-oriented testing: measurements tied to a clinical task, paired metric and reader assessments, checks that do not depend entirely on a reference image, robustness tests across acquisition shifts, and methods aimed specifically at detecting hallucinations. They also argue for pathology-bearing benchmarks that combine raw k-space data, acute vascular disease, lesion annotations, independent truth and permission to redistribute the data.
The project's reported counts, screening decisions, appraisal worksheets and analysis files are publicly deposited through its project records, with code copies in the accompanying repositories. The work was supported by RMIT University research-computing infrastructure and scholarship funding, with no external commercial or industry funding or vendor support reported.
Paper data and sources
Original title: Evaluating the Safety of Deep Learning-Based Brain MRI Reconstruction
Authors: Dat Tat Mai, Thai Viet Pham, Thu Nguyen Thi Dang, James Jin Kang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text