Abstract:
Accelerated MRI acquisition can improve the cost and efficiency of clinical imaging, thereby making high-quality treatment accessible to a larger portion of the population. However, accelerated MRI acquisitions commonly involve a trade-off: the quality and interpretability of the acquired images diminish, making subsequent manual clinical diagnosis unreliable. Modern Machine Learning (ML) methods can reconstruct accelerated MRI images and generate reconstructions with high average similarity scores to the fully sampled image. However, common similarity metrics are only a poor proxy for the clinical utility of an MRI image: Even though a reconstruction appears similar to the fully sampled image, its diagnostic content can be vastly different if, for example, a small pathology is not reconstructed correctly.
This insight leads to crucial changes in perspective regarding the robustness and reliability of commonly utilized ML-based reconstruction methods. As small changes in the reconstructed image can lead to diagnostic errors, we need to challenge common assumptions within the MRI reconstruction literature. First, we show that worst-case measurement noise should be evaluated in terms of its effect on clinical diagnosis rather than on similarity scores. We show that previous papers therefore overestimate the robustness of ML-based reconstruction methods. Second, we show that the clinical utility of a reconstruction is not accurately reflected in terms of its similarity scores, as high-quality reconstructions are not necessarily required to predict high-quality segmentations from accelerated MRI images. Surprisingly, we observe that classical reconstruction methods with low similarity scores can lead to the highest segmentation scores.
Additionally, we show that evaluating reconstruction methods in terms of their clinical utility, such as segmentation and pathology detection, necessitates consideration of the inherent reconstruction ambiguity. MRI reconstruction is an ill-posed problem, for which an infinite set of possible and plausible solutions exists. ML-based MRI reconstruction methods tend to reconstruct the most probable images or those that lead to the Minimum Mean Square Error (MMSE). In both cases, rare and small pathologies or image abnormalities, such as small cartilage volumes, are unlikely to be reconstructed correctly at high acceleration factors. We introduce methods to mitigate this problem by generating semantically diverse reconstructions that are both data-consistent and plausible, yet less probable. These methods not only provide interpretable uncertainty estimates but also detect more pathologies.