Could an emerging deep learning model have an impact for automated prediction of BI-RADS classifications based on dynamic contrast-enhanced (DCE) breast MRI?
For a retrospective multicenter study, recently published in the European Journal of Radiology, researchers assessed the use of the deep learning model in 158 women (median age of 50) who had breast MRI exams. Twenty-one exams had malignant findings, according to the study. The study authors noted that the deep learning model was trained on 12,241 breast MRI exams.
Overall, the researchers found in the external validation testing that radiologist assessment of breast MRI exams achieved a 91.3 percent AUC in comparison to 76.7 percent for the deep learning model.
Radiologist assessment yielded 95.2 percent sensitivity and 77.4 percent specificity in the external validation cohort in contrast to 90.5 percent sensitivity and 54.7 percent specificity for deep learning assessment, according to the study authors.
“A deep learning-based model trained to predict BI-RADS assessment from breast MRI showed lower discrimination than routine radiologist assessment during external validation,” noted lead study author Peter Brader, MD, who is affiliated with Diagnostikum Linz in Linz, Austria, and colleagues.
Three Key Takeaways
• Radiologists still outperform the deep learning model on external validation data. In external validation, radiologist assessment had a higher AUC than the deep learning model (91.3 percent vs. 76.7 percent). Sensitivity was similar (95.2 percent vs. 90.5 percent), but the model's specificity was much lower (54.7 percent vs. 77.4 percent). In practice, the model would generate substantially more false positives so it isn't ready to replace or independently match radiologist BI-RADS assessment.
• The model’s most promising role may be triaging low-risk exams. Using a BI-RADS 3 threshold, the model achieved a 97.4 percent negative predictive value while classifying nearly half of benign exams as low risk. That suggests it could eventually help prioritize worklists or flag likely benign studies. The authors stress, however, that this is only a suggested operating point for future study, not a demonstrated workflow benefit.
• The evidence is preliminary. The external cohort was small with only 21 malignant exams among 158. The model was also trained on radiologist-assigned BI-RADS categories rather than pathology-proven outcomes so it may inherit radiologists' interpretive variability. Prospective implementation and reader studies are needed before the deep learning model could be used clinically.
Employing a BI-RADS category 3 threshold in the internal validation cohort, researchers noted a 91.3 percent AUC for differentiating between low-risk (BI-RADS < 3) and high-risk (BI-RADS > 3) assessments. The study authors also pointed out that the deep learning model demonstrated a negative predictive value (NPV) of 97.4 percent for nearly 50 percent of benign exams assessed as low-risk presentations.
“This finding identifies a potential operating point for future study but does not directly demonstrate a workflow benefit. Prospective implementation and reader studies are required to determine whether AI-based prioritization can improve efficiency without compromising diagnostic safety,” added Brader and colleagues.
(Editor’s note: For related content, see “Breast MRI Study Reveals 29 Percent Reduction in Scan Time with Deep Learning Reconstruction and Multi-Shot DWI,” “Can MRI-Derived Tumor Habitat Subregions Bolster Recurrence Risk Stratification in Stage 2 and 3 Breast Cancer Cases?” and “Video: Stamatia Destounis, MD, FACR, Discusses Key Changes for the Updated BI-RADS System”)
In regard to study limitations, the authors acknowledged a relatively small cohort for external validation (with 21 malignant exams) and conceded that the development of the AI model was dependent on BI-RADS category assessments by radiologists.