| Home > Publications database > Robustness of radiomics-based machine learning to real-world MRI distribution shifts: a phantom study. |
| Journal Article | DKFZ-2026-01983 |
; ; ;
2026
IOP Publ.
Bristol
Abstract: This study systematically investigates the robustness of radiomics-based machine learning models to distribution shifts caused by variations in MRI acquisition protocols and image segmentation strategies, and evaluates strategies for maintaining predictive performance and uncertainty calibration under such shifts. Approach. Using a controlled phantom comprising 16 fruits of four types, we acquired images across five MRI sequences (T2-HASTE, T2-TSE, T2-MAP, T1-TSE, T2-FLAIR) with multiple scans, observers, and segmentation variants (full, partial, rotated). XGBoost classifiers were trained using three feature selection strategies: eight protocol-invariant features, sequence-specific robust features, and all 107 available features. Model performance and calibration were evaluated under in-domain, cross-protocol, segmentation-induced, and compound distribution shift scenarios using a nested, group-wise (leave-one-row-out) cross-validation scheme that prevents instance-level data leakage. Because all data derive from a controlled phantom comprising only 16 independent physical objects, reported variability reflects repeated technical evaluation conditions (scans, observers, segmentation variants) rather than uncertainty across independent biological samples. Main results. Under in-domain conditions, all three feature strategies achieved comparable performance. Under distribution shifts, protocol-invariant features retained substantially more predictive power than sequence-specific or all-feature models, and multi-protocol training further improved generalization, though residual degradation persisted. Calibration proved far more vulnerable than accuracy, with ECE reaching 0.4 under compound shifts. Training-time feature-space augmentation yielded more consistent improvements in calibration than post-hoc temperature scaling, though its effect on predictive performance was mixed. Significance. These findings establish protocol-invariant feature selection, multi-protocol training, and feature-space augmentation as complementary strategies for improving radiomics model resilience, and identify calibration as a vulnerable dimension of model reliability under distribution shift. The controlled phantom framework isolates specific failure modes and enables quantitative benchmarking of robustness strategies prior to clinical validation. As this study is based on a controlled phantom, these findings are hypothesis-generating and require independent validation in clinical cohorts before clinical translation.
Keyword(s): distribution shift ; machine learning ; radiomics ; robustness analysis
|
The record appears in these collections: |