Using Relative Risk Rankings to Understand Information Differences in Multimodal Prediction Models

Published in medRxiv, 2025

Recommended citation: Chanhwi Kim, WonJin Yoon, Hoonick Lee, Jung-Oh Lee, Majid Afshar, Jaewoo Kang, and Timothy Miller. 2025. Using Relative Risk Rankings to Understand Information Differences in Multimodal Prediction Models. In medRxiv. https://www.medrxiv.org/content/medrxiv/early/2026/04/07/2025.10.30.25339162.full.pdf

Abstract:

Recent multimodal models increasingly combine heterogeneous clinical data, yet in practice raw modalities are often replaced by expert-written summaries for convenience. However, whether such representational substitution, such as replacing medical images with their expert-written reports, preserves prognostic information has not been systematically characterized. Quantifying such gaps is essential for identifying which representational modality is most informative for predictive modeling of patient outcomes. We investigated this question by comparing the predictive utility of chest radiographs (CXRs) and their paired radiology report text for 30-day post-discharge mortality prediction using visionlanguage models (VLMs). Using the discharge note summary as global clinical context at discharge, we augmented it with either the latest pre-discharge CXR or the corresponding radiology report. On a linked MIMIC-IV/MIMIC-CXR subset with paired inputs (n = 1,360), the discharge-note-plus-CXR model achieved the best performance (AUROC = 0.864), compared with the discharge-note-only model (AUROC = 0.831) and the discharge-note-plus-report model (AUROC = 0.813). To quantify the effect of modality substitution on predictive behavior, we measured Kendall’s τ-based distances between predicted risk rankings. Inter-modality distances exceeded intra-modality distances, indicating that replacing CXRs with reports changes risk prioritization rather than merely reducing overall discrimination. Post hoc review by a radiologist further suggested that reports, as clinically oriented summaries, may not exhaustively document visually available …