Comparing prognostic performance and reasoning between large language models and physicians

Published in medRxiv, 2026

Recommended citation: Megan Gjertsen, WonJin Yoon, Majid Afshar, Brandon Temte, Brandon Leding, Stephen Halliday, Kaitlin Bradley, Joseph Kim, Juliana Mitchell, Anna K Sanders, Emma Croxford, John Caskey, Matthew Churpek, Anoop Mayampurath, Yanjun Gao, Timothy Miller, and Jacqueline M Kruser. 2026. Comparing prognostic performance and reasoning between large language models and physicians. In medRxiv. https://www.medrxiv.org/content/medrxiv/early/2026/04/25/2026.04.17.26350898.full.pdf

Abstract:

Importance Physicians routinely prognosticate to guide care delivery and shared decision making, particularly when caring for patients with critical illnesses. Yet, these physician estimates are prone to inaccuracy and uncertainty. Artificial intelligence, including large language models (LLMs), show promise in supporting or improving this prognostication. However, the performance of contemporary LLMs in prognosticating for the heterogeneous population of critically ill patients remains poorly understood. Objective To characterize and compare the performance of LLMs and physicians when predicting 6-month mortality for hospitalized adults who survived critical illness. Design Embedded mixed methods study with elicitation and comparison of prognostic estimates and reasoning from LLMs and practicing physicians. Setting The publicly available, deidentified Medical Information Mart for Intensive Care (MIMIC)-IV v2.2 dataset. Participants We randomly selected 100 hospitalizations of adult survivors of critical illness. Four contemporary LLMs (Open AI GPT-4o, o3- and o4-mini, and DeepSeek-R1) and 7 physicians provided independent prognostic estimates for each case (1,100 total estimates; 400 LLM and 700 physician). Main outcomes and measures For each case, LLMs and physicians used the hospital discharge summary and demographics to predict 6-month mortality (yes/no) and provide their reasoning (free text). We assessed prognostic performance using accuracy, sensitivity, and specificity, and used inductive, qualitative content analysis to characterize reasonings. Results Mean physician accuracy for predicting mortality was 70.1% (95% CI …