Comparing prognostic performance and reasoning between physicians and large language models

Published in Journal of Pain and Symptom Management, 2026

Recommended citation: Megan Gjertsen, WonJin Yoon, Majid Afshar, Emma Croxford, John R Caskey, Yanjun Gao, Timothy A Miller, and Jacqueline M Kruser. 2026. Comparing prognostic performance and reasoning between physicians and large language models. In Journal of Pain and Symptom Management. https://www.jpsmjournal.com/article/S0885-3924(26)00333-7/fulltext

Abstract:

Background Physicians routinely prognosticate to guide shared decision making, particularly when caring for patients with critical illnesses (1). Yet, these physician estimates are prone to inaccuracy and uncertainty (2,3). Artificial intelligence, including large language models (LLMs), shows promise in supporting or improving prognostication (4). However, the performance of contemporary LLMs in prognosticating for the heterogeneous population of critically ill patients remains poorly understood. Objective To compare the performance and reasoning of LLMs and physicians when predicting 6-month mortality for hospitalized, critically ill adults. Methods We randomly selected 100 patient admissions with ICU stays from MIMIC-IV, a publicly available, single-center dataset of hospitalizations. For each case, 4 contemporary LLMs and 7 practicing physicians were asked to use the discharge summary and demographics to …