Evaluating retrieval-augmented generation versus long-context input for clinical reasoning over electronic health records
Published in Journal of the American Medical Informatics Association, 2026
Recommended citation: Skatje Myers, Dmitriy Dligach, Timothy A Miller, Samantha Barr, James Landefeld, Yanjun Gao, Matthew M Churpek, Anoop Mayampurath, and Majid Afshar. 2026. Evaluating retrieval-augmented generation versus long-context input for clinical reasoning over electronic health records. In Journal of the American Medical Informatics Association.
Abstract:
Objective To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Materials and Methods We defined 3 EHR-based tasks that are replicable across health systems and vary in reasoning complexity: (1) extracting imaging procedures (modality, date, and anatomic site), (2) generating timelines of therapeutic antibiotic use, and (3) identifying the key diagnoses for a hospitalization. Using real inpatient clinical notes from a US academic health system, we evaluated 3 large language models (GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1) with varying amounts of provided context, comparing targeted retrieval to using the most recent clinical notes. Results For Imaging Procedures, RAG strongly outperformed recent-note inputs and …