Detecting stigmatizing language in clinical notes with large language models for addiction care
Published in npj Health Systems, 2026
Recommended citation: Rohan Sethi, John Caskey, Yanjun Gao, Matthew M Churpek, Timothy A Miller, Anoop Mayampurath, Elizabeth Salisbury-Afshar, Majid Afshar, and Dmitry Dligach. 2026. Detecting stigmatizing language in clinical notes with large language models for addiction care. In npj Health Systems. https://www.nature.com/articles/s44401-026-00069-0
Abstract:
Intensive care units (ICU) produce numerous progress notes that may contain stigmatizing language that perpetuate negative biases and punitive approaches against patients. Patients with substance use disorders are particularly vulnerable to stigma. This study examined the performance of Large Language Models (LLMs) in the identification of stigmatizing language. We annotated a dataset with over 77,000 stigmatizing and non-stigmatizing notes from the MIMIC-III database. We utilized Meta’s Llama-3 8B Instruct LLM to run the following experiments for stigma detection: zero-shot; in-context learning; in-context learning with a selective retrieval; supervised fine-tuning (SFT); and keyword search. All approaches were evaluated on a held-out test set and external validation (University of Wisconsin Health System). SFT had the best performance with 97.2% accuracy, followed by in-context learning. The LLMs with …