Natural language processing to build a computable phenotype library for adults with congenital heart disease
Published in International Journal of Medical Informatics, 2026
Recommended citation: Spencer Thomas, Angus Dawson, Hifsa Chaudhry, Sidra Ahmad, Xiyu Ding, Sarah A Hummel, David M Leone, Rohith Vanam, Angela J Weingarten, Eric Farber-Eger, Lauren Lee Shaffer, Benjamin P Frischhertz, Sydney St Clemmons, Sunil J Ghelani, Fernando Baraona Reyes, Tzu-Chun Wu, Danny TY Wu, Alexander R Opotowsky, and Timothy A Miller. 2026. Natural language processing to build a computable phenotype library for adults with congenital heart disease. In International Journal of Medical Informatics. https://www.sciencedirect.com/science/article/pii/S1386505626004144
Abstract:
Objective Our objective was to build classifiers for multiple phenotypes that categorize a cohort of adults with congenital heart disease (ACHD), that can be used to populate variables in a biobank. Materials and methods A dataset of 1492 ACHD patients, with expert-created labels for eight phenotypes, was created and used to train classifiers with three different architectures. A larger unlabeled dataset containing 15,869 patients was used to pre-train the classifiers, and a 20 % subset of the unlabeled dataset was used to validate the classifier predictions. Results On held out labeled data, F1 scores for the eight target phenotypes of interest ranged from 0.66 to 1. Of those, the six phenotypes with best classification performance were then validated on unlabeled data, where positive predictive value ranged from 81.5 % to 100 %. Discussion We were able to classify six out of eight phenotypes with satisfactory …