Toward a clinical text encoder: pretraining for clinical natural language processing with applications to substance misuse
Published in Journal of the American Medical Informatics Association, 2019
Recommended citation: Dmitriy Dligach, Majid Afshar, and Timothy Miller. 2019. Toward a clinical text encoder: pretraining for clinical natural language processing with applications to substance misuse. In Journal of the American Medical Informatics Association. https://pmc.ncbi.nlm.nih.gov/articles/PMC6798566/pdf/ocz072.pdf
Abstract:
Objective Our objective is to develop algorithms for encoding clinical text into representations that can be used for a variety of phenotyping tasks. Materials and Methods Obtaining large datasets to take advantage of highly expressive deep learning methods is difficult in clinical natural language processing (NLP). We address this difficulty by pretraining a clinical text encoder on billing code data, which is typically available in abundance. We explore several neural encoder architectures and deploy the text representations obtained from these encoders in the context of clinical text classification tasks. While our ultimate goal is learning a universal clinical text encoder, we also experiment with training a phenotype-specific encoder. A universal encoder would be more practical, but a phenotype-specific encoder could perform better for a specific task. Results We …