TY - JOUR
T1 - Adapted large language models can outperform medical experts in clinical text summarization
AU - Van Veen, Dave
AU - Van Uden, Cara
AU - Blankemeier, Louis
AU - Delbrouck, Jean-Benoit
AU - Aali, Asad
AU - Bluethgen, Christian
AU - Pareek, Anuj
AU - Polacin, Malgorzata
AU - Reis, Eduardo Pontes
AU - Seehofnerová, Anna
AU - Rohatgi, Nidhi
AU - Hosamani, Poonam
AU - Collins, William
AU - Ahuja, Neera
AU - Langlotz, Curtis P
AU - Hom, Jason
AU - Gatidis, Sergios
AU - Pauly, John
AU - Chaudhari, Akshay S
N1 - © 2024. The Author(s), under exclusive licence to Springer Nature America, Inc.
PY - 2024/4
Y1 - 2024/4
N2 - Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language processing (NLP) tasks, their effectiveness on a diverse range of clinical summarization tasks remains unproven. Here we applied adaptation methods to eight LLMs, spanning four distinct clinical summarization tasks: radiology reports, patient questions, progress notes and doctor-patient dialogue. Quantitative assessments with syntactic, semantic and conceptual NLP metrics reveal trade-offs between models and adaptation methods. A clinical reader study with 10 physicians evaluated summary completeness, correctness and conciseness; in most cases, summaries from our best-adapted LLMs were deemed either equivalent (45%) or superior (36%) compared with summaries from medical experts. The ensuing safety analysis highlights challenges faced by both LLMs and medical experts, as we connect errors to potential medical harm and categorize types of fabricated information. Our research provides evidence of LLMs outperforming medical experts in clinical text summarization across multiple tasks. This suggests that integrating LLMs into clinical workflows could alleviate documentation burden, allowing clinicians to focus more on patient care.
AB - Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language processing (NLP) tasks, their effectiveness on a diverse range of clinical summarization tasks remains unproven. Here we applied adaptation methods to eight LLMs, spanning four distinct clinical summarization tasks: radiology reports, patient questions, progress notes and doctor-patient dialogue. Quantitative assessments with syntactic, semantic and conceptual NLP metrics reveal trade-offs between models and adaptation methods. A clinical reader study with 10 physicians evaluated summary completeness, correctness and conciseness; in most cases, summaries from our best-adapted LLMs were deemed either equivalent (45%) or superior (36%) compared with summaries from medical experts. The ensuing safety analysis highlights challenges faced by both LLMs and medical experts, as we connect errors to potential medical harm and categorize types of fabricated information. Our research provides evidence of LLMs outperforming medical experts in clinical text summarization across multiple tasks. This suggests that integrating LLMs into clinical workflows could alleviate documentation burden, allowing clinicians to focus more on patient care.
KW - Humans
KW - Semantics
KW - Documentation
KW - Electronic Health Records
KW - Natural Language Processing
KW - Physician-Patient Relations
UR - https://www.scopus.com/pages/publications/85186201615
U2 - 10.1038/s41591-024-02855-5
DO - 10.1038/s41591-024-02855-5
M3 - Journal article
C2 - 38413730
SN - 1078-8956
VL - 30
SP - 1134
EP - 1142
JO - Nature Medicine
JF - Nature Medicine
IS - 4
ER -