EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
Description
Discharge summaries in Electronic Health Records (EHRs) are crucial forclinical decision-making, but their length and complexity make informationextraction challenging, especially when dealing with accumulated summariesacross multiple patient admissions. Large Language Models (LLMs) show promisein addressing this challenge by efficiently analyzing vast and complex data.Existing benchmarks, however, fall short in properly evaluating LLMs'capabilities in this context, as they typically focus on single-noteinformation or limited topics, failing to reflect the real-world inquiriesrequired by clinicians. To bridge this gap, we introduce EHRNoteQA, a novelbenchmark built on the MIMIC-IV EHR, comprising 962 different QA pairs eachlinked to distinct patients' discharge summaries. Every QA pair is initiallygenerated using GPT-4 and then manually reviewed and refined by threeclinicians to ensure clinical relevance. EHRNoteQA includes questions thatrequire information across multiple discharge summaries and covers eightdiverse topics, mirroring the complexity and diversity of real clinicalinquiries.
Citations (0)
No citations found
Mentions (0)
No mentions found
Metrics Over Time
Publication Details
Subfield
Plant Science
Field
Agricultural and Biological Sciences
Domain
Life Sciences
Confidence Score
58%
Source
Open Alex