Version 1.0.1

EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries

Kweon, Sunjun;Kim, Jiyoun;Kwak, Heeyoung;Cha, Dongchul;Yoon, Hangyul;Kim, Kwang Hyun;Yang, Jeewon;Won, Seunghyun;Choi, Edward

Description

Discharge summaries in Electronic Health Records (EHRs) are crucial forclinical decision-making, but their length and complexity make informationextraction challenging, especially when dealing with accumulated summariesacross multiple patient admissions. Large Language Models (LLMs) show promisein addressing this challenge by efficiently analyzing vast and complex data.Existing benchmarks, however, fall short in properly evaluating LLMs'capabilities in this context, as they typically focus on single-noteinformation or limited topics, failing to reflect the real-world inquiriesrequired by clinicians. To bridge this gap, we introduce EHRNoteQA, a novelbenchmark built on the MIMIC-IV EHR, comprising 962 different QA pairs eachlinked to distinct patients' discharge summaries. Every QA pair is initiallygenerated using GPT-4 and then manually reviewed and refined by threeclinicians to ensure clinical relevance. EHRNoteQA includes questions thatrequire information across multiple discharge summaries and covers eightdiverse topics, mirroring the complexity and diversity of real clinicalinquiries.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.5

FAIR Score

73%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

PhysioNet

Assigned Domain

Subfield

Plant Science

Field

Agricultural and Biological Sciences

Domain

Life Sciences

Confidence Score

58%

Source

Open Alex

Normalization Factors

FT

53.85

CTw

1.00

MTw

1.00