Description
SHIELD (Synthetic Human-annotated Identifier-replaced Entries for Learning and De-identification) is a diverse clinical note dataset designed for the development and evaluation of de-identification systems for electronic health records. The dataset contains 1,381 de-identified clinical notes with 10,229 gold-standard PHI (Protected Health Information) spans across 9 entity categories, sourced from the Stanford Medicine STARR-OMOP clinical data warehouse. Notes were selected via set-cover diversity sampling across demographic and document-type strata, ensuring broad representation of clinical language variation. All notes underwent human-in-the-loop annotation with expert adjudication to establish high-quality gold-standard PHI boundaries.SHIELD is intended for the evaluation of PII detection and de-identification systems on clinical text. Every note has been de-identified using a cryptographic surrogate-replacement pipeline, which replaces each PHI span with a type-appropriate surrogate that preserves the linguistic structure and readability of the original text. The dataset is provided in Parquet format with train/validation/test splits and is compatible with standard NLP evaluation frameworks.
Citations (0)
No citations found
Mentions (0)
No mentions found
Metrics Over Time
Publication Details
Subfield
Molecular Biology
Field
Biochemistry, Genetics and Molecular Biology
Domain
Life Sciences
Confidence Score
40%
Source
Scholar Data Model