SHIELD Dataset

Stanford Medicine Healthcare Data for Research

Description

SHIELD (Synthetic Human-annotated Identifier-replaced Entries for Learning and De-identification) is a diverse clinical note dataset designed for the development and evaluation of de-identification systems for electronic health records. The dataset contains 1,381 de-identified clinical notes with 10,229 gold-standard PHI (Protected Health Information) spans across 9 entity categories, sourced from the Stanford Medicine STARR-OMOP clinical data warehouse. Notes were selected via set-cover diversity sampling across demographic and document-type strata, ensuring broad representation of clinical language variation. All notes underwent human-in-the-loop annotation with expert adjudication to establish high-quality gold-standard PHI boundaries.SHIELD is intended for the evaluation of PII detection and de-identification systems on clinical text. Every note has been de-identified using a cryptographic surrogate-replacement pipeline, which replaces each PHI span with a type-appropriate surrogate that preserves the linguistic structure and readability of the original text. The dataset is provided in Parquet format with train/validation/test splits and is compatible with standard NLP evaluation frameworks.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.5

FAIR Score

85%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Redivis

Assigned Domain

Subfield

Molecular Biology

Field

Biochemistry, Genetics and Molecular Biology

Domain

Life Sciences

Confidence Score

40%

Source

Scholar Data Model

Keywords

file

Normalization Factors

FT

53.85

CTw

1.00

MTw

1.00