Version 1

Data from: Building the graph of medicine from millions of clinical narratives

Finlayson, Samuel G.;LePendu, Paea;Shah, Nigam H.

Description

Electronic health records (EHR) represent a rich and relatively untapped resource for characterizing the true nature of clinical practice and for quantifying the degree of inter-relatedness of medical entities such as drugs, diseases, procedures and devices. We provide a unique set of co-occurrence matrices, quantifying the pairwise mentions of 3 million terms mapped onto 1 million clinical concepts, calculated from the raw text of 20 million clinical notes spanning 19 years of data. Co-frequencies were computed by means of a parallelized annotation, hashing, and counting pipeline that was applied over clinical notes from Stanford Hospitals and Clinics. The co-occurrence matrix quantifies the relatedness among medical concepts which can serve as the basis for many statistical tests, and can be used to directly compute Bayesian conditional probabilities, association rules, as well as a range of test statistics such as relative risks and odds ratios. This dataset can be leveraged to quantitatively assess comorbidity, drug-drug, and drug-disease patterns for a range of clinical, epidemiological, and financial applications.

Citations (0)

Mentions (0)

Metrics

Dataset Index

838.9

FAIR Score

77%

Citations

3

Mentions

1,485

Metrics Over Time

Publication Details

DOI

Publisher

Dryad

License

Creative Commons Zero v1.0 Universal

Assigned Domain

Subfield

Statistics and Probability

Field

Mathematics

Domain

Physical Sciences

Confidence Score

56%

Source

Scholar Data Model

Keywords

biomedical informaticsData miningelectronic health records

Normalization Factors

FT

44.23

CTw

1.00

MTw

1.00