Sanjay Soundarajan

Sanjay Soundarajan

California Medical Innovations Institute

Current S-Index

12.3

Sum of Dataset Indices for all datasets

Average Dataset Index per Dataset

0.8

Average Dataset Index per dataset

Total Claimed Datasets

16

Total datasets claimed by the user

Average FAIR Score

81.5%

Average FAIR Score per dataset

Total Citations

6

Total citations to the user's datasets

Total Mentions

4

Total mentions of the user's datasets

S-Index Interpretation

S-Index Over Time

Cumulative Citations Over Time

Cumulative Mentions Over Time

Datasets

S-index Real-World Testing and Validation Dataset (Version: 1.0.0)

OverviewThis dataset contains dataset metadata, FAIR scores, citations, mentions, and research field data collected/generated during our real world testing and validation of the S-index. The analysis of this dataset was included in our NIH S-index Challenge Phase 2 proposal. The code used to collect this data and additional details about the data collection methods and data structure are available in the related GitHub repository. We refer to the S-index Hub for more information about our S-index, the Challenge, and other related resources. DetailsAll data files are in NDJSON format. Data was collected between October 2025 and January 2026. A DuckDB file combining this data is available here.FolderDescriptionmetadataContains reduced metadata of 49M+ datasets from the Datacite database and 51k+ datasets from the Electron Microscopy Data Bank (EMDB). Includes all datasets registered/published on or before September 30th, 2025. Data is in NDSON format following the DataCite Metadata Schema.fair_scoresContains FAIR scores of all the datasets calculated using F-UJI (between October 2025 and January 2026). citationsContains 7.7M + unique citations to our datasets collected from the Make Data Count Data Citation Corpus, OpenAlex, and DataCite. mentionsContains 91k+ mentions to our datasets collected from the README files of GitHub repositories using Software Data, Model Cards of AI/ML models on Hugging Face, and patents from the USPTO. research_fieldsContains research fields assigned to our datasets following the OpenAlex Topics taxonomy. Fields assigment comes from OpenAlex and a custom classifier we developed.

Authors

  • Patel, Bhavesh ;
  • Soundarajan, Sanjay
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.186280672026

S-index Real-World Testing and Validation Dataset (Version: 1.0.0)

OverviewThis dataset contains dataset metadata, FAIR scores, citations, mentions, and research field data collected/generated during our real world testing and validation of the S-index. The analysis of this dataset was included in our NIH S-index Challenge Phase 2 proposal. The code used to collect this data and additional details about the data collection methods and data structure are available in the related GitHub repository. We refer to the S-index Hub for more information about our S-index, the Challenge, and other related resources. DetailsAll data files are in NDJSON format. Data was collected between October 2025 and January 2026. A DuckDB file combining this data is available here.FolderDescriptionmetadataContains reduced metadata of 49M+ datasets from the Datacite database and 51k+ datasets from the Electron Microscopy Data Bank (EMDB). Includes all datasets registered/published on or before September 30th, 2025. Data is in NDSON format following the DataCite Metadata Schema.fair_scoresContains FAIR scores of all the datasets calculated using F-UJI (between October 2025 and January 2026). citationsContains 7.7M + unique citations to our datasets collected from the Make Data Count Data Citation Corpus, OpenAlex, and DataCite. mentionsContains 91k+ mentions to our datasets collected from the README files of GitHub repositories using Software Data, Model Cards of AI/ML models on Hugging Face, and patents from the USPTO. research_fieldsContains research fields assigned to our datasets following the OpenAlex Topics taxonomy. Fields assigment comes from OpenAlex and a custom classifier we developed.

Authors

  • Patel, Bhavesh ;
  • Soundarajan, Sanjay
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.186280662026

Dataset: Poster Sharing Reuse Paper (Version: 1.0.0)

AboutScientific posters are one of the most common forms of scholarly communication, with millions presented at conferences each year. They contain early-stage insights that, if shared beyond the conference, could accelerate scientific discovery. This dataset contains the data files related to our analysis of poster sharing and reuse. See this inventory for all related resources, including the paper.Standards followedThis dataset is structured according to the SPARC Data Structure v3.0.2 (c.f. related manuscript and documentation) and curated using the SPARC data curation software SODA v17.0.0.Using the datasetSimply download this dataset. The main data files are in the primary folder. The derivative folder contains data files derived from the primary files. There is a manifest.xlsx file at the root level that describes all the files. See this inventory for a link to the Jupyter notebook we developed to analyze and visualize the data.LicenseThis work is licensed under a Creative Commons Attribution 4.0 International License (see LICENSE file).How to citeIf you use this dataset, please follow the citation instruction presented on Zenodo.

Authors

  • Gasimova, Aydan ;
  • Mensah-Kane, Paapa ;
  • Blake, Gerard F. ;
  • Soundarajan, Sanjay ;
  • O'Neill, James ;
  • Patel, Bhavesh
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.187501312026

S-index Real-World Testing and Validation DB Data (Version: 1.0.0)

OverviewThis dataset contains the DuckDB file where we combined dataset metadata, FAIR scores, citations, mentions, and research field data collected/generated during our real world testing and validation of the S-index. It also contains computed Dataset Index and S-index. The analysis of this dataset was included in our NIH S-index Challenge Phase 2 proposal. The code used to collect this data and additional details about the data collection methods and data structure are available in the related GitHub repository. We refer to the S-index Hub for more information about our S-index, the Challenge, and other related resources.Related ItemsThe NDJSON files of dataset metadata, FAIR scores, citations, mentions, and research field available here were loaded in this DuckDB fileThe code used to compute the Dataset Index and S-index is available in the related GitHub repositoryThe code to analyze this data is available in this GitHub repository.

Authors

  • Patel, Bhavesh ;
  • Sanjay, Soundarajan
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.186291052026

Dataset: Poster Sharing Reuse Paper (Version: 1.0.0)

AboutScientific posters are one of the most common forms of scholarly communication, with millions presented at conferences each year. They contain early-stage insights that, if shared beyond the conference, could accelerate scientific discovery. This dataset contains the data files related to our analysis of poster sharing and reuse. See this inventory for all related resources, including the paper.Standards followedThis dataset is structured according to the SPARC Data Structure v3.0.2 (c.f. related manuscript and documentation) and curated using the SPARC data curation software SODA v17.0.0.Using the datasetSimply download this dataset. The main data files are in the primary folder. The derivative folder contains data files derived from the primary files. There is a manifest.xlsx file at the root level that describes all the files. See this inventory for a link to the Jupyter notebook we developed to analyze and visualize the data.LicenseThis work is licensed under a Creative Commons Attribution 4.0 International License (see LICENSE file).How to citeIf you use this dataset, please follow the citation instruction presented on Zenodo.

Authors

  • Gasimova, Aydan ;
  • Mensah-Kane, Paapa ;
  • Blake, Gerard F. ;
  • Soundarajan, Sanjay ;
  • O'Neill, James ;
  • Patel, Bhavesh
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.187501302026

S-index Real-World Testing and Validation DB Data (Version: 1.0.0)

OverviewThis dataset contains the DuckDB file where we combined dataset metadata, FAIR scores, citations, mentions, and research field data collected/generated during our real world testing and validation of the S-index. It also contains computed Dataset Index and S-index. The analysis of this dataset was included in our NIH S-index Challenge Phase 2 proposal. The code used to collect this data and additional details about the data collection methods and data structure are available in the related GitHub repository. We refer to the S-index Hub for more information about our S-index, the Challenge, and other related resources.Related ItemsThe NDJSON files of dataset metadata, FAIR scores, citations, mentions, and research field available here were loaded in this DuckDB fileThe code used to compute the Dataset Index and S-index is available in the related GitHub repositoryThe code to analyze this data is available in this GitHub repository.

Authors

  • Patel, Bhavesh ;
  • Sanjay, Soundarajan
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.186291042026

Dataset: FAIR AMD OCT Datasets Paper (Version: 1.2.0)

This is the dataset associated with the paper titled "Publicly Available Imaging Datasets for Age-related Macular Degeneration: Evaluation according to the Findable, Accessible, Interoperable, Reproducible (FAIR) Principles". Age-related macular degeneration (AMD), a leading cause of vision loss among older adults, affects more than 200 million people worldwide. In this paper, We evaluated openly available AMD-related datasets containing optical coherence tomography (OCT) data against the FAIR principles. This is an archive of the repository that contains the data related to our evaluation. See this inventory for all related resources, including the paper. This dataset is maintained from https://github.com/fairdataihub/FAIR-AMD-OCT-paper-dataset.

Authors

  • Gim, Nayoon ;
  • Ferguson, Alina ;
  • Blazes, Marian ;
  • Soundarajan, Sanjay ;
  • Gasimova, Aydan ;
  • Patel, Bhavesh ;
  • Lee, Cecilia
1 Citation0 Mentions88% FAIR0.9 Dataset Index
10.5281/zenodo.149267622025

Dataset: FAIR AMD OCT Datasets Paper (Version: 1.2.0)

This is the dataset associated with the paper titled "Publicly Available Imaging Datasets for Age-related Macular Degeneration: Evaluation according to the Findable, Accessible, Interoperable, Reproducible (FAIR) Principles". Age-related macular degeneration (AMD), a leading cause of vision loss among older adults, affects more than 200 million people worldwide. In this paper, We evaluated openly available AMD-related datasets containing optical coherence tomography (OCT) data against the FAIR principles. This is an archive of the repository that contains the data related to our evaluation. See this inventory for all related resources, including the paper. This dataset is maintained from https://github.com/fairdataihub/FAIR-AMD-OCT-paper-dataset.

Authors

  • Gim, Nayoon ;
  • Ferguson, Alina ;
  • Blazes, Marian ;
  • Soundarajan, Sanjay ;
  • Gasimova, Aydan ;
  • Patel, Bhavesh ;
  • Lee, Cecilia
0 Citations0 Mentions73% FAIR0.5 Dataset Index
10.5281/zenodo.126696512025

Dataset: Dataset Documentation for AI Paper (Version: 1.0.0)

This is the dataset associated with our paper titled "Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets". Artificial Intelligence (AI) is rapidly transforming healthcare, but also raising concerns about algorithmic biases that mostly stem from the training data. It is widely supported that transparent dataset documentation is key to enabling responsible AI development. Several standardized dataset documentation approaches have been established, such as Datasheet, Dataset Nutrition Label, Accountability Documentation, Healthsheet, and Data Card. However, their suitability and usage for health datasets remain unclear. In this paper, we compared all five approaches and evaluated their alignment with the STANDING Together Recommendations for Documentation of Health Datasets. We also investigated their real-world usage and gathered insights from generators and consumers of health datasets.The paper will be linked here when published: https://github.com/AI-READI/dataset-documentation-paper-inventory

Authors

  • Heinke, Anna ;
  • Huang, Lingling ;
  • Simpkins, Kyongmi U. ;
  • Kalaw, Fritz Gerald P. ;
  • Karsolia, Apoorva ;
  • Singh, Kiratjit ;
  • Soundarajan, Sanjay ;
  • Patel, Bhavesh
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.173087422025

Dataset: Dataset Documentation for AI Paper (Version: 1.0.0)

This is the dataset associated with our paper titled "Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets". Artificial Intelligence (AI) is rapidly transforming healthcare, but also raising concerns about algorithmic biases that mostly stem from the training data. It is widely supported that transparent dataset documentation is key to enabling responsible AI development. Several standardized dataset documentation approaches have been established, such as Datasheet, Dataset Nutrition Label, Accountability Documentation, Healthsheet, and Data Card. However, their suitability and usage for health datasets remain unclear. In this paper, we compared all five approaches and evaluated their alignment with the STANDING Together Recommendations for Documentation of Health Datasets. We also investigated their real-world usage and gathered insights from generators and consumers of health datasets.The paper will be linked here when published: https://github.com/AI-READI/dataset-documentation-paper-inventory

Authors

  • Heinke, Anna ;
  • Huang, Lingling ;
  • Simpkins, Kyongmi U. ;
  • Kalaw, Fritz Gerald P. ;
  • Karsolia, Apoorva ;
  • Singh, Kiratjit ;
  • Soundarajan, Sanjay ;
  • Patel, Bhavesh
2 Citations0 Mentions85% FAIR1.2 Dataset Index
10.5281/zenodo.173087432025