Automated Author ProfileL.Á, CuellarOlmedo
L.Á, CuellarOlmedo
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets for this author
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the author's datasets
Total Mentions
Total mentions of the author's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 1.6 (sum of 2 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
Background: Persons with type 2 diabetes mellitus (T2DM) attending hospitals frequently experience major complications. We assessed the potential use of unstructured free-text data extracted from electronic health records (EHRs) using natural language processing (NLP) and machine learning (ML) to develop a predictive model for chronic kidney disease (CKD) in T2DM.Methods: This multicenter retrospective study included data from eight Spanish hospitals (2013–2018), extracted using NLP and ML techniques (EHRead®) based on SNOMED CT terminology. From a cohort of individuals with T2DM, we identified those with and without CKD at inclusion. Among individuals without CKD, we trained and validated a two-year predictive model for CKD development. The model showing the best balance between performance and clinical interpretability was selected for integration into a web-based tool to support early detection and risk stratification.Results: Of 588,786 individuals with T2DM, 316,597 were included for model development [training: 291,429 (92.1%); validation: 25,168 (7.9%); CKD incidence: 15.4% and 18.4%, respectively]. A high proportion of missing data was observed in key clinical variables. Among models evaluated, logistic regression (LR) achieved the best performance (AUC-ROC 0.72) using 27 predictors. Both a reduced 10-predictor model and a clinically refined 8-predictor model showed comparable performance to the full model in training and validation cohorts. The clinically refined model was selected for implementation in the web-based tool.Conclusions: Unstructured EHR data enabled the development of a predictive model for two-year CKD risk in persons with T2DM. Improving EHR data completeness remains essential to enhance future predictive modeling.
Authors
- karger, figshare admin ;
- J.F., NavarroGonzález ;
- L., PerezdeIsla ;
- G., CánovasMolina ;
- M.Á, Brito-Sanfiel ;
- D.E., BarajasGalindo ;
- L.Á, CuellarOlmedo ;
- D., Mauricio ;
- S., ToféPovedano ;
- J.A., BalsaBarro ;
- M., RubioAlmanza ;
- J.J., AparicioSánchez ;
- M., SequeraMutiozabal ;
- B., Pimentel ;
- A., PérezDomínguez ;
- C., Arias-Cabrales ;
- V., Fanjul ;
- A.J., Blanco-Carrasco ;
- J.F., MerinoTorres
Background: Persons with type 2 diabetes mellitus (T2DM) attending hospitals frequently experience major complications. We assessed the potential use of unstructured free-text data extracted from electronic health records (EHRs) using natural language processing (NLP) and machine learning (ML) to develop a predictive model for chronic kidney disease (CKD) in T2DM.Methods: This multicenter retrospective study included data from eight Spanish hospitals (2013–2018), extracted using NLP and ML techniques (EHRead®) based on SNOMED CT terminology. From a cohort of individuals with T2DM, we identified those with and without CKD at inclusion. Among individuals without CKD, we trained and validated a two-year predictive model for CKD development. The model showing the best balance between performance and clinical interpretability was selected for integration into a web-based tool to support early detection and risk stratification.Results: Of 588,786 individuals with T2DM, 316,597 were included for model development [training: 291,429 (92.1%); validation: 25,168 (7.9%); CKD incidence: 15.4% and 18.4%, respectively]. A high proportion of missing data was observed in key clinical variables. Among models evaluated, logistic regression (LR) achieved the best performance (AUC-ROC 0.72) using 27 predictors. Both a reduced 10-predictor model and a clinically refined 8-predictor model showed comparable performance to the full model in training and validation cohorts. The clinically refined model was selected for implementation in the web-based tool.Conclusions: Unstructured EHR data enabled the development of a predictive model for two-year CKD risk in persons with T2DM. Improving EHR data completeness remains essential to enhance future predictive modeling.
Authors
- karger, figshare admin ;
- J.F., NavarroGonzález ;
- L., PerezdeIsla ;
- G., CánovasMolina ;
- M.Á, Brito-Sanfiel ;
- D.E., BarajasGalindo ;
- L.Á, CuellarOlmedo ;
- D., Mauricio ;
- S., ToféPovedano ;
- J.A., BalsaBarro ;
- M., RubioAlmanza ;
- J.J., AparicioSánchez ;
- M., SequeraMutiozabal ;
- B., Pimentel ;
- A., PérezDomínguez ;
- C., Arias-Cabrales ;
- V., Fanjul ;
- A.J., Blanco-Carrasco ;
- J.F., MerinoTorres