Automated Organization ProfileSt. Petersburg Department of the Steklov Mathematical Institute
St. Petersburg Department of the Steklov Mathematical Institute
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets in this organization
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the organization's datasets
Total Mentions
Total mentions of the organization's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 3.8 (sum of 2 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
Text mining of scientific libraries and social media has already proven itself as a reliable tool for
drug repurposing and hypothesis generation. The task of mapping a disease mention to a concept
in a controlled vocabulary, typically to the standard thesaurus in the Unified Medical Language
System (UMLS), is known as medical concept normalization. This task is challenging due to the
differences in medical terminology between health care professionals and social media texts coming
from the lay public. To bridge this gap, we use sequence learning with recurrent neural networks
and semantic representation of one- or multi-word expressions: we develop end-to-end architectures
directly tailored to the task, including bidirectional Long Short-Term Memory and Gated Recurrent
Units with an attention mechanism and additional semantic similarity features based on UMLS.
Our evaluation over a standard benchmark shows that recurrent neural networks improve results
over an effective baseline for classification based on convolutional neural networks. A qualitative
examination of mentions discovered in a dataset of user reviews collected from popular online health
information platforms as well as quantitative evaluation both show improvements in the semantic
representation of health-related expressions in social media.
Authors
- Tutubalina, Elena ;
- Miftahutdinov, Zulfat ;
- Nikolenko, Sergey ;
- Malykh, Valentin
Text mining of scientific libraries and social media has already proven itself as a reliable tool for
drug repurposing and hypothesis generation. The task of mapping a disease mention to a concept
in a controlled vocabulary, typically to the standard thesaurus in the Unified Medical Language
System (UMLS), is known as medical concept normalization. This task is challenging due to the
differences in medical terminology between health care professionals and social media texts coming
from the lay public. To bridge this gap, we use sequence learning with recurrent neural networks
and semantic representation of one- or multi-word expressions: we develop end-to-end architectures
directly tailored to the task, including bidirectional Long Short-Term Memory and Gated Recurrent
Units with an attention mechanism and additional semantic similarity features based on UMLS.
Our evaluation over a standard benchmark shows that recurrent neural networks improve results
over an effective baseline for classification based on convolutional neural networks. A qualitative
examination of mentions discovered in a dataset of user reviews collected from popular online health
information platforms as well as quantitative evaluation both show improvements in the semantic
representation of health-related expressions in social media.
Authors
- Tutubalina, Elena ;
- Miftahutdinov, Zulfat ;
- Nikolenko, Sergey ;
- Malykh, Valentin