Automated Author ProfileWeng, Yufei
University of California, San Diego
Weng, Yufei
Current S-Index
Sum of Dataset Indices for all datasets
Average Dataset Index per Dataset
Average Dataset Index per dataset
Total Datasets
Total datasets for this author
Average FAIR Score
Average FAIR Score per dataset
Total Citations
Total citations to the author's datasets
Total Mentions
Total mentions of the author's datasets
S-Index Interpretation
The S-Index (Sharing Index) is a comprehensive metric that represents the cumulative impact of all your datasets. It is calculated as the sum of Dataset Index scores across all your claimed datasets.
What it means:
- A higher S-index indicates greater overall impact of your datasets relative to typical datasets in their fields of research
- The S-Index grows as you add more datasets or as existing datasets gain more citations and mentions
- It provides a single number to track your research data impact over time
Current S-Index: 0.5 (sum of 3 datasets Dataset Index scores)
More information here.
S-Index Over Time
Cumulative Citations Over Time
Cumulative Mentions Over Time
Datasets
OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript DataThe Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents. Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.
Authors
- Parkes Allen, Jonathan ;
- Mullan, John ;
- Nigst, Lorenz ;
- Barber, Mathew ;
- Shahid Khan, Taimoor ;
- Seydi, Masoumeh ;
- Chen, Danlu ;
- Weng, Yufei ;
- Vogler, Nikolai ;
- Murel, Jacob ;
- Eshera, Osama ;
- Berg-Kirkpatrick, Taylor ;
- Smith, David ;
- Bowen Savant, Sarah ;
- Thomas Miller, Matthew ;
- Verkinderen, Peter
OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript DataThe Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents. Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.
Authors
- Parkes Allen, Jonathan ;
- Mullan, John ;
- Nigst, Lorenz ;
- Barber, Mathew ;
- Shahid Khan, Taimoor ;
- Seydi, Masoumeh ;
- Chen, Danlu ;
- Weng, Yufei ;
- Vogler, Nikolai ;
- Murel, Jacob ;
- Eshera, Osama ;
- Berg-Kirkpatrick, Taylor ;
- Smith, David ;
- Bowen Savant, Sarah ;
- Thomas Miller, Matthew ;
- Verkinderen, Peter
OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript DataThe Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents.Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.
Authors
- Parkes Allen, Jonathan ;
- Mullan, John ;
- Nigst, Lorenz ;
- Barber, Mathew ;
- Shahid Khan, Taimoor ;
- Seydi, Masoumeh ;
- Chen, Danlu ;
- Weng, Yufei ;
- Vogler, Nikolai ;
- Murel, Jacob ;
- Eshera, Osama ;
- Berg-Kirkpatrick, Taylor ;
- Smith, David ;
- Bowen Savant, Sarah ;
- Thomas Miller, Matthew ;
- Verkinderen, Peter