Automated Author Profile

Weng, Yufei

University of California, San Diego

Current S-Index

0.5

Sum of Dataset Indices for all datasets

Average Dataset Index per Dataset

0.2

Average Dataset Index per dataset

Total Datasets

3

Total datasets for this author

Average FAIR Score

84.6%

Average FAIR Score per dataset

Total Citations

0

Total citations to the author's datasets

Total Mentions

0

Total mentions of the author's datasets

S-Index Interpretation

S-Index Over Time

Cumulative Citations Over Time

Cumulative Mentions Over Time

Datasets

OpenITI MAKHZAN (Version: 2026.1.2)

OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript DataThe Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents. Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.

Authors

  • Parkes Allen, Jonathan ;
  • Mullan, John ;
  • Nigst, Lorenz ;
  • Barber, Mathew ;
  • Shahid Khan, Taimoor ;
  • Seydi, Masoumeh ;
  • Chen, Danlu ;
  • Weng, Yufei ;
  • Vogler, Nikolai ;
  • Murel, Jacob ;
  • Eshera, Osama ;
  • Berg-Kirkpatrick, Taylor ;
  • Smith, David ;
  • Bowen Savant, Sarah ;
  • Thomas Miller, Matthew ;
  • Verkinderen, Peter
0 Citations0 Mentions85% FAIR0.4 Dataset Index
10.5281/zenodo.164101052026

OpenITI MAKHZAN (Version: 2026.1.2)

OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript DataThe Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents. Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.

Authors

  • Parkes Allen, Jonathan ;
  • Mullan, John ;
  • Nigst, Lorenz ;
  • Barber, Mathew ;
  • Shahid Khan, Taimoor ;
  • Seydi, Masoumeh ;
  • Chen, Danlu ;
  • Weng, Yufei ;
  • Vogler, Nikolai ;
  • Murel, Jacob ;
  • Eshera, Osama ;
  • Berg-Kirkpatrick, Taylor ;
  • Smith, David ;
  • Bowen Savant, Sarah ;
  • Thomas Miller, Matthew ;
  • Verkinderen, Peter
0 Citations0 Mentions85% FAIR0.4 Dataset Index
10.5281/zenodo.198619122026

OpenITI MAKHZAN (Version: 2025.1.1)

OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript DataThe Open Islamicate Texts Initiative (OpenITI) was founded in 2017 with the goal of building the digital infrastructure for the study of the premodern Islamic world. The underperformance of automatic transcription via Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) for Arabic-script languages long stymied the production of large-scale corpora in Islamicate languages, such as Persian, Arabic, Ottoman Turkish, and Urdu.The OpenITI MAKHZAN dataset is a large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu print and handwritten (manuscript) documents.Note on Release Numbering: Version 2025.1.2—where 2025 is the year of the release, the first dotted number—.1—is the ordinal release number in 2025, and the second dotted number—.2—is the overall release number; the first dotted number will reset every year, while the second one will continue on increasing.

Authors

  • Parkes Allen, Jonathan ;
  • Mullan, John ;
  • Nigst, Lorenz ;
  • Barber, Mathew ;
  • Shahid Khan, Taimoor ;
  • Seydi, Masoumeh ;
  • Chen, Danlu ;
  • Weng, Yufei ;
  • Vogler, Nikolai ;
  • Murel, Jacob ;
  • Eshera, Osama ;
  • Berg-Kirkpatrick, Taylor ;
  • Smith, David ;
  • Bowen Savant, Sarah ;
  • Thomas Miller, Matthew ;
  • Verkinderen, Peter
0 Citations0 Mentions85% FAIR0.5 Dataset Index
10.5281/zenodo.164101062025