Version V1

A Chinese-Uighur comparable corpus

Tao, Feng;Miao, Li;Yichao, Cao;Weihui, Zeng

Description

The dataset is composed of comparable corpus of Chinese and Uigur, obtained from the Internet. Chinese and Uigur language pairs are textually corresponding. The dataset is mainly from news, including news headlines, time and text. The dataset contains two data files: ch_corpus.zip and uy_corpus.zip. Each package contains four documents, namely document_1, document_2, document_3 and document_4. Each document contains two folders: uy and ch, where uy represents Uyghur, ch represents Chinese, and each folder contains multiple text documents. Uighur and Chinese language pairs are organized correspondingly according to their names.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.1

FAIR Score

15%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Science Data Bank

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Artificial Intelligence

Field

Computer Science

Domain

Physical Sciences

Confidence Score

40%

Source

Scholar Data Model

Keywords

Information science and systems sciencecorpus constructioncomparable corpusChinese- Uighurdata mining

Normalization Factors

FT

57.69

CTw

1.00

MTw

1.00