Turkic basic vocabularies

Savelyev, Alexander;Robbeets, Martine

Description

The dataset represents basic vocabulary data across 32 Turkic languages. The basic vocabulary list merges the Leipzig-Jakarta 200 list (Haspelmath and Tadmor 2009) with the Jena 200 list (Anderson and Heggarty n.d.) and contains 254 different concepts. For each word in the dataset we provide an etymological analysis to establish cognacy classes on the basis of regular sound correspondences. Borrowings that can be identified using clearcut historical comparative criteria are excluded to provide a clearer phylogenetic signal. We deal with cases of synonymy in that we allow more than one word with a certain basic meaning in our dataset unless there is evidence that it is less basic than one of its synonyms. Singletons are removed from the dataset in case they have a non-singleton synonym that fits the criteria for basic status. The dataset also contains the tsv file edited in the EDICTOR tool (List 2017) as well as the file fed to BEAST in the original nexus format and in the XML format required by BEAST.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.4

FAIR Score

77%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Zenodo

License

Creative Commons Attribution 4.0 International

Open Access

Assigned Domain

Subfield

Language and Linguistics

Field

Arts and Humanities

Domain

Social Sciences

Confidence Score

100%

Source

Open Alex

Keywords

Turkic languages, genealogical classification, Proto-Turkic, Bayesian phylogenetic linguistics

Normalization Factors

FT

57.69

CTw

1.00

MTw

1.00