300-Dimensional Word Embeddings for Nepali Language

Lamsal, Rabindra

Description

This pre-trained Word2Vec model has 300-dimensional vectors for more than 0.5 million Nepali words and phrases. A separate Nepali language text corpus was created using the news contents freely available in the public domain. The text corpus contained more than 100 million running words.Word2Vec model details: Embeddings Dimension: 300, Architecture: Continuous - BOW, Training algorithm: Negative sampling = 15, Context (window) size: 10, Token minimum count: 2, Encoded in UTF-8.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.3

FAIR Score

58%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

IEEE DataPort

License

Creative Commons Attribution

Assigned Domain

Subfield

Language and Linguistics

Field

Arts and Humanities

Domain

Social Sciences

Confidence Score

43%

Source

Scholar Data Model

Keywords

OtherNepali Word2VecNepali Word Embeddings

Normalization Factors

FT

57.69

CTw

1.00

MTw

1.00