BanglaBlend: A Large-Scale Nobel Dataset of Bangla Sentences Categorized by Saint(Sadhu) and Common(Cholito) Form of Bengali Language

View Dataset
ayman, Umme Ayman;Saha, Chayti ;Mawa, Zannatul

Description

This BanglaBlend dataset is a comprehensive collection of Bangla (Bengali) sentences meticulously categorized based on two specific forms: Saint(Sadhu) and Common(Cholito). This dataset is comprised of a total 7350 annotated Bangla sentences as well as it is preprocessed dataset where several data preprocessing techniques have been applied. This dataset is designed to facilitate research and development in natural language processing (NLP) and computational linguistics, particularly for Bangla, a widely spoken language in Bangladesh and parts of India. Specially, this dataset can be leveraged for several natural language processing task such as text summarization, text classification, sentiment analysis, automatic language translation.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.9

FAIR Score

65%

Citations

1

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Mendeley Data

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Language and Linguistics

Field

Arts and Humanities

Domain

Social Sciences

Confidence Score

52%

Source

Scholar Data Model

Keywords

Data ScienceNatural Language ProcessingMachine LearningBengali LanguageSentence Processing

Normalization Factors

FT

43.27

CTw

1.00

MTw

1.00