BhasaBodh: Bridging Bangla Dialects and Romanized Forms through Machine Translation

View Dataset
Bhuiyan, Md Tofael Ahmed;Rahman, Md Abdur;Masum, Abdul Kadar Muhammad

Description

While machine translation has made significant strides for high-resource languages, many regional languages and their dialects, such as the Bangla variants Chittagong and Sylhet, remain underserved. Existing resources are often insufficient for robust sentence-level evaluation and overlook the widespread real-world practice of romanization, the common practice of typing native languages using the Latin script in digital communication. To address these gaps, we introduce BhasaBodh, a comprehensive benchmark for Bangla dialectal machine translation. We construct and release a sentence-level parallel dataset for Chittagong and Sylhet dialects aligned with Standard Bangla and English, create a novel romanized version of the dialectal data to facilitate evaluation in realistic multi-script scenarios, and provide the first comprehensive performance baselines by fine-tuning two powerful multilingual models, NLLB-200 and mBART-50, on seven distinct translation tasks. However, complex cross-lingual and cross-script translation remains a significant challenge. BhasaBodh lays the groundwork for future research in low-resource dialectal NLP, offering a valuable resource for developing more inclusive and practical translation systems.Citation: Bhuiyan, Md Tofael Ahmed, Md Abdur Rahman, and Abdul Kadar Muhammad Masum. "BhasaBodh: Bridging Bangla Dialects and Romanized Forms through Machine Translation." In Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025), pp. 113-118. 2025.

Citations (0)

Mentions (0)

Metrics

Dataset Index

0.4

FAIR Score

69%

Citations

0

Mentions

0

Metrics Over Time

Publication Details

DOI

Publisher

Mendeley Data

License

Creative Commons Attribution 4.0 International

Assigned Domain

Subfield

Communication

Field

Social Sciences

Domain

Social Sciences

Confidence Score

49%

Source

Scholar Data Model

Keywords

Natural Language ProcessingMachine Translation

Normalization Factors

FT

63.46

CTw

1.00

MTw

1.00