BhasaBodh: Bridging Bangla Dialects and Romanized Forms through Machine Translation
View DatasetDescription
While machine translation has made significant strides for high-resource languages, many regional languages and their dialects, such as the Bangla variants Chittagong and Sylhet, remain underserved. Existing resources are often insufficient for robust sentence-level evaluation and overlook the widespread real-world practice of romanization, the common practice of typing native languages using the Latin script in digital communication. To address these gaps, we introduce BhasaBodh, a comprehensive benchmark for Bangla dialectal machine translation. We construct and release a sentence-level parallel dataset for Chittagong and Sylhet dialects aligned with Standard Bangla and English, create a novel romanized version of the dialectal data to facilitate evaluation in realistic multi-script scenarios, and provide the first comprehensive performance baselines by fine-tuning two powerful multilingual models, NLLB-200 and mBART-50, on seven distinct translation tasks. However, complex cross-lingual and cross-script translation remains a significant challenge. BhasaBodh lays the groundwork for future research in low-resource dialectal NLP, offering a valuable resource for developing more inclusive and practical translation systems.Citation: Bhuiyan, Md Tofael Ahmed, Md Abdur Rahman, and Abdul Kadar Muhammad Masum. "BhasaBodh: Bridging Bangla Dialects and Romanized Forms through Machine Translation." In Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025), pp. 113-118. 2025.
Citations (0)
No citations found
Mentions (0)
No mentions found
Metrics Over Time
Publication Details
Subfield
Communication
Field
Social Sciences
Domain
Social Sciences
Confidence Score
49%
Source
Scholar Data Model