Research
Leveraging the Cross-Domain & Cross-Linguistic Corpus for Low Resource NMT: A Case Study On Bhili-Hindi-English Parallel Corpus
Overview Research area: Low-resource neural machine translation (NMT) and multilingual large language models (LLMs), focused on an endangered Indic tribal language (Bhili). Technical level: Intermedia
- arXiv
- 2511.00486
- Published
- 2025-11-01
- Authors
- Pooja Singh, Shashwat Bhardwaj, Vaibhav Sharma, Sandeep Kumar
AI summary
Overview
- Research area: Low-resource neural machine translation (NMT) and multilingual large language models (LLMs), focused on an endangered Indic tribal language (Bhili).
- Technical level: Intermediate. Readers will get the most from this paper with familiarity with machine translation, fine-tuning versus in-context learning, and MT evaluation metrics, though the writing is largely accessible.
- Scope: The paper introduces a 110,000-sentence Bhili-Hindi-English parallel corpus, benchmarks a wide range of open-source and proprietary models on four translation directions, and studies cross-domain generalization and distributional divergence.
What This Paper Is About
Bhili, a Western Indo-Aryan language written in Devanagari and spoken by roughly 13 million people, has almost no digitized parallel data, so no effective machine translation system exists for it. The authors build the Bhili-Hindi-English Parallel Corpus (BHEPC) by having professional translators render Hindi source sentences into Bhili, then use it to test whether existing multilingual models can translate into and out of Bhili under fine-tuning, in-context learning, and cross-domain conditions. The goal is both to fill a resource gap for one language and to demonstrate a reusable corpus-construction workflow for other under-documented languages.
Key Contributions
- The BHEPC corpus. The authors introduce what they describe as the first and largest parallel corpus worldwide for Bhili: 110,000 sentences across Bhili, Hindi, and English, curated with expert human translators and spanning education, administration, and news domains (split into 1,08,000 training, 1,000 validation, and 1,000 test sentences).
- A broad Bhili MT benchmark. They evaluate mT5, Qwen3, DeepSeek-V3, Gemma-2-9B, Mistral-7B-v0.1, BLOOMZ, Llama-2, Llama-3, Llama-4-Scout-17B-16E, Gemini 2.0 Flash, Gemini 2.5 Flash, GPT-3.5 Turbo, GPT-4o-0513, and GPT-4.5, plus IndicTrans2 and NLLB-200, across Hindi↔Bhili and English↔Bhili.
- Cross-domain generalization analysis. They quantify distributional divergence between training and test domains using Jensen-Shannon Divergence (JSD) and measure how domain shift degrades translation quality.
- A scalable data-creation template. They propose and describe a hybrid seed-corpus-plus-post-editing workflow that combines model-assisted generation with native speaker post-editing to reduce manual translation effort, framed as adaptable to other endangered languages.
Main Findings
- Fine-tuned NLLB-200 (600M) is the strongest model overall. NLLB-200 distilled 600M achieves the best fine-tuned results in every direction: spBLEU 11.30 / chrF++ 42.27 (Hindi→Bhili), 37.59 / 60.62 (Bhili→Hindi), 7.85 / 35.18 (English→Bhili), and 27.00 / 53.00 (Bhili→English).
- mT5 is a close second. mT5 base reaches spBLEU 11.68 / chrF++ 42.83 on Hindi→Bhili, slightly ahead of NLLB-200 on that pair, and mT5 small also generalizes well. Outside Hindi→Bhili, NLLB-200 significantly outperforms all other models at p < 0.005 under paired bootstrap resampling.
- IndicTrans2 underperforms despite Indic pretraining. It scores only spBLEU 9.29 / chrF++ 35.67 on Hindi→Bhili, which the authors attribute to specialization and overfitting on its 22 pre-training languages and poor generalization to Bhili.
- Larger decoder-only models do not automatically win. Mixtral-7B-v0.1, DeepSeek-V3, Qwen3-8B, and Gemma-2-9B show limited fine-tuned performance, while Llama-3-8B and BLOOMZ-7B1 do relatively better, which the authors read as evidence that multilingual pretraining quality matters more than sheer model size.
- Proprietary models dominate in-context learning. GPT-4.5 is the strongest ICL model, scoring chrF++ 27.47 / 29.75 / 31.23 on Hindi→Bhili, 45.72 / 46.68 / 49.65 on Bhili→Hindi, 23.35 / 25.67 / 27.12 on English→Bhili, and 42.16 / 43.78 / 45.16 on Bhili→English at 0, 5, and 10 in-context examples respectively. GPT-4o-0513, GPT-3.5 Turbo, and Gemini 2.5 Flash follow in that order.
- Gains from more in-context examples are uneven. Open-source models improve substantially from 0 to 5 shots, but improvements from 5 to 10 shots are marginal, especially for Hindi→Bhili and English→Bhili.
- Translating out of Bhili is easier than translating into it. chrF++ scores are higher when Bhili is the source language. Gemma-2-9B reaches 30.50 chrF++ on Bhili→English at 0 shots, while Llama-4-Scout-17B-16E scores only 6.35 on English→Bhili at 0 shots.
- Mixed open-source winners across directions and shot counts. Llama-4-Scout-17B-16E is the best open-source model for English→Bhili at all shot counts and for Hindi→Bhili at 0 shots; Gemma leads at 5 shots; BLOOMZ-560M is better at 10 shots. BLOOMZ-3.1B and BLOOMZ-7B1 lead on Bhili→Hindi.
- In-domain fine-tuning beats cross-domain fine-tuning. For Bhili→Hindi, a model fine-tuned on Govt/PMI reaches chrF++ 60.09 in-domain but drops to 33.40 when tested on NCERT. The single highest reported cell in the cross-domain table is 87.35 chrF++ for the NCERT-fine-tuned model tested in-domain on English→Bhili.
- Domain similarity tracks quality. JSD heatmaps show that models fine-tuned on domains with lower JSD to the test domain translate better, while highly divergent corpora lead to poor adaptation.
- chrF++ aligns with human judgment better than spBLEU. Segment-level correlations with human MQM scores are consistently higher for chrF++ across all four directions, but both metrics correlate weakly when Bhili is the target (Hindi→Bhili: chrF++ τ = 0.20, ρ = 0.30; English→Bhili: τ = 0.15, ρ = 0.22).
- Human annotation was reliable. Reported inter-annotator agreement coefficients were 0.60 (Hindi→Bhili), 0.66 (Bhili→Hindi), 0.53 (English→Bhili), and 0.57 (Bhili→English), with 0.59 for the manually translated Hindi→Bhili gold references.
- Error analysis of the best model reveals four recurring failure types. Language mixing (Gujarati words and verb inflections due to lexical overlap), hallucination and omission, polysemy and lexical ambiguity, and failure on domain-specific terminology and formal registers.
- Script similarity does not guarantee transfer. Because Bhili shares Devanagari with Hindi yet still degrades sharply in generation, the authors argue that script-level similarity alone is insufficient for semantic transfer.
Methodology in Plain English
The authors first assembled Hindi source sentences from existing resources: the Bharat Parallel Corpus Collection (BPCC), publicly available government documents including Legislative Assembly Speeches, the PMIndia corpus, and NCERT textbooks. A team of on average 10 professional translators, working from May 2024 to March 2025 for a total of 27,000 hours, translated this Hindi data into Bhili. The English side was not manually translated; it was generated by running the Hindi sentences through the IndicTrans2 model.
Preprocessing removed noise and duplication: Bhili homophones were normalized, extraneous characters deleted, English lowercased, near-identical sentences de-duplicated (including cosine-similarity filtering of near-identical source-target pairs), and sentences shorter than 6 words or longer than 80 words discarded. Tokenization used a SentencePiece model. Personal information, hate speech, and redundancy were screened out before segmentation.
For evaluation, the team built a balanced test set through stratified sampling, plus a domain-specific benchmark of 288 NCERT sentences, 487 Government/PMI sentences, and 1,063 mass-media sentences with gold translations. The remaining corpus was split into training and validation at a 99:1 ratio while preserving domain proportions.
Models were tested in three ways: fine-tuning (both full fine-tuning and LoRA-based parameter-efficient fine-tuning, with hyperparameters chosen by grid search over batch sizes 8/16/32 and learning rates from 5e-3 down to 1e-5), in-context learning at 0, 5, and 10 shots using a uniform prompt and a decoding temperature of 0.1, and cross-domain transfer where training and testing domains differ. Performance was measured with chrF++ and spBLEU on the automatic side. Separately, eight models' outputs were judged by bilingual native-speaker experts using MQM and Direct Assessment guidelines over 250 segments per direction, yielding 1,000 annotated segments. Domain divergence between corpora was measured with Jensen-Shannon Divergence, computed over token frequency distributions built with bert-base-multilingual-cased. All experiments ran on NVIDIA A100 GPUs, with training times between 6 and 48 hours depending on model size.
Why This Matters
This work addresses a resource gap for a language spoken by about 13 million people that is absent from IndicTrans2, NLLB, mT5, the Llama, Gemma, Mixtral, DeepSeek, Qwen3, and BLOOMZ families, and also from commercial systems such as Google Translate and Microsoft Translator. It provides both the data and a measured baseline that future work on Bhili can build on, and it challenges the common assumption that sharing a script with a high-resource language (Hindi, in Devanagari) yields useful transfer.
Real-world applications:
- Government service delivery. The corpus deliberately covers administration and legislative speech, supporting translation of policy directives, welfare schemes, and official communication into Bhili.
- Education. NCERT textbook material is one of the corpus domains, which could feed translated or bilingual learning materials for Bhili-speaking students.
- News and mass media. Mass media is the largest of the three domains in the corpus (64k fine-tuning sentences), relevant to local-language journalism and information access.
- Cultural and linguistic preservation. The corpus is described as deliberately encoding orthography, idioms, and community practices, making it a cultural record as well as an NLP resource.
Industry relevance: The results give a practical decision rule for teams with limited budgets. Fine-tuning a compact model such as NLLB-200 distilled 600M outperforms much larger general-purpose models, while in-context learning with large proprietary models is a competitive alternative that avoids expensive fine-tuning. The paper suggests a hybrid strategy: fine-tune smaller models and use ICL for larger ones. The cross-domain results also warn that a system trained on one domain (for example, news) should not be deployed on another (for example, government text) without expectation of a large quality drop.
Future Directions
- Unsupervised and semi-supervised methods. The authors state that the scarcity of monolingual Bhili data limited them to supervised fine-tuning, leaving unsupervised and semi-supervised approaches untested.
- Data augmentation. Back-translation and pivot-based transfer were not experimented with, and the authors identify them as effective techniques in other low-resource NMT settings that remain to be tried here.
- Scalability of corpus creation. The manual effort behind a 110k-sentence corpus raises concerns about extending the approach to the thousands of other low-resource languages. The proposed seed-corpus-plus-post-editing workflow is the authors' partial answer, but they describe broader generalizability and cross-domain robustness as open challenges.
- Purpose-built models for low-resource languages. Since script-level similarity did not produce reliable transfer, the paper argues for dedicated models tailored to languages like Bhili rather than relying on existing multilingual systems.
Target Audience
Researchers and practitioners in low-resource and multilingual machine translation; NLP groups working on Indic languages; teams building datasets for endangered or under-documented languages; and linguists or community organizations interested in a documented workflow for creating gold-standard parallel corpora with native-speaker translators. Policy and public-sector technology groups concerned with digital inclusion for tribal language communities will also find the resource and benchmark results relevant.
Authors’ abstract
The linguistic diversity of India poses significant machine translation challenges, especially for underrepresented tribal languages like Bhili, which lack high-quality linguistic resources. This paper addresses the gap by introducing Bhili-Hindi-English Parallel Corpus (BHEPC), the first and largest parallel corpus worldwide comprising 110,000 meticulously curated sentences across Bhili, Hindi, and English. The corpus was created with the assistance of expert human translators. BHEPC spans critical domains such as education, administration, and news, establishing a valuable benchmark for research in low resource machine translation. To establish a comprehensive Bhili Machine Translation benchmark, we evaluated a wide range of proprietary and open-source Multilingual Large Language Models (MLLMs) on bidirectional translation tasks between English/Hindi and Bhili. Comprehensive evaluation demonstrates that the fine-tuned NLLB-200 distilled 600M variant model outperforms others, highlighting the potential of multilingual models in low resource scenarios. Furthermore, we investigated the generative translation capabilities of multilingual LLMs on BHEPC using in-context learning, assessing performance under cross-domain generalization and quantifying distributional divergence. This work bridges a critical resource gap and promotes inclusive natural language processing technologies for low-resource and marginalized languages globally.