Research
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
Overview Research area: Natural Language Processing — low-resource and endangered language resources, multilingual machine translation, culturally grounded benchmarking, and community-based data colle

- arXiv
- 2512.01077
- Published
- 2025-11-30
- Authors
- Neha Joshi, Pamir Gogoi, Aasim Mirza, Aayush Jansari, Aditya Yadavalli, Ayushi Pandey, Arunima Shukla, Deepthi Sudharsan, Kalika Bali, Vivek Seshadri
AI summary
Overview
Research area: Natural Language Processing — low-resource and endangered language resources, multilingual machine translation, culturally grounded benchmarking, and community-based data collection.
Technical level: Intermediate. The dataset construction and community-governance discussion is accessible to beginners; the translation evaluation setup (model selection, judge ensembles, Likert-scale rubric) assumes some familiarity with MT evaluation practice.
Scope: The paper introduces ELR-1000, a multimodal corpus of 1,060 recipes in 10 endangered Eastern Indic languages, and uses it to evaluate six state-of-the-art LLMs on recipe translation into English.
What This Paper Is About
Indic-language NLP has largely digitized India's constitutionally recognized, high-resource languages, leaving tribal and indigenous languages without digital resources. The authors crowdsource traditional recipes from rural speakers of 10 endangered languages as a culturally grounded, multimodal dataset, and then test whether current LLMs can translate those recipes into English without erasing cultural meaning. The goal is both to release a benchmark for underrepresented languages and to show what breaks when general-purpose models encounter culturally specific content.
Key Contributions
-
Release of ELR-1000: 1,060 recipes in 10 endangered languages of Eastern India (Bodo, Assamese, Meitei, Kaman-Mishmi, Khortha, Santhali, Ho, Sadri, Mundari, and Khasi), released under the Karya Public License (KPL), with text, image, and audio modalities.
-
A parallel translation subset: A representative subset of the corpus was manually translated into English by native speakers and released as a parallel corpus for LLM-enabled translation evaluation.
-
An LLM translation benchmark with two conditions: Six state-of-the-art LLMs were evaluated on translating traditional recipes both with no context and with targeted context (language background, few-shot examples, cultural-preservation guidelines, and recipe-structure instructions).
-
Documentation of collection challenges and ethics: The paper records the practical difficulties of community-rooted data collection — trust building, seasonality of foraged ingredients, fair payment across recipes of varying complexity, and quality control across languages — as a guide for similar efforts.
Main Findings
-
Context is the decisive factor: Providing context dramatically improved translation quality across all models. For Mistral, context was described as the difference between nonsensical output (scoring 1.0) and usable translations (scoring 4.0 or higher).
-
Gemini 2.5 Flash was the top performer: It achieved the highest average scores in both conditions, frequently receiving perfect or near-perfect ratings in the contextual setting, with particular strength in Santhali, Meitei, and Assamese translations.
-
A second performance tier emerged: Claude Sonnet 4, GPT-4o, and Llama-4-Scout-17B improved substantially with context. Llama-4-Scout was noted for Sadri and Ho, and Claude Sonnet 4 for Khortha and Ho.
-
"Fluent falsehood": Across nearly all models and languages, Fluency and Comprehensibility scores consistently exceeded Adequacy and Cultural Appropriateness scores — models produced readable English that often did not match the source meaning.
-
Hallucination on the hardest languages: For Bodo and Kaman Mishmi, models systematically generated entirely different recipes, replacing ingredients such as silkworms or regional vegetables with mushrooms, chicken curry, or unrelated dishes.
-
Systematic ingredient and tool errors: "Jhingi" was consistently mistranslated as "prawns/shrimp." Models omitted indigenous implements such as the mortar and pestle, and Gemini introduced "chopping board" in both contextual and non-contextual outputs — a tool the authors state is almost never used in indigenous kitchens.
-
Language-level variation: Translation quality was notably better for Khortha and Sadri, while Kaman Mishmi, Bodo, Ho, Santhali, and Mundari prompted near-complete translation failures across most models.
-
NMT results were weak: For Assamese, the one relatively resource-rich language in the set supported by leading NMT systems, both BLEU and chrF scores were significantly lower than expected, which the authors attribute to the specialized recipe domain and limited culturally specific training data.
-
Human translations remained superior: Even with context, human reference translations preserved cultural nuance better than any model.
-
Multimodal contribution patterns: 54% of recipe steps included all three modalities, 29% used image-text, 83.4% had featured text, and 64.5% included audio narration.
Methodology in Plain English
The authors began with a pilot in the Sadri-speaking community, running a demonstration session with approximately 30 Sadri-speaking women from 3 remote tribal villages in Jharkhand, each of whom recorded one recipe. Based on that success, they selected 10 languages from the UNESCO Endangered Languages list — five mainly spoken in Jharkhand and Bihar and the rest from Northeast India — and aimed to balance higher- and lower-resource languages.
One local coordinator was recruited per language through local NGOs. Coordinators mobilized 30–50 rural women per language (aged 15–45), and Karya ran in-person training and app demonstrations. Participants used a mobile application designed for first-time digital workers: minimal text, icons and audio instructions in local languages, the ability to review and edit entries before submission, separated photo capture and annotation steps, and offline data entry for poor connectivity. Contributors were paid $8.68 per recipe and typically submitted two to five recipes, earning between $17.50 and $43.50. Local coordinators validated each submission, and project managers conducted spot checks across languages.
Because contributors documented recipes in their own style, the raw corpus was structurally heterogeneous. The authors normalized it into a modular, array-based structure separating text, image, and audio content stored in parallel directories.
For evaluation, they selected three proprietary models (Gemini 2.5 Flash, GPT-4o, Claude Sonnet 4) and three open-source models (Llama 4 Scout 17B-16E, Mistral Small 3.1 (25.03), CohereLabs Aya Expanse 8B). Each model translated recipes under two conditions: no context (source recipe only) and contextual (language background, few-shot examples, cultural-preservation guidelines, and recipe-structure instructions). Human translations from native speakers served as references. Scoring used a 5-point Likert scale across four dimensions — Adequacy, Fluency, Comprehensibility, and Cultural and Contextual Appropriateness. The paper states that Gemini 2.5 Flash was used as the evaluation judge on default settings, and also describes a two-judge LLM ensemble of Gemini 2.5 Pro and OpenAI GPT-5 for automated evaluation against human references, followed by human oversight on a sample of evaluations, especially those with score discrepancies.
Why This Matters
Research impact. The paper argues that standard evaluation metrics are inadequate for endangered-language translation, and that translation failures in culturally specific domains are not merely linguistic but epistemic — models substitute globally dominant tools, practices, and ingredients for indigenous ones. It calls for culturally grounded datasets and evaluation frameworks rather than fluency-based scoring alone.
Real-world applications:
- Building translation and speech tools that serve speakers of non-standardized, endangered languages excluded from current digital services.
- Creating cultural archives of oral culinary, agricultural, and medicinal knowledge before it disappears from community memory.
- Informing equitable data-collection and compensation practices for crowdsourcing with low-digital-literacy communities.
- Providing a benchmark for developing culturally aware evaluation methods beyond surface-level fluency.
Industry relevance. The findings are directly relevant to teams deploying multilingual translation, speech, and question-answering systems in India and similar multilingual contexts, where silent failures — fluent but inaccurate outputs — can misrepresent content rather than simply refuse it. The paper's central claim is that enabling models to move from syntactic fluency to contextual fidelity requires both culturally grounded data and human oversight.
Future Directions
-
Scale language coverage: The authors state they want to extend beyond the 10 endangered languages to cover more of the linguistic diversity present in Eastern India.
-
Grow the dataset beyond benchmarking: They note that roughly 100 recipes per language may benchmark existing LLMs but is not enough to improve them, and they plan to crowdsource more recipes with rural communities.
-
Broaden topics beyond cuisine: Suggested expansion into agricultural and livestock farming practices to make the benchmark more valuable.
-
Probe cultural awareness directly: The paper proposes moving beyond surface evaluation toward question-answering tasks or internal probing methods that examine how cultural concepts are represented in model embeddings.
-
Address the alignment limitation: The authors note the dataset is multimodal (audio, text, images) but is not aligned to support benchmarking for advanced tasks such as knowledge graph construction or multimodal reasoning.
Target Audience
NLP researchers working on low-resource and endangered languages, machine translation evaluation, and culturally grounded benchmarks; practitioners building multilingual products for Indic and other underrepresented language communities; and researchers or organizations designing community-based, ethically compensated data collection with rural or low-digital-literacy participants.
Authors’ abstract
We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered languages. These recipes, rich in linguistic and cultural nuance, were collected using a mobile interface designed for contributors with low digital literacy. Endangered Language Recipes (ELR)-1000 -- captures not only culinary practices but also the socio-cultural context embedded in indigenous food traditions. We evaluate the performance of several state-of-the-art large language models (LLMs) on translating these recipes into English and find the following: despite the models' capabilities, they struggle with low-resource, culturally-specific language. However, we observe that providing targeted context -- including background information about the languages, translation examples, and guidelines for cultural preservation -- leads to significant improvements in translation quality. Our results underscore the need for benchmarks that cater to underrepresented languages and domains to advance equitable and culturally-aware language technologies. As part of this work, we release the ELR-1000 dataset to the NLP community, hoping it motivates the development of language technologies for endangered languages.