Skip to content
AI.info

Research

HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples

Overview Research area: Multilingual and multimodal natural language processing — specifically benchmarking vision-language models (VLMs) in Indian languages against English. Technical level: Intermed

HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples
arXiv
2511.15183
Published
2025-11-19
Authors
Rishikant Chigrupaatii, Ponnada Sai Tulasi Kanishka, Lalit Chandra Routhu, Martin Patel Sama Supratheek Reddy, Divyam Gupta, Dasari Srikar, Krishna Teja Kuchimanchi, Rajiv Misra, Rohun Tripathi

AI summary

Overview

  • Research area: Multilingual and multimodal natural language processing — specifically benchmarking vision-language models (VLMs) in Indian languages against English.
  • Technical level: Intermediate. Readers should be comfortable with VQA benchmarks, translation pipelines and LLM evaluation methodology, but the paper is written accessibly.
  • Scope: The paper proposes a semi-automated framework for building human-verified, culturally grounded VQA benchmarks, uses it to construct HinTel-AlignBench for Hindi and Telugu with English-aligned samples, and reports a large evaluation of open-weight and proprietary VLMs on it.

What This Paper Is About

India has enormous linguistic diversity, but most state-of-the-art vision-language models are English-centric, and existing multilingual evaluations for Indian languages are small, narrow in scope, built on unverified machine translations, or lacking cultural grounding. The paper's goal is to build a scalable, human-verified framework for creating VQA benchmarks in Indian languages and to release a Hindi-Telugu benchmark with English-aligned samples so that a model's performance can be compared exactly across languages on the same questions and images. Using that benchmark, the authors quantify how much current VLMs regress when moving from English to Hindi or Telugu.

Key Contributions

  1. A framework. A semi-automated, scalable pipeline combining machine translation, back-translation-based filtering, text-only LLM question generation, and human verification for producing multilingual vision-language evaluation sets.
  2. A benchmark. HinTel-AlignBench, described as the most comprehensive vision-language benchmark for Hindi and Telugu, mixing adapted English datasets (VQAv2, RealWorldQA, CLEVR-Math) with native Indic datasets (JEE-Vision for STEM, VAANI for cultural grounding), with approximately 4k QA per language.
  3. Comprehensive evaluation. An in-depth analysis of state-of-the-art open-weight and closed-source VLMs, showing a regression from English to Indic languages in 4 out of 5 tasks across all models, with up to a 15.6-point drop in the worst case.
  4. A failure taxonomy. A categorization of common failure modes (missing Indian context, visual grounding errors, visual perception failures, failure to ground in Indian cultural context) to point to concrete areas for improvement.

Main Findings

  • Systematic English-to-Indic regression: Averaging across models and tasks, there is a gap of 8.3 points from English to Hindi and 5.5 points from English to Telugu, with regression on 4 out of 5 tasks.
  • Regression hits frontier models too: Even for GPT-4.1, the paper reports a 3.8-point drop from English to Hindi and an 8.6-point drop from English to Telugu.
  • Hindi and Telugu behave similarly: On the aligned subsets evaluated for GPT-4.1, Gemini 2.5 Flash and Chitrarth, average performance is 61.51 in Hindi, 62.53 in Telugu and 68.6 in English — within 1 point of each other, indicating a similar English-to-Indic gap for both languages.
  • Best-performing models: Gemini 2.5 Flash is the overall best performing model on both Hindi and Telugu in the paired datasets. Chitrarth is the leader on the VQAv2 subset in English and Telugu, attributed to its training on VQAv2 in multiple languages.
  • Language inconsistency within models: GPT-4.1 leads on the RealWorldQA subset for English and Hindi but performs much more poorly on the Telugu version, illustrating that strong performance in one language does not generalize to another.
  • Largest gaps by task: CLEVR-Math and RealWorldQA have the largest deltas between English and the Indic language sets; VAANI has the smallest.
  • VAANI reverses the trend: On VAANI-T, average performance is only marginally worse than English, and on VAANI-H the average is 2.09 points higher in Hindi than in English. The authors attribute this to English QA that failed to fully capture option meaning (including cases where visible text in the image was in the Indic language and got translated), and to text-only LLMs producing weak distractors, letting models exploit statistical patterns.
  • Chain-of-Thought helps English more than Hindi: On JEE-H, CoT boosts English by 13.8 points and Hindi by 2.22 points. Overall, CoT benefits the English subset by 3.78 points versus 0.79 points for Hindi, which the authors suggest may reflect more English CoT training data.
  • Caption-only is worse: Feeding a generated dense caption instead of the image underperforms direct image input on all tasks.
  • Failure analysis on VAANI-T (GPT-4.1): Visual grounding errors 49% (English) / 50% (Telugu); visual perception failures 19% / 18%; lack of knowledge about India 17% / 16%; failure to ground in Indian culture 15% / 16%.
  • Cost: Open-source models were run on H100 and A100 GPUs via the Akash Console Network (~100 USD), plus about 60 USD on OpenAI APIs — roughly 160 USD total.

Methodology in Plain English

The authors work in three parts. First, they build a translation pipeline for three English datasets — they sample 1,000 examples from VQAv2, extend RealWorldQA (765 QA), and sample 1,000 from CLEVR-Math — and translate them into Hindi and Telugu. To pick a translation tool, they review 50 diverse samples per language using IndicTrans, Google Translate, Azure and AWS; AWS Translate won for Telugu and Azure for Hindi. Every translated sample is then manually reviewed by co-authors who are native or fluent speakers, checking semantic accuracy, linguistic style, and fluency/readability. They deliberately avoid selecting samples for review via back-translation, because they found such subsets were biased toward the translation system's high-confidence mistakes.

Second, they create native Indic datasets. For VAANI-H and VAANI-T, they take broadly related image-caption subsets from the VAANI image-caption corpus and use text-only GPT-4.1 to generate multiple-choice questions from the captions; an LLM then filters out questions answerable from text alone, and humans remove low-quality items. For JEE-Vision, they collect diagram-dependent problems from the JEE-Advanced B.TECH exam (192 questions with images, Math/Physics/Chemistry, English and Hindi) and the JEE-Main B.TECH exam (325 questions with images, English and Telugu), using OCR with Gemini-2.5-Flash and manual correction of OCR errors.

Third, they evaluate. They only report a model on a target Indian language if the model claims prior proficiency in it. For Hindi they evaluate Gemma3-4B/12B/27B, Chitrarth-8B, Aya-8B, Qwen2.5VL-7B, LLaMA 3.2 Vision 11B, Gemini-2.5 Flash and GPT-4.1; for Telugu, where fewer VLMs are supported, they evaluate Gemini 1.5 Flash, 2.0 Flash and 2.5 Flash, GPT-4.1 and Chitrarth-8B. Multiple-choice and integer sets are scored by accuracy; VQAv2 and CLEVR-Math use a hybrid scheme that tries exact match first (using a modified official VQA evaluation script extended to Hindi and Telugu) and falls back to GPT-4.1 as an LLM judge. JEE-H uses regex-based answer extraction and rule-based scoring, including partial credit for multi-correct MCQs. They also ablate prompting strategies on Gemma-27B: standard prompting, chain-of-thought, and a caption-only baseline.

Why This Matters

Impact on research. The paper argues that unreliable multilingual evaluation — auto-translated, small, narrow, and culturally thin — slows progress toward equitable AI for low-resource languages. By releasing aligned English/Hindi/Telugu samples, it lets researchers separate task knowledge from language understanding and measure cross-lingual generalization exactly. The benchmark's scale (about 4k QA per language, with more than 1K images and more than 1K QA per language culturally sourced, which the authors say is 5 to 20 times larger than culturally relevant subsets of prior datasets) makes per-language conclusions more statistically meaningful than benchmarks with 50 or 135 QA per language.

Real-world applications:

  • Deploying multilingual customer-facing or government-service assistants that must interpret images and questions in Hindi or Telugu for Indian users.
  • Educational technology that helps students work through diagram-based STEM problems from exams like JEE in their own language.
  • Accessibility and assistive tools that describe images and answer questions about them in Indian languages.
  • Cultural heritage, media and e-commerce systems that need to recognize and reason about region-specific objects, traditions and artifacts.

Industry relevance. Model developers get a diagnostic signal that strong English performance does not transfer automatically: open-weight models generally lag proprietary ones, specific model families show large language-specific drops (Qwen2.5VL-7B drops 25.92 points from English to Hindi), and prompt strategies like chain-of-thought pay off far less in Hindi than English. The failure taxonomy also gives product teams concrete categories — visual grounding, visual perception, Indian knowledge, cultural grounding — to target.

Future Directions

  • Improving distractor quality in VAANI-generated MCQs, since text-only LLMs appear to produce distractors weak enough for models to exploit statistically; the authors propose using LLM-based quality control to align questions across languages and adopting Multi-Binary Accuracy for the VAANI subsets.
  • Expanding the JEE data source, because the JEE-Advanced set contains only 2 image-based Mathematics questions; the authors plan to include JEE-Mains questions.
  • Extending beyond static images to video understanding of culturally diverse, multilingual, multimodal content.
  • Extending beyond Hindi and Telugu to the other major Indian languages, with contributions from native speakers of each.

Target Audience

This paper is most useful for multimodal and multilingual NLP researchers and benchmark builders, VLM developers targeting South Asian languages, evaluation and responsible-AI teams assessing language coverage gaps, and practitioners building image-and-text products for Indian users who want to know where current models fail and how to construct reliable low-resource-language evaluation data. Researchers interested in annotation methodology will also benefit from the concrete details on translation tool selection, human review pass rates, and how much faster translation-plus-verification is than generating QA from scratch.

Authors’ abstract

With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to drive progress toward equitable AI for low-resource languages. Current multilingual VLM evaluations suffer from four major limitations: reliance on unverified auto-translations, narrow task/domain coverage, limited sample sizes, and lack of cultural and natively sourced Question-Answering (QA). To address these gaps, we present a scalable framework to evaluate VLMs in Indian languages and compare it with performance in English. Using the framework, we generate HinTel-AlignBench, a benchmark that draws from diverse sources in Hindi and Telugu with English-aligned samples. Our contributions are threefold: (1) a semi-automated dataset creation framework combining back-translation, filtering, and human verification; (2) the most comprehensive vision-language benchmark for Hindi and and Telugu, including adapted English datasets (VQAv2, RealWorldQA, CLEVR-Math) and native novel Indic datasets (JEE for STEM, VAANI for cultural grounding) with approximately 4,000 QA pairs per language; and (3) a detailed performance analysis of various State-of-the-Art (SOTA) open-weight and closed-source VLMs. We find a regression in performance for tasks in English versus in Indian languages for 4 out of 5 tasks across all the models, with an average regression of 8.3 points in Hindi and 5.5 points for Telugu. We categorize common failure modes to highlight concrete areas of improvement in multilingual multimodal understanding.

Read the original paper