Research
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
Overview Research area: Multimodal NLP — specifically vision-language models (VLMs) applied to table question answering, with a focus on multilingual evaluation and robustness to visual degradation. T

- arXiv
- 2511.17238
- Published
- 2025-11-21
- Authors
- Anshul Singh, Rohan Chaudhary, Gagneet Singh, Abhay Kumary
AI summary
Overview
Research area: Multimodal NLP — specifically vision-language models (VLMs) applied to table question answering, with a focus on multilingual evaluation and robustness to visual degradation.
Technical level: Intermediate. The paper is a benchmark-and-evaluation study; it requires some familiarity with VLM/table-QA literature and metrics like Exact Match and BLEU, but the core argument is accessible.
Scope: The paper introduces MirageTVQA, a benchmark of nearly 60,000 table question-answer pairs spanning 24 languages with both clean and visually noisy table images, and reports an empirical evaluation of leading open-source VLMs on it.
What This Paper Is About
Existing table QA benchmarks such as WikiTableQuestions and FinQA are overwhelmingly monolingual (English) and present tables as digitally perfect, clean text — a gap between what models are tested on and what tables actually look like in scanned documents, financial reports, or photographs around the world. The paper builds a benchmark that stresses both dimensions at once (many languages and realistic visual noise) and then measures how state-of-the-art VLMs fail. The goal is to show that scores on pristine, English-only data are an unreliable predictor of real-world table reasoning.
Key Contributions
- MirageTVQA, a new benchmark. The authors describe it as the first large-scale visual question-answering benchmark combining massive multilingual support (24 languages) with visually realistic table images, totaling a final valid set of 58,480 QA pairs and nearly 60,000 QA pairs overall.
- A multilingual table corpus built via a translate–refine–filter pipeline. Starting from a curated seed of 250 English tables (50 from arXiv, 100 from Wikipedia, 100 from other sources) drawn from an initial pool of 3,000 English tables, translations were produced with Qwen3-32B, refined with Gemini 2.5 Pro against the original English table, back-translated, and filtered by back-translation BLEU score — reducing 30 candidate target languages to 24.
- A two-stage rendering pipeline that injects realistic noise. Tables are rendered to clean PNGs from HTML using 40+ CSS themes, then degraded with the
imgauglibrary to simulate camera capture and scanning (rotations, skew, perspective transforms, Gaussian blur, variable JPEG compression, salt-and-pepper noise, scan lines, corner shadowing). - An extensive empirical evaluation of leading open-source VLMs across languages and across clean versus noisy images, including an analysis of model scale.
Main Findings
- Large accuracy gap overall. On clean images across all languages, the best model evaluated, Qwen2.5-VL-72B, achieves an average Exact Match (EM) of only 13.57%. Performance correlates strongly with model scale: the largest models consistently beat smaller ones (e.g., Qwen3-8B at 8.01% average EM versus Qwen3-30B at 9.45% versus Qwen2.5-72B at 13.57%).
- Visual noise causes severe degradation, most acutely in the strongest model. For the English subset, Qwen2.5-VL-72B drops from 25.52% EM on clean images to 16.50% EM on noisy images — a 35.3% drop. Other models degrade less: Qwen 2.5 VL 32B goes from 23.15% to 20.36% (−12.1%), Qwen 2.5 VL 8B from 17.53% to 16.62% (−5.2%), and Gemma-3 27B-IT from 13.79% to 12.87% (−6.7%).
- A consistent English-first bias. Across all model scales, performance peaks on English (5.23% EM for Pangea-7B, up to 25.52% for Qwen2.5-VL-72B on the clean English subset). Scores drop sharply for other high-resource languages, degrade further for languages with different scripts, and become negligible for many low-resource languages (e.g., Hokkien at 5.47% or 0.12% depending on the model).
- Scale helps but does not fix the underlying problems. Greater parameter counts improve multi-step reasoning on clean images, yet the paper notes that performance on pristine, synthetic data is an unreliable predictor of performance on realistic, visually imperfect data.
- Contamination-style artifact observed. InternVL3-14B produces 0.00 EM on several languages (Spanish, French, Indonesian variants, and Korean) while scoring 12.72% on English, an extreme instance of the language skew.
Methodology in Plain English
The authors started by gathering 3,000 English tables from four sources: Wikipedia (via WikiSQL), financial documents from FinQA, scientific papers from arXiv, and GitHub. They filtered these down to 250 "seed" tables judged suitable for translation based on the median word character count per cell.
They then built a multilingual corpus with a three-step pipeline per language: translate the table's text with Qwen3-32B, have Gemini 2.5 Pro refine the translation while cross-checking the original English table for context and data integrity, and back-translate to English — keeping only languages whose back-translation BLEU score was high enough.
For the visual side, each table in each language was rendered from HTML into a clean PNG using 40+ hand-designed CSS themes. From each clean image, the imgaug library produced several noisy variants through a stochastic pipeline of geometric distortions (minor rotations, skew, perspective transforms), quality degradation (Gaussian blur, variable JPEG compression), and scanning artifacts (salt-and-pepper noise, scan lines, corner shadowing). Each table therefore has one clean image and multiple noisy counterparts, with metadata for each applied transformation.
Question-answer pairs were built with a combined human-LLM approach: human annotators wrote one high-quality QA pair per table as a seed, and an LLM (Google Gemini) expanded each seed into 10 additional QA pairs covering 10 reasoning types (comparative, numerical aggregation, multi-hop, temporal, conditional, proportional/ratio, hypothetical, correlation inference, structural/metadata, and outlier detection) and two question types (value and open-ended reasoning). This yielded 11 QA pairs per table-language combination and a total of 80,520 QA pairs (244 tables × 30 languages × 11 QA pairs). LLMs then validated the generated pairs, flagging misclassified ones, and three human annotators corrected the flagged pairs, leaving a final valid set of 58,480 QA pairs across 24 languages. Both tables and QA pairs were translated into all target languages.
Finally, models were evaluated by Exact Match, comparing clean versus noisy images and comparing performance across languages.
Why This Matters
Impact on research. The paper argues that existing evaluation practice — clean, English-only table benchmarks — systematically overstates VLM capability. MirageTVQA provides a measurement instrument for two axes that the literature had treated separately (visual complexity and linguistic diversity) and makes the case that the strongest models are also the most brittle: the 72B model's advantage on clean images largely collapses under noise.
Real-world applications:
- Financial document processing, where tables arrive as scanned or photographed reports in many languages (the paper draws FinQA-style financial tables into its source pool).
- Scientific literature extraction, where tables from arXiv papers and similar sources must be parsed from PDF images rather than clean markup.
- Healthcare records and enterprise databases, which the introduction cites as domains where tables are the backbone of information storage and where degraded scans are common.
- Global/multilingual document workflows, since the failure of cross-lingual transfer means non-English users cannot rely on these models even when the same table would be handled well in English.
Industry relevance. Any organization deploying document AI for automated data extraction, RAG over scanned reports, or agentic table reasoning should treat the reported drop from 25.52% to 16.50% EM under noise as a warning against benchmarking on synthetically clean data. The benchmark and code are released at https://github.com/anshulsc/MirageTVQA.
Future Directions
- Interpretability of failure modes. The authors state explicitly that they did not establish interpretability methods to explain why noise causes degradation, nor ways to reduce it — they flag this as an important area for future work.
- Evaluating proprietary models. The study covers only open models; the limitations section notes that top-tier proprietary models may behave differently under noise and should be tested.
- Wider language coverage. The authors acknowledge that their cross-lingual experiments were limited in scope (they refer to 25 languages in the limitations section and report a final valid set of 24 languages), and propose introducing a wider array of languages.
- Robustness mitigation. Beyond measuring degradation, the paper points toward methods for understanding and addressing it — implying training or inference-time techniques that transfer reasoning to non-English and visually degraded inputs.
Target Audience
Researchers and practitioners working on vision-language models, multimodal document understanding, and table question answering; benchmark builders interested in multilingual evaluation design; and applied teams building document AI pipelines for finance, science, or enterprise data who need to know where current VLMs break. Readers with a basic grasp of QA metrics (Exact Match, BLEU) and multimodal model families (Qwen, Gemma, InternVL) will get the most out of it.
Authors’ abstract
The impressive performance of VLMs is largely measured on benchmarks that fail to capture the complexities of real-world scenarios. Existing datasets for tabular QA, such as WikiTableQuestions and FinQA, are overwhelmingly monolingual (English) and present tables in a digitally perfect, clean format. This creates a significant gap between research and practice. To address this, we present \textbf{MirageTVQA}, a new benchmark designed to evaluate VLMs on these exact dimensions. Featuring nearly 60,000 QA pairs across 24 languages, MirageTVQA challenges models with tables that are not only multilingual but also visually imperfect, incorporating realistic noise to mimic scanned documents. Our evaluation of the leading VLMs reveals two primary failure points: a severe degradation in performance (over 35\% drop for the best models) when faced with visual noise and a consistent English-first bias where reasoning abilities fail to transfer to other languages. MirageTVQA provides a benchmark for measuring and driving progress towards more robust VLM models for table reasoning. The dataset and the code are available at: https://github.com/anshulsc/MirageTVQA.