Research
Unexplored flaws in multiple-choice VQA make benchmarking unreliable
Overview Research area: Evaluation methodology for Multimodal Large Language Models (MLLMs), specifically multiple-choice Visual Question Answering (MC-VQA) benchmarking. Technical level: Intermediate

- arXiv
- 2511.22341
- Published
- 2025-11-27
- Authors
- Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagdonat, Stephan Günnemann, Leo Schwinn
AI summary
Overview
Research area: Evaluation methodology for Multimodal Large Language Models (MLLMs), specifically multiple-choice Visual Question Answering (MC-VQA) benchmarking.
Technical level: Intermediate. The paper combines large-scale empirical benchmarking with mechanistic interpretability analyses (tokenization and attention), so some familiarity with LLM decoding and prompt design helps, but the core argument is accessible.
Scope: A large-scale study showing that semantically neutral prompt format choices (option ID sets, delimiters, separators) make MC-VQA rankings unstable and therefore unreliable as a benchmark.
What This Paper Is About
Multiple-choice VQA is a standard way to benchmark multimodal models, and prior work has already shown that models are biased by the order in which answer options appear. The authors argue that existing fixes for option-order bias are not enough: even when order effects are controlled, results still swing wildly based on cosmetic prompt formatting choices that carry no semantic meaning, such as whether options are labeled "A/B/C/D" or "1/2/3/4" or whether a colon, dot, or bracket delimits each option.
The goal is to quantify how much these format choices distort MC-VQA rankings across models and datasets, to explain the mechanism behind the instability, and to test whether MC-VQA results are a valid proxy for open-ended model capability.
Key Contributions
-
Large-scale demonstration of rank instability: Across seven MLLMs, five MC-VQA datasets, and 48 semantically equivalent prompt formats, the authors show frequent rank reversals and large accuracy shifts even under order-invariant (circular) evaluation.
-
Mechanistic explanation: They trace instability to low-level language modeling effects, including tokenizer-induced fusion or removal of option ID tokens, generalization of option ID sets beyond training exposure, and attention head behavior during option selection.
-
Evidence of implicit format awareness: They show that most MLLMs can select prompt formats that perform better than random, achieving average ranks between 5.4 and 14.4 across datasets versus an expected random rank of 24.
-
Validity check against open-ended evaluation: They demonstrate weak to moderate correlation between MC-VQA rankings and open-ended evaluation (Kendall's τ of 0.09 on A-OKVQA, 0.52 on HRBench-4K, and 0.67 on V*Bench), questioning MC-VQA as a proxy for general multimodal capability.
Main Findings
-
Rank reversals are frequent. Across 48 prompt formats per dataset, rankings vary substantially and frequently invert. On A-OKVQA and MME-RealWorld-Lite no model exhibits a stable rank; on MMBench and V*Bench the top-ranked model stays stable while lower-ranked models fluctuate; HRBench-4K shows the most stable rankings.
-
Extreme single-model shifts. LLaVA-OV moves from the second-lowest to the top rank under a changed prompt format, with an accuracy difference of nearly 50 percentage points (pp). LLaVA-1.5 and Qwen-2-VL also show considerable accuracy differences across formats.
-
Order control is not sufficient. Circular evaluation effectively mitigates option order preference, yet format-driven instability persists, and all prompt format variations produce statistically significant accuracy changes.
-
Scaling does not fix the problem. On V*Bench, Qwen2.5-VL-32B shows a 2.6 pp deviation versus 7.1 pp for the 7B variant, but the 7B model still outperforms the 32B model under some formats (7 out of 48 formats favor the 7B; 41 out of 48 favor the 32B). Gemma-3-27B shows increased sensitivity at 5.9 pp deviation versus 3.9 pp for the smaller model. GPT-5.2 variants also display substantial prompt-induced variability.
-
Full accuracy ranges on V*Bench (Table 2). Qwen2.5-VL-7B: average 77.44%, min 72.65%, max 81.71%. Qwen2.5-VL-32B: average 80.52%, min 78.85%, max 81.54%. Gemma-3-4B: average 34.41%, min 32.31%, max 36.25% (higher accuracy in 0 of 48 formats). Gemma-3-27B: average 45.09%, min 41.78%, max 47.65% (48 of 48). GPT-5.2-nano: average 59.70%, min 57.21%, max 61.41% (0 of 48). GPT-5.2-mini: average 70.05%, min 67.61%, max 71.81% (48 of 48).
-
Tokenization is a direct cause. When a tokenizer fuses the option ID and delimiter into a corrupted token so the bare option ID never appears (for example, splitting "(A)" into "(A" and ")"), some models lose accuracy. Gemma-3 consistently preserves atomic option IDs across all evaluated formats and correspondingly shows the lowest accuracy variation. LLaVA-1.5, LLaVA-OV, and Qwen2-VL show significant accuracy decreases when the bare option ID token is missing, while Phi-4 and Qwen2.5-VL remain stable.
-
Embedding similarity tracks the effect (Table 3). Models harmed by non-atomic tokenization show low cosine similarity between fused and bare option ID embeddings: Phi-3.5 at 0.0110 with −5.0% average accuracy deviation; LLaVA-1.5 at 0.0341 with −43.4%; Qwen2-VL at 0.2617 with −18.3%; LLaVA-OV at 0.2672 with −30.6%. Unaffected models: Qwen2.5-VL at 0.2390 with +0.9%; Phi-4 at 0.6139 with +4.2%; Gemma-3 has no negative effect.
-
Training exposure does not explain ID preferences. Evaluating Phi-4 and Qwen2.5-VL on V*Bench with Greek letter words, directional words, and concrete nouns, Greek-letter IDs outperform numbers and Roman numerals and are comparable to letter-based IDs, despite being unlikely to appear in multiple-choice training data.
-
Attention margin correlates with accuracy (Table 4). On A-OKVQA, average attention margins and accuracy are: lowercase 0.609 / 87.51%, uppercase 0.596 / 87.64%, numbers 0.526 / 85.73%, Roman 0.509 / 86.19%, with Pearson ρ = 0.930. On V*Bench: lowercase 0.651 / 80.10%, uppercase 0.648 / 79.58%, Greek 0.620 / 76.69%, numbers 0.543 / 75.45%, Roman 0.499 / 74.34%, with Pearson ρ = 0.934.
-
Models select favorable formats (Table 5). Gemma-3 selects uppercase/colon/linebreak with average rank 14.4 and relative gap −1.61 to optimal; LLaVA-1.5 the same format with rank 7.0 and gap −3.12; LLaVA-OV uppercase/bracket/linebreak with rank 5.4 and gap −0.79; Phi-3.5 uppercase/colon/linebreak with rank 9.6 and gap −1.42; Phi-4 numbers/dot/comma with rank 34.0 and gap −4.21; Qwen2-VL uppercase/colon/linebreak with rank 11.6 and gap −1.08; Qwen2.5-VL uppercase/colon/linebreak with rank 7.4 and gap −0.92.
-
Weak alignment with open-ended evaluation. Grounded multiple-choice (giving models their own open-ended answers) improves accuracy for most models but does not restore original rankings and does not eliminate prompt sensitivity.
-
Prompt-format transfer is unreliable (Table 6). Selecting a model-specific format on HRBench-4K or V*Bench and applying it elsewhere produces large swings on A-OKVQA (for example, +65.17 pp for LLaVA-1.5, +48.21 pp for LLaVA-OV, +40.96 pp for Qwen-2-VL with HRBench-4K selection; averages of 22.28 pp and 21.60 pp across models) alongside near-zero or small changes on other datasets.
Methodology in Plain English
The authors built an exhaustive grid of prompt formats rather than testing a few hand-picked variants. They decomposed a multiple-choice prompt into three independent factors: the set of option IDs (uppercase letters, lowercase letters, numbers, Roman numerals), the delimiter that links an ID to its option text (colon, dot, bracket, double brackets), and the separator between options (line break, comma, semicolon). Four ID sets times four delimiters times three separators yields 48 semantically identical prompt formats.
They evaluated seven open models — Gemma-3, LLaVA-1.5, LLaVA-OV, Phi-3.5, Phi-4, Qwen-2-VL, and Qwen2.5-VL — in a zero-shot setting with sampling disabled for deterministic outputs, on five datasets: HRBench-4K, V*Bench, A-OKVQA, MME-RealWorld-Lite, and MMBench. To remove the known confound of option order, they used circular evaluation, effectively averaging over positions. They then compared how model rankings shift across the 48 formats, checked whether larger and proprietary models behave differently, and inspected tokenizer output and attention weights to explain why the shifts occur. Finally, they compared MC-VQA rankings against open-ended generation scored by an LLM-as-a-judge, and tested a simple mitigation where a format chosen on one dataset is reused on others.
Why This Matters
Impact on research: The paper argues that a widely used benchmark class is confounded by a variable most papers never report: the prompt format. Rankings produced by MC-VQA can change by tens of percentage points from choices that do not change the task, which undermines claims of model progress and makes cross-paper comparisons fragile unless the exact format is disclosed.
Real-world applications:
- Medical imaging model selection, where multimodal models are evaluated with multiple-choice question sets and a wrong benchmark conclusion could steer deployment decisions.
- Autonomous driving, where visual question answering and multimodal evaluation are used to assess perception and reasoning components.
- High-resolution and specialized imaging pipelines that rely on benchmark scores to choose which model to integrate.
- Model procurement and vendor comparison, where buyers use published leaderboard ranks to decide between providers.
Industry relevance: The risk is directly commercial. Companies that pick models based on leaderboard position may be selecting on option-selection dynamics rather than reasoning ability. The authors are affiliated with the Technical University of Munich, Volkswagen AG, Helmholtz AI, and ELLIS, and the paper carries a disclaimer that its conclusions are not necessarily those of Volkswagen Aktiengesellschaft.
Future Directions
- Develop format-controlled evaluation protocols that explicitly report and vary option IDs, delimiters, and separators instead of fixing them silently.
- Investigate whether the findings hold under few-shot prompting and stochastic decoding, which the authors explicitly list as a limitation of their zero-shot, deterministic setup.
- Improve tokenizer and embedding design so option IDs remain atomic and identity-preserving across formats, since Gemma's tokenizer behavior correlated with the lowest accuracy variation in this study.
- Build more reliable mitigations. Validation-based prompt-format selection reduces extreme failures but generalizes poorly across datasets and remains computationally expensive, and the open-ended evaluation used here relies on an LLM-as-a-judge protocol that the authors do not systematically validate and that uses models from the same families as some tested MLLMs.
Target Audience
Researchers and practitioners who design, run, or consume MLLM benchmarks will get the most from this paper, particularly those working on multimodal evaluation, VQA leaderboards, or benchmark methodology. It is also valuable for model developers concerned with tokenizer design and for teams in industry that select multimodal models based on published rankings and need to know which evaluation artifacts to distrust.
Authors’ abstract
Previous works identify sensitivity to option order as a key issue in multiple-choice VQA (MC-VQA) evaluation and propose protocols to mitigate this effect. We show that such mitigation is insufficient to ensure the validity of MC-VQA as a reliable benchmark for Multimodal Large Language Model (MLLMs): performance remains highly sensitive to semantically neutral prompt format choices that are not controlled by current benchmarks. In a large-scale study spanning seven MLLMs and five MC-VQAs datasets, we find frequent rank reversals even under order-invariant evaluation. These reversals arise when we systematically vary option ID sets, delimiters, and separators, yielding 48 semantically equivalent prompt formats. Mechanistic analyses trace this instability to low-level language modeling effects: tokenizer-induced fusion or removal of option ID tokens introduces corrupted option ID tokens into the input sequence, while the choice of option ID sets directly affects the reliability of attention patterns for option selection. Accordingly, MC-VQA rankings correlate weakly with open-ended evaluation, indicating that MC-VQA reflects option-selection dynamics in addition to multimodal reasoning. These findings identify prompt formatting as a major, previously under-controlled confounder in MC-VQA benchmarking and motivate evaluation protocols that explicitly control prompt format sensitivity.