Skip to content
AI.info

Research

Eye-Q: A Multilingual Benchmark for Visual Word Puzzle Solving and Image-to-Phrase Reasoning

Overview Research area: Computer Vision / Vision-Language Model evaluation (multimodal reasoning benchmarks). Technical level: Intermediate. The paper is a benchmark-and-evaluation study; it requires

Eye-Q: A Multilingual Benchmark for Visual Word Puzzle Solving and Image-to-Phrase Reasoning
arXiv
2601.03400
Published
2026-01-06
Authors
Ali Najar, Alireza Mirrokni, Arshia Izadyari, Sadegh Mohammadian, Amir Homayoon Sharifizade, Asal Meskin, Mobin Bagherian, Ehsaneddin Asgari

AI summary

Overview

Research area: Computer Vision / Vision-Language Model evaluation (multimodal reasoning benchmarks).

Technical level: Intermediate. The paper is a benchmark-and-evaluation study; it requires familiarity with vision-language models, prompting protocols, and accuracy metrics, but no specialized mathematical background.

Scope: A single-sentence scope: the paper introduces Eye-Q, a 1,343-puzzle multilingual benchmark that tests whether vision-language models can infer a hidden target word or short phrase from cue-implicit, distractor-rich images in English, Persian, Arabic, and cross-lingual (English–Persian) settings.

What This Paper Is About

Current vision-language benchmarks often let models succeed through surface recognition, OCR-style text reading, or retrieval of content that has circulated on the web, rather than through genuine abstract reasoning. The authors argue that visual word puzzles are a better stress test because they require discovering implicit visual cues, forming and revising hypotheses, and mapping perceptual evidence onto non-literal concepts such as puns, idioms, phonetic blends, and conventional phrases. Eye-Q is their benchmark for this task: given one image and a short prompt, a model must output the specific hidden word or phrase rather than select from candidates.

Key Contributions

  1. Task formulation. The authors formalize visual word puzzle solving as a vision-language reasoning task that demands multimodal cue integration, multi-step reasoning and search, and multilingual / cross-lingual generalization.
  2. Benchmark dataset. They release Eye-Q, a multilingual benchmark of 1,343 puzzles spanning English, Persian, Arabic, and cross-lingual settings (English 300, Persian 671, Arabic 50, cross-lingual 322).
  3. Cue-implicit design. The puzzles are deliberately unstructured, conceptually dense, and anti-OCR: cues involve orientation, counts, color, relative size, spatial relations (containment, overlap, inversion), material, posture, and affect, with distractors present.
  4. Human-aligned evaluation protocol. They propose an open-ended protocol probing hypothesis formation and revision through lightweight assistance: answer-length hints, few-shot examples, iterative refinement, and partial character reveal.
  5. Empirical findings. They benchmark six state-of-the-art LVLMs and report large performance gaps, especially on abstract and cross-lingual puzzles, with a maximum accuracy of 60.27%.

Main Findings

  • Performance is low overall. The peak accuracy across every model, subset, and prompt variant is 60.27%, achieved on the English subset by Grok 4.1 Fast (reasoning) under Partial Character Reveal.
  • Per-subset best results. English 60.27% (Grok 4.1 Fast, Partial Character Reveal), Persian 43.03% (Gemini 2.5 Pro, Partial Character Reveal), Arabic 19.15% (Gemini 2.5 Pro, Few-Shot CoT), and cross-lingual 29.15% (Gemini 2.5 Pro, Partial Character Reveal).
  • Proprietary models dominate. GPT-5.2, Gemini 2.5 (Flash and Pro), and Grok 4.1 Fast outperform the open-source Llama 4 Scout (17B) and Qwen 3 VL (235B), which show near-zero accuracy on Arabic and cross-lingual subsets in multiple settings (e.g., Llama 4 Scout scores 0.00% on Arabic under both Basic and Iterative Refinement; Qwen 3 VL scores 0.00% on Arabic and cross-lingual under Basic).
  • Assistance helps but does not close the gap. Averaged across models and language subsets, iterative refinement raises accuracy from 11.53% (Basic) to 15.63%, and Partial Character Reveal raises it further to 18.11%. Arabic stays below 10% on average even with Partial Character Reveal, and cross-lingual performance remains around 12% on average.
  • The failure mode is upstream, not surface-level. A cosine-similarity analysis using OpenAI text-embedding-3-large embeddings over failure cases shows density concentrated at low similarity across all four language subsets, meaning errors are typically not near-miss paraphrases. The authors conclude the bottleneck is cue selection and abstraction rather than output-space brittleness.
  • Scale helps within a model family but does not saturate. For the Qwen3-VL family (8B, 32B, 235B-A22B) on the English subset, larger models consistently perform better, with the biggest gains under assistance variants that encourage revision or constrain output space; accuracy still remains far from saturated.
  • Languages are tightly coupled. Across models, subset-level accuracies correlate highly: Persian and cross-lingual accuracy are the most tightly coupled (Pearson 0.98, Spearman 1.00), while Arabic shows weaker coupling (around 0.88 in multiple pairings), consistent with Arabic being the hardest subset.
  • Reasoning effort scales non-monotonically. Under a five-attempt iterative refinement setting on the English subset, the average number of attempts conditioned on solved puzzles first decreases from 8B to 32B and then increases for the 235B model; averaged over all puzzles (unsolved assigned five attempts), it decreases monotonically with model size.
  • Decoding temperature matters and differs by model. In the Few-Shot CoT English setup, Gemini 2.5 Pro declines steadily with temperature (47.14% at T=0.01, 44.11% at T=1.0, 42.76% at T=2.0), and Grok 4.1 Fast drops sharply (46.13%, 44.11%, then 0.34% at T=2.0). Weaker or more generation-sensitive models show an inverted-U: Gemini 2.5 Flash and Qwen 3 VL peak at T=1.0, and Llama 4 Scout breaks down completely at T=5.0 (0.00%).

Methodology in Plain English

The authors assembled the benchmark from two sources. Persian and cross-lingual puzzles come from Aftabe, a popular Iranian puzzle game released in 2014, used with permission and with the game-provided intended solution as ground truth; ambiguous items were manually filtered out. English and Arabic puzzles were authored by the team: they first chose a target word or short phrase, wrote a scene description intended to lead a human solver to it, and rendered the scene with text-to-image models (GPT and Nano Banana); all generated images were manually reviewed and only clear, artifact-free ones retained.

Each puzzle is an image plus a short prompt stating the game rule and requesting a single answer. Solving requires three steps the authors name explicitly: cue discovery (which elements are informative versus distractors), relational abstraction (reasoning over relations and transformations, not isolated objects), and linguistic association (mapping the inferred concept to an idiom, pun, phonetic resemblance, or conventional phrase in the target language). Cross-lingual puzzles require bridging English cues to Persian answers, or vice versa. Illustrative cases in the paper include a cup of "tea" on a "chair" yielding "teacher", "rain" plus a "bow" yielding "rainbow", and Persian examples such as "ماه" (moon) on a swing ("تاب") yielding "مهتاب".

Evaluation is open-ended exact-match. Six models were queried with one image and a fixed base template, varying only four prompting modules: Basic (an orthographic hint giving the answer length in characters), Few-Shot Chain-of-Thought (three solved demonstrations of the same subset, excluding the test instance, fixed across models for fairness), Iterative Refinement (re-query with the previous guess and a revision instruction, up to two revisions for three attempts total, counted correct if any attempt matches), and Partial Character Reveal (a randomly selected 25% of the answer's non-space characters is revealed, positions fixed by a random seed). All runs use each model's default decoding configuration and default visual preprocessing. Scoring trims whitespace and removes Persian and Arabic diacritics (A'rab), with additional minimal surface-form cleanup such as punctuation trimming and lowercasing. The authors also ran controlled temperature sweeps and embedding-based semantic near-miss analysis.

Why This Matters

Impact on research. Eye-Q offers an evaluation where success is hard to obtain through OCR, template-matching, or memorized web content, and it extends multimodal puzzle evaluation beyond English and Latin scripts. It also introduces what the authors describe as the first systematic evaluation of visual word puzzles with Persian answers and of cross-lingual puzzles bridging English visual cues with Persian solutions. The finding that failures are not near-miss paraphrases reframes the low scores as a cue-selection and abstraction problem rather than a string-matching artifact.

Real-world applications (from the paper's framing):

  • Evaluating and stress-testing multimodal assistants before deployment on tasks where literal grounding is insufficient.
  • Cross-lingual and culturally grounded applications, particularly English–Persian bridging, non-Latin scripts, and idiomatic expression.
  • Assessing whether additional model scale or additional reasoning steps actually buy better abstraction, relevant to model-selection and procurement decisions.
  • Diagnostic use in human-AI collaboration settings where a system must generate and revise hypotheses from ambiguous visual evidence.

Industry relevance. The gaps the paper reports bear on products that advertise "general-purpose" visual reasoning, on multilingual deployment in markets using Arabic and Persian, and on decisions about inference cost (iterative refinement adds queries for modest gains) and decoding configuration (temperature sensitivity varies sharply between models). The benchmark and code are released publicly at https://huggingface.co/datasets/llm-lab/Eye-Q and https://github.com/llm-lab-org/Eye-Q.

Future Directions

  • Extend to more languages and varieties. The authors note that dataset generation depends on contributors fluent in the target languages, which limits scaling to additional languages, dialects, or low-resource varieties while keeping difficulty and style consistent.
  • Build more reliable multilingual human evaluation. Fair correctness judgments ideally require native speakers or annotators of comparable proficiency per subset; the paper flags coordination cost and variability between annotator groups, especially for borderline cases.
  • Reduce ambiguity in ground truth. Eye-Q prioritizes the intended ground truth when alternative interpretations exist; a next step would be systematically characterizing and modeling these alternatives.
  • Attack the cue-discovery bottleneck directly. Since scaling and added refinement steps do not resolve the core difficulty, and since larger Qwen3-VL models already sustain longer productive reasoning chains, the open question is what training or inference method would improve the construction and search over conceptual representations for flexible image-to-phrase inference.

Target Audience

Researchers and engineers working on vision-language models, multimodal reasoning benchmarks, and multilingual evaluation; practitioners selecting or stress-testing VLMs for deployment in English-, Persian-, and Arabic-language products; and linguists or dataset builders interested in culturally grounded wordplay, non-Latin scripts, and cross-lingual image-to-phrase reasoning.

Authors’ abstract

Vision-Language Models (VLMs) have achieved strong performance on standard vision-language benchmarks, yet often rely on surface-level recognition rather than deeper reasoning. We propose visual word puzzles as a challenging alternative, as they require discovering implicit visual cues, generating and revising hypotheses, and mapping perceptual evidence to non-literal concepts in ways that are difficult to solve via literal grounding, OCR-heavy shortcuts, or simple retrieval-style matching. We introduce Eye-Q, a multilingual benchmark designed to assess this form of complex visual understanding. Eye-Q contains 1,343 puzzles in which a model observes a conceptually dense scene with a brief description and must infer a specific target word or phrase. The puzzles are intentionally unstructured and cue-implicit, with distractors and contextual relationships that demand selective attention, abstraction, and associative inference. The benchmark spans English, Persian, Arabic, and cross-lingual puzzles. We evaluate state-of-the-art VLMs using an open-ended, human-aligned protocol that probes hypothesis formation and revision under lightweight assistance. Results reveal substantial performance gaps, especially on abstract and cross-lingual puzzles, highlighting limitations in current models' ability to construct and search over appropriate conceptual representations for flexible image-to-phrase inference; maximum accuracy reaches only 60.27%.

Read the original paper