Skip to content
AI.info

Research

VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage

Overview Research area: Computer vision / multimodal evaluation, specifically Visual Question Answering (VQA) benchmarks for the cultural heritage and visual art domain. Technical level: Intermediate.

VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
arXiv
2510.12750
Published
2025-10-14
Authors
A. Alfarano, L. Venturoli, D. Negueruela del Castillo

AI summary

Overview

Research area: Computer vision / multimodal evaluation, specifically Visual Question Answering (VQA) benchmarks for the cultural heritage and visual art domain.

Technical level: Intermediate. The paper sits at the intersection of multimodal large language model (MLLM) evaluation and digital art history. The core ideas are accessible, but familiarity with VQA, multiple-choice benchmarking, and terms such as grounding and hallucination helps.

Scope: The paper introduces VQArt-Bench, a 14,463-question multiple-choice VQA benchmark for art and cultural heritage built with a multi-agent LLM pipeline, and reports an evaluation of 14 state-of-the-art MLLMs against it.

What This Paper Is About

Existing VQA benchmarks for art are largely built with rule-based or template-driven methods that paste structured metadata into fixed sentence patterns, producing shallow, repetitive, and sometimes factually incorrect questions. This incentivizes models to exploit statistical shortcuts and linguistic priors rather than actually looking at the artwork. The paper's goal is to build a semantically richer, agent-generated VQA benchmark for art and cultural heritage, and to measure how current MLLMs perform on it across seven visual reasoning dimensions.

Key Contributions

  1. A diagnosis of existing art VQA benchmarks. The authors show how rule-based generation (exemplified by AQUA, described as the first and only publicly available benchmark for art-focused VQA) produces shallow questions, incorrect terminology, and skewed answer distributions — for example, in the AQUA test set "Human" and "Person" alone account for more than 30% of all correct replies.

  2. A novel multi-agent question-generation pipeline. A sequential workflow combining agentic data cleaning with four question-crafting agents — a Topic Selector, a Question Generator, a Question Refiner, and a Judge — designed to produce linguistically complex, context-aware questions that test nuanced visual reasoning.

  3. VQArt-Bench, a new large-scale dataset. A benchmark of 14,463 high-quality multiple-choice questions in the cultural heritage domain, structured across seven evaluation dimensions.

  4. An extensive evaluation. A study of 14 state-of-the-art models on the new benchmark across multiple dimensions, reporting capabilities and limitations in visual art analysis.

Main Findings

  • Most MLLMs perform poorly on art. Most analyzed models fail to reach 50% accuracy overall, though they do better than random guessing on four-choice questions (approximately 25%).

  • Counting is a surprising weakness, reasoning is a strength. All evaluated models score significantly below their own overall accuracy on Instance Counting, yet excel on Visual-Inspired Reasoning — a result the authors call counterintuitive, since humans find counting simpler than interpreting an entire artwork. The authors attribute this to Visual-Inspired Reasoning requiring less specific instance knowledge and more generalization, and to Instance Counting questions being built with more attractive distractors.

  • Art is harder than standard benchmarks. InstructBLIP-Vicuna scored just below the 25% random-guessing value on VQArt-Bench, while having achieved the best overall score on the SEEDBench dataset (59%) under similar evaluation.

  • Closed-source models lead. Gemini 2.5 surpasses all current baselines on every evaluation metric, with an overall accuracy of 0.71. GPT-4o reached 0.64 and GPT-4o mini 0.57.

  • Kimi-VL is the best open-source model. Its overall accuracy of 0.67 outperforms even the closed-source GPT-4o and GPT-4o mini. The authors suggest this may stem from its native-resolution vision encoder (MoonViT) and a training strategy targeting visual reasoning rather than standard tasks like image captioning, and note it suggests MoE models can improve performance while reducing active parameters.

  • Larger models generally perform better. Gemma 3 27B improved by +9% over the 4B version and by +6% over the 12B version, without any fine-tuning; the authors cite GPT-4o's improvement over its mini version as another example.

  • Human validation found high quality. In a manual review of a random 25% sample of generated image-question pairs, over 98% of reviewed questions were confirmed free from hallucinations, and the questions covered most of the salient information in the source descriptions.

  • Dataset diversity is documented statistically. An automated analysis using Gemini 2.5 as a state-of-the-art visual LLM characterized the collection: a focus on scenes with 1-3 human figures alongside large crowds; a majority of male figures; Religious art and Portraiture as the most prominent genres; a significant presence of vegetation but mostly no animals; more indoor settings and natural landscapes than urban scenes, with most outdoor scenes set during the day; times of day focused on midday and afternoon; weather predominantly clear or cloudy; a strong emphasis on single clear light sources over diffuse lighting; and a tendency toward balanced or cooler color palettes.

Methodology in Plain English

The authors start from the observation that asking an LLM to invent a question directly from an image tends to produce superficial questions in an art context, because "relevance" is hard to operationalize. Instead, they ground questions in expert-curated descriptions.

Their data source is Wikipedia, chosen for its large repository of images with human-written collaborative descriptions. They downloaded approximately 30,000 images and their descriptions, then applied a length filter that discarded any image-text pair whose article contained fewer than 400 words. Because these descriptions interleave visual analysis with non-visual context (artist biographies, historical background) and sometimes describe other artworks, they use an LLM pre-processing step to isolate only the sentences relevant to the target artwork. They deliberately do not pass the image itself into this cleaning step, noting that images can help discrimination but also risk inducing hallucinated descriptions not present in the text.

Questions are then generated by a sequential pipeline of four specialist agents:

  1. Topic Selector — analyzes the cleaned visual description and proposes candidate topics, citing the minimal text snippet from the description that supports each one, to keep later generation factually grounded.

  2. Question Generator — turns grounded topics into nuanced, open-ended questions that are "informed" by the text but answerable by observing the image, avoiding questions that only the metadata could answer.

  3. Question Refiner — converts open-ended questions into multiple-choice format, deliberately designing plausible distractors based on likely visual misinterpretations of that specific image.

  4. Judge — a quality gatekeeper that verifies each question is unambiguously answerable from the image, non-trivial, aligned with the evaluation dimensions, and linguistically sound. Only questions passing this check enter the benchmark.

The seven evaluation dimensions, derived from prior work (SEEDBench), are Instance Identity, Instance Attribute, Instance Location, Instance Counting, Spatial Relation, Instance Interaction, and Visual-Inspired Reasoning. The final distribution is: Instance Identity 2031, Instance Attribute 2598, Instance Location 2100, Instance Counting 1710, Spatial Relation 2067, Instance Interaction 1794, Visual-Inspired Reasoning 2163, totaling 14,463.

For evaluation, the authors tested 14 models, including three Gemma 3 variants (4B, 12B, 27B) to test scaling, plus Aria, Aya Vision, Kimi-VL, Phi-4, Pixtral 12B, LLaVA, LLaVA-NeXT, InstructBLIP-Vicuna, Gemini 2.5, GPT-4o, and GPT-4o mini. Accuracy is the proportion of correctly answered multiple-choice questions.

Why This Matters

Impact on research. The paper argues that current art VQA benchmarks fail to test genuine visual literacy, and that rule-based generation actively rewards shortcut exploitation rather than visual reasoning. It connects computational benchmarking to a long scholarly tradition — Wölfflin on visual systems, Riegl's Kunstwollen, Warburg's Mnemosyne atlas, and Panofsky's three levels of interpretation (iconographic recognition of motifs, iconographic identification of subjects and symbols, and iconological analysis of cultural and ideological models). The authors position VQArt-Bench as complementing ontology-driven efforts such as ICON and IICONGRAPH by shifting focus from structured representation to active reasoning. It also provides a public leaderboard-style comparison across 14 models, exposing a clear closed-source versus open-source gap.

Real-world applications:

  • Museum and gallery digital cataloguing: systems that can describe, interpret, and answer visitor questions about artworks at a level beyond surface attributes.
  • Accessibility tools: richer image descriptions for visually impaired users that go beyond object labels to symbolic meaning and narrative.
  • Education and art history teaching: automated question generation for formative assessment using the seven reasoning dimensions.
  • Cultural heritage search and discovery: querying collections by symbolic content, action, or compositional relationship rather than only by metadata terms.

Industry relevance. For developers of MLLMs, the benchmark provides a harder, art-specific evaluation target that current models fail to clear even at 50% overall accuracy for most of them. The finding that a mixture-of-experts open-source model (Kimi-VL) can beat closed-source GPT-4o and GPT-4o mini on this task signals that specialized training strategies and native-resolution vision encoders matter more than raw parameter count alone. The consistent scaling trend (Gemma 3 27B > 12B > 4B) also gives product teams a concrete signal about model sizing.

Future Directions

  • Closing the counting gap. The unexpected weakness in Instance Counting is unexplained beyond the distractor-design hypothesis; whether better prompting, counting-specific training, or visual grounding mechanisms fix it remains open.

  • Broadening data sources. The pipeline is described as source-agnostic and Wikipedia was chosen here; applying it to museum catalogues, scholarly databases, or other repositories could test whether the approach generalizes.

  • Improving open-source models for art. The paper explicitly calls on the open-source community to intensify efforts on models tailored to artistic interpretation, given the persistent gap to Gemini 2.5.

  • Extending toward deeper interpretive layers. The authors frame their benchmark as complementing Panofsky-inspired ontology work, leaving open how far benchmarks can push from iconographic recognition toward iconological analysis of cultural and ideological models.

Target Audience

Researchers and practitioners working on multimodal large language models and VQA evaluation; digital humanities and cultural heritage computing scholars; art historians interested in computational interpretation; and machine learning engineers who need a domain-specific, non-saturating benchmark to test visual reasoning. It is also useful for anyone evaluating the practical limits of current MLLMs on specialized, out-of-distribution visual domains.

Authors’ abstract

Multimodal Large Language Models (MLLMs) have demonstrated significant capabilities in joint visual and linguistic tasks. However, existing Visual Question Answering (VQA) benchmarks often fail to evaluate deep semantic understanding, particularly in complex domains like visual art analysis. Confined to simple syntactic structures and surface-level attributes, these questions fail to capture the diversity and depth of human visual inquiry. This limitation incentivizes models to exploit statistical shortcuts rather than engage in visual reasoning. To address this gap, we introduce VQArt-Bench, a new, large-scale VQA benchmark for the cultural heritage domain. This benchmark is constructed using a novel multi-agent pipeline where specialized agents collaborate to generate nuanced, validated, and linguistically diverse questions. The resulting benchmark is structured along relevant visual understanding dimensions that probe a model's ability to interpret symbolic meaning, narratives, and complex visual relationships. Our evaluation of 14 state-of-the-art MLLMs on this benchmark reveals significant limitations in current models, including a surprising weakness in simple counting tasks and a clear performance gap between proprietary and open-source models.

Read the original paper