Research
VISaGE: Understanding Visual Generics and Exceptions
Overview Research area: Natural Language Processing / Vision-Language Model evaluation, specifically the semantics of generics (unquantified generalizations such as "cats have four legs") and how visu

- arXiv
- 2510.12548
- Published
- 2025-10-14
- Authors
- Stella Frank, Emily Allaway
AI summary
Overview
Research area: Natural Language Processing / Vision-Language Model evaluation, specifically the semantics of generics (unquantified generalizations such as "cats have four legs") and how visual grounding interacts with conceptual knowledge.
Technical level: Intermediate. The paper is readable without deep background, but it assumes familiarity with vision-language models, prompting-based evaluation, and the notion of generics from formal semantics.
Scope: The paper introduces VISaGE, a dataset of typical and "exceptional" images paired with category-attribute norms, and uses it to test whether open-weight VLMs can separate conceptual (generic) knowledge from instance-level visual evidence.
What This Paper Is About
VLMs learn general conceptual knowledge during training but are almost always queried about individual images, and most prior benchmarks use typical images to stand in for a concept. This conflates instance-level understanding with conceptual knowledge, leaving untested what happens when an image is an exception to a generic (for example, a tripod cat for the norm "cats have four legs"). The paper builds a dataset of such exceptions and measures how two competing model priors — a pragmatic prior that text and image are both relevant and congruent, and a semantic prior that a category-attribute generalization generally holds — trade off against each other.
Key Contributions
-
A new evaluation dataset, VISaGE (Visual Generics and Exceptions). It contains 1601 exceptional image examples for 437 exception subcategories, derived from 296 category-attribute relations (generics/conceptual norms) covering 171 categories, balanced with the same number of typical images. Dataset and code are at github.com/scfrank/visage1601.
-
Experimental evidence that VLM conceptual representations are grounded only in typical or generic instances and are not robust to within-category variation, based on carefully balanced experiments across multiple open-weight VLMs.
-
A demonstration that incongruency between image and text damages conceptual understanding more than the semantic prior damages instance-level recognition. The paper reports that the effect of pragmatic prior violation is stronger than violations of the semantic prior.
-
A feature-attribution analysis using Shapley values (MM-SHAP) that quantifies how much the image versus the text contributes to model predictions across the conceptual and instance query conditions.
Main Findings
-
Conceptual accuracy collapses when the image is incongruent with the text. In Experiment 1, asking "Do lions have manes?" with an exceptional image (condition 1b) rather than a typical one (1a) produces a large accuracy drop across models. For example, smolvlm2 falls from 0.8657 (1a) to 0.4372 (1b), and llava-next falls from 0.9319 to 0.6758. This is interpreted as the pragmatic prior overriding correct retrieval of conceptual knowledge.
-
Incongruency hurts less when the query is about the exception subcategory and the image is typical (condition 1d versus 1c). The paper suggests this may be because models process atypical images more attentively than typical ones.
-
Instance queries do not rescue performance on exceptional images. Experiment 2 shows a drop from typical images (2a) to exceptional images (2b) even though the query is instance-level ("Does this lion have a mane?"). If models prioritized visual features, accuracy would be stable across these conditions; it is not.
-
Naming the exception does not fix the problem. Condition (2c), which uses the exception name with an exceptional image (text and image congruent), yields accuracy comparable to (2b), indicating that supplying the correct semantic information is of limited use for improving instance attribute recognition in atypical cases.
-
Possible explanations offered: exception categories are lower frequency than general categories, so attribute knowledge about them may be less developed; and many exception names contain the category name (e.g., lioness, tripod cat), which could semantically prime the general category and its attributes.
-
Not all models show the semantic bias. Four of the evaluated VLMs are more accurate on exception images than typical images in instance queries, which the authors speculate reflects sensitivity to image (a)typicality and to the pragmatics of the query.
-
Pragmatic prior violation is the stronger effect. In Figure 3iii, when visual input and category name are congruent, conceptual and instance query performance is similar (difference near zero). Differences appear only with incongruent inputs, where conceptual queries suffer relative to instance queries — which is why exceptional examples are needed to reveal the interaction.
-
Shapley analysis on smolvlm2 shows a general pragmatic bias plus quantitative shifts. There is no qualitative difference in image use between concept and instance conditions, but V-SHAP increases for instance queries relative to the matching concept queries (significant at p < 0.01; Cohen's d = 0.3 for typical image conditions and d = 0.19 for exception image conditions, both small effects). Exception image conditions have higher V-SHAP than typical image conditions (p < 0.01; d = 0.62 for concept queries and d = 0.56 for instance queries, both medium effects), suggesting models recruit visual information to supplement or counteract incongruent conceptual information.
-
Full per-model accuracy numbers for every experiment condition are given in the paper's Appendix G table, covering smolvlm2, deepseek-vl-v2-tiny, paligemma2, phi3-v, gemma3-4B, phi4_mm, llava-next, molmo, idefics3, internvl3 and gemma3-12B.
Methodology in Plain English
The authors first build the dataset in text. They intersect the category-attribute lists from the XCSLB norms and the McRae norms with categories in the THINGS object image dataset, producing a set of conceptual norms expressed as generics. For each generic they generate candidate exceptions (subcategories that violate the norm) using the LM prompting framework from Allaway et al. (2024) with GPT-3.5 (gpt-3.5-turbo-0613), keeping the top 5 candidates ranked by perplexity and removing false ones.
They then pair each norm and exception with images. Exception images are retrieved from Bing Image Search by querying the exception name, and typical images come from the THINGS dataset, which was collected specifically to contain typical object instances. Human annotators — the paper's authors — validate three things for each tuple: that the typical images exhibit the conceptual norm, that each generated exception is genuinely an exception to that norm (filtering hallucinated items such as "strawberry blonde cheetah" and incorrectly related ones), and that the retrieved exception images are correct in both category and style. During validation, annotators could also add new exceptions and new category-attribute relations; this revision process added 121 new norm-and-exception tuples and nearly doubled the dataset, from 872 tuples to the final 1601. The final acquisition mode is 4 images per exception.
For the experiments, the authors query open-weight VLMs with yes/no questions, varying three factors: whether the question is conceptual or instance-level, whether the image is typical or exceptional, and whether the noun phrase refers to the category or to the exception subcategory. Accuracy is the percentage of correct responses, scored on the first output token. Conceptual prompts use the plural form ("Answer yes or no. Do {concept-pl} have {attribute}?"), and instance prompts use the singular ("Answer yes or no. Does this {concept-sg} have {attribute}?"). The correct answer depends on the condition. Models were run with the vllm library (version 0.8.5.post1) with transformers v4.52.0.dev0 and torch v2.6.0, at generation temperature 0.05, on Nvidia A100 or A4500 GPUs, with roughly 15 minutes per single-model, single-condition evaluation.
For interpretability, the authors compute Shapley values following MM-SHAP, treating the image as a single feature and text tokens as individual features, using the ExactExplainer from the shap library (version 0.48.0). They mask the image by removing it entirely and mask language tokens with "_". V-SHAP is the proportion of Shapley values coming from the image and T-SHAP the proportion from text tokens, aggregated over the dataset. This analysis was run on smolvlm2.
Why This Matters
The paper identifies a specific failure mode that standard VLM benchmarks miss: because those benchmarks rely on typical instances, they cannot expose the conflict between conceptual knowledge and visual evidence that arises when an input is atypical. The dataset makes that conflict measurable and shows that current models neither reliably attend to an exception instance while ignoring the semantic prior, nor reliably ignore distractor images when answering generic conceptual queries.
Real-world applications:
- Assistive and accessibility systems that describe images to users need to report what is actually in the image rather than what is typical for the category, especially when the depicted object is unusual.
- Medical and diagnostic imaging, where an image may show an atypical presentation that contradicts the general rule for a condition, and where trusting the generic over the image would be harmful.
- Content moderation and safety pipelines that reason about whether an image contains a prohibited category typically associated with a generalization.
- E-commerce and inventory tagging, where a product photo that deviates from the norm for its category must still be labeled correctly.
Industry relevance: Any deployment where a model must answer about a specific image using category knowledge — automated captioning, visual question answering, robotics perception, document and photo organization — inherits the failure modes reported here. The finding that models lean on textual category cues even under instance queries is directly relevant to how prompts are written and how much trust to place in VLM outputs about unusual inputs.
Future Directions
-
Extending VISaGE beyond American English conceptual norms, since conceptual spaces are language-dependent and other languages make different conceptual distinctions and attend to different attributes. The authors state they believe the general pattern of results would hold across languages and models but do not test this.
-
Closing the recall gap in exception coverage: the data collection prioritized quality over recall and omitted exceptions that are rare, hard to see, or unlikely to be photographed, such as an "insomniac owl" as an exception for "owls sleep in the day" or a "cheetah with a broken leg" as an exception for "cheetahs are fast."
-
Investigating why some models are more accurate on exceptional images than typical ones, and whether that reflects genuine sensitivity to image atypicality or an artifact of query pragmatics.
-
Applying the framework to categories that group people, where the paper notes that failing to distinguish exceptional instances from conceptual generalizations can lead to stereotyping, and where understanding VLM capabilities is framed as a step toward mitigating that risk.
-
Addressing the paper's conclusion directly: whether VLMs can be built or prompted to both attend to exception instances while overriding the semantic prior, and to ignore distractor images when answering generic conceptual queries.
Target Audience
Researchers working on vision-language models, evaluation benchmark design, and computational semantics of generics will find the dataset and experimental protocol most directly useful. The paper is also relevant to practitioners building applications that must reason about atypical inputs, and to anyone studying how models weigh in-context evidence against in-weights knowledge. A reader interested in the overlap between formal semantics (quantification, generics, exceptions) and multimodal model evaluation will get the most from it; a reader looking for training recipes or architectural solutions will not find them here, since the paper is an evaluation and analysis study.
Authors’ abstract
While Vision Language Models (VLMs) learn conceptual representations, in the form of generalized knowledge, during training, they are typically used to analyze individual instances. When evaluation instances are atypical, this paradigm results in tension between two priors in the model. The first is a pragmatic prior that the textual and visual input are both relevant, arising from VLM finetuning on congruent inputs; the second is a semantic prior that the conceptual representation is generally true for instances of the category. In order to understand how VLMs trade off these priors, we introduce a new evaluation dataset, VISaGE, consisting of both typical and exceptional images. In carefully balanced experiments, we show that conceptual understanding degrades when the assumption of congruency underlying the pragmatic prior is violated with incongruent images. This effect is stronger than the effect of the semantic prior when querying about individual instances.