Research
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
Overview Research area: Multimodal AI evaluation — specifically benchmarks for large vision-language models (LVLMs), drawing on topology as a way to test global visual perception. Technical level: Int

- arXiv
- 2511.11831
- Published
- 2025-11-14
- Authors
- Wenhao Zhou, Hao Zheng, Rong Zhao
AI summary
Overview
Research area: Multimodal AI evaluation — specifically benchmarks for large vision-language models (LVLMs), drawing on topology as a way to test global visual perception.
Technical level: Intermediate. The paper is readable without deep mathematics, but it assumes familiarity with LVLM architecture (visual encoder, projection module, LLM) and standard evaluation metrics such as accuracy, precision, recall, and F1.
Scope in one sentence: The paper introduces TopoPerception, a benchmark built from synthetic images with pure topological properties, and uses it to show that state-of-the-art LVLMs fail at global visual perception even at the easiest difficulty level.
What This Paper Is About
Current LVLM benchmarks test semantically rich tasks like visual question answering and captioning, but these tasks contain "local shortcuts" — models can answer correctly by picking up on a single object, a co-occurring detail, or even just the question text, without ever perceiving the image as a whole. That means benchmark scores can overestimate how well a model actually sees. This paper builds a test where those shortcuts are structurally impossible, by asking models to classify images based on topological properties (which depend only on global structure and are invariant to local features), and finds that top models perform no better than random guessing.
Key Contributions
-
A shortcut-free benchmark (TopoPerception). The benchmark defines multiple-choice tasks based on topological properties of synthetic images across a scalable hierarchy of perceptual granularities, providing "a scalable difficulty hierarchy for stress-testing LVLMs."
-
A comprehensive evaluation of SOTA LVLMs. The authors test seven models (GPT-4o, o4-mini, o3, Claude-sonnet-4-0, Claude-opus-4-0, Gemini-2.5-flash, Gemini-2.5-pro) and report "systematic and severe deficiencies" in global visual perception, with results at or near random guessing even at the coarsest granularity.
-
Discovery of a counter-intuitive scaling trend. Within the same model family, larger models with stronger reasoning capabilities tend to score lower on TopoPerception, suggesting that scaling up reasoning may interfere with or override the visual signal.
-
A taxonomy of benchmark shortcuts. The paper classifies shortcuts at a statistical level and a semantic level, and separately as Type 1 (solvable from text alone) and Type 2 (a shortcut inside the image itself), explaining how TopoPerception eliminates both.
Main Findings
-
Near-random performance at the easiest level. At Level 0 (a resolution of 29 × 29, comparable to MNIST's 28 × 28), accuracy was GPT-4o 22.00%, o4-mini 19.67%, o3 12.00%, Claude-sonnet-4-0 30.00%, Claude-opus-4-0 24.33%, Gemini-2.5-flash 33.33%, and Gemini-2.5-pro 30.67%. The paper reports that most models hover around the 20% random baseline, and no model's performance exceeds the 33.3% threshold that would result from a perfect bias toward one of the three correct options.
-
Full metric table at Level 0. F1 scores were 24.07 (GPT-4o), 21.59 (o4-mini), 16.38 (o3), 28.07 (Claude-sonnet-4-0), 22.06 (Claude-opus-4-0), 26.34 (Gemini-2.5-flash), and 27.96 (Gemini-2.5-pro). Precision was 33.40, 28.28, 35.36, 42.51, 21.50, 48.44, and 33.11 respectively. Recall equals accuracy in every row. Precision, recall, and F1 are weighted averages of per-class metrics, weighted by the number of true samples in each class.
-
Stronger reasoning correlates with lower accuracy. The OpenAI family follows GPT-4o > o4-mini > o3; in the Anthropic family Claude-sonnet-4-0 outperforms Claude-opus-4-0; in the Google family Gemini-2.5-flash beats Gemini-2.5-pro. The authors conclude that merely scaling up models is insufficient and may even exacerbate the deficit.
-
Answers are driven by intrinsic bias, not the image. Confusion matrices show stable option preferences that repeat within model families — both Claude 4 models favor option C, followed by A and D; both Gemini 2.5 models strongly prefer option C. GPT-4o favors A and B, then C; o4-mini is closer to uniform (consistent with its 19.67% accuracy); o3 shows a very strong bias toward option A. Prediction distributions were nearly identical across ground-truth categories, meaning the selection strategy does not change with the input image.
-
Bias is a stable model characteristic. Because the authors deliberately retained each model's default temperature rather than setting it to 0, they argue the observed preference patterns reflect stable probabilistic tendencies rather than deterministic artifacts.
-
The failure is architectural, not a capacity problem. The paper attributes the deficit to a systemic information bottleneck: lossy tokenization, inductive biases like fixed input resolutions and aspect ratios requiring resizing or patching, token-reduction techniques, and cross-modal alignment mismatch — plus training objectives that reward descriptive language over retention of all visual information.
Methodology in Plain English
The authors needed a task where you cannot cheat by spotting a local detail. Topology is ideal for this: whether a shape is connected, how many holes it has, and what is inside versus outside are all properties that survive stretching and bending and do not depend on local features.
To generate stimuli, they build a uniform spanning tree on a connected graph, where each graph node corresponds to a 3 × 3 pixel block; a graph with n × n nodes yields an image of resolution (4n+1) × (4n+1). Difficulty is set by the partitioning granularity, defined as (n−7)/2, so coarser partitions make weaker demands on a model's information retention and finer partitions make stronger ones.
The sample space grows asymptotically exponentially, roughly exp(hn²), where h ≈ 1.16624 is the lattice tree entropy constant — so even if the benchmark were later used in training, extended difficulty levels could not be memorized.
Each question uses a fixed text prompt and a fixed set of five options. Options B, C, and D correspond to the three topological categories; A and E are distractors that keep the question closed-form and help diagnose guessing versus option bias. A score near 20% suggests random guessing across all five options; near 33.3% suggests a preference within the valid set.
For each difficulty level and each category, they randomly sampled 100 images, giving 300 samples per granularity level. Models were queried through standard API calls with default inference settings and were allowed open-ended text responses rather than being forced to output a single letter.
Why This Matters
Impact on research: The paper argues that the field's benchmark ecosystem measures semantic interpretation, which conflates perception, reasoning, and language generation, and can mask fundamental perceptual weaknesses. TopoPerception separates global perception from textual reasoning and offers a diagnostic that is not vulnerable to the local shortcuts the authors describe in shape-based and maze-based benchmarks. The finding that stronger reasoning models score lower raises concerns about the interplay between chain-of-thought prompting and visual grounding, and suggests that prompting step-by-step reasoning is not always beneficial for tasks requiring strict perceptual fidelity.
Real-world applications:
- Medical imaging and radiology, where a diagnosis may depend on the overall spatial relationship between structures rather than a single salient region.
- Autonomous driving and robotics, where understanding the global layout of a scene matters for safe navigation.
- Document and diagram understanding, where the meaning often resides in connectivity and containment relationships rather than in local glyph features.
- Satellite and aerial imagery analysis, where global structure of terrain and networks carries critical information.
Industry relevance: The results suggest that simply attaching a fixed visual encoder to an LLM through a minimal interface may be insufficient for tasks requiring deeper image understanding. Companies building multimodal products on current architectures may be shipping models whose advertised visual understanding rests on biases and language priors rather than genuine global perception.
Future Directions
- Reconsidering multimodal architecture. The paper calls for more expressive or iterative visual encoders, or mechanisms that let a model re-examine the image during reasoning, rather than "stitching" a fixed encoder to an LLM through a minimal interface.
- Calibrating reasoning against perception. Since stronger reasoning can degrade performance, the authors suggest a mechanism for "fact-checking" the model's reasoning against the visual input at each step, and note that a more direct vision-to-decision mapping, perhaps akin to a classification head, might be more effective than step-by-step prompting for strict perceptual fidelity.
- Separating capacity from training objectives. The paper finds no clear correlation suggesting larger models perform better, so an open question is which training paradigms — rather than which scale — would preserve global visual structure.
- Extending the difficulty hierarchy. Because the synthetic image sample space grows asymptotically exponentially and cannot be memorized, TopoPerception can be extended to arbitrary difficulty levels, leaving open how far any improved model could climb.
Target Audience
Researchers and engineers working on multimodal models, particularly those designing or auditing vision-language benchmarks, as well as teams selecting LVLM architectures for applications where global scene understanding matters. It is also relevant to anyone studying the relationship between chain-of-thought reasoning and visual grounding. Readers without a background in vision-language modeling will still follow the core argument, though they may need to look up terms such as visual encoder, projection module, and token reduction.
Authors’ abstract
Large Vision-Language Models (LVLMs) typically align visual features from an encoder with a pre-trained Large Language Model (LLM). However, this makes the visual perception module a bottleneck, which constrains the overall capabilities of LVLMs. Conventional evaluation benchmarks, while rich in visual semantics, often contain unavoidable local shortcuts that can lead to an overestimation of models' perceptual abilities. Here, we introduce TopoPerception, a benchmark that leverages topological properties to rigorously evaluate the global visual perception capabilities of LVLMs across various granularities. Since topology depends on the global structure of an image and is invariant to local features, TopoPerception enables a shortcut-free assessment of global perception, fundamentally distinguishing it from semantically rich tasks. We evaluate state-of-the-art models on TopoPerception and find that even at the coarsest perceptual granularity, all models perform no better than random chance, indicating a profound inability to perceive global visual features. Notably, a consistent trend emerge within model families: more powerful models with stronger reasoning capabilities exhibit lower accuracy. This suggests that merely scaling up models is insufficient to address this deficit and may even exacerbate it. Progress may require new training paradigms or architectures. TopoPerception not only exposes a critical bottleneck in current LVLMs but also offers a lens and direction for improving their global visual perception. The data and code are publicly available at: https://github.com/Wenhao-Zhou/TopoPerception.