Research
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models Overview Research area: Computer vision, specifically the evaluation of spatial reasoning in Visual Found

- arXiv
- 2601.11729
- Published
- 2026-01-16
- Authors
- Turhan Can Kargin, Wojciech Jasiński, Adam Pardyl, Bartosz Zieliński, Marcin Przewięźlikowski
AI summary
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation ModelsOverview
- Research area: Computer vision, specifically the evaluation of spatial reasoning in Visual Foundation Models (VFMs).
- Technical level: Intermediate. The paper assumes familiarity with Vision Transformers, self-supervised learning paradigms (Joint-Embedding, Masked Image Modeling), and probing protocols, but the core question it asks is conceptually simple.
- One-sentence scope: The paper introduces a synthetic, Unreal Engine 5-based benchmark that tests whether frozen VFM representations encode the directional relation ("left", "right", "front", "back") between two objects, either from the camera's point of view or from a third object's point of view.
What This Paper Is About
Visual Foundation Models such as DINO and CLIP are strong at recognizing what is in an image, but it is unclear whether they understand where things are relative to one another, which matters for embodied agents that must act and navigate. Recent work has tried to fix this by adding 3D objectives (depth, camera pose, surface normals) during training, yet improvements are inconsistent across tasks, raising the question of whether these models genuinely acquire spatial awareness or merely overfit to the specific geometric objectives they were trained on. SpaRRTa addresses this by isolating a simpler, abstract capability: recognizing the static directional relation between two visible objects from a specified viewpoint.
Key Contributions
- The Tool: SpaRRTa, a photorealistic, synthetically generated evaluation environment built on Unreal Engine 5 for probing directional spatial relation recognition in VFMs, with fully controllable object arrangements and freely accessible spatial annotations.
- The Benchmark: Two standardized challenge sets — SpaRRTa-ego (camera-centric viewpoint) and SpaRRTa-allo (allocentric viewpoint, requiring implicit perspective-taking) — designed to disentangle spatial logic from semantic recognition.
- The Insights: A benchmark of VFM families across supervision regimes (MIM, Joint-Embedding, Supervised, Vision-Language) characterizing which models carry spatial information in their representations and how those representations are structured.
- The Probing Analysis: Evidence that spatial information lives primarily at the patch level and is largely obscured by global pooling, motivating structured aggregation for recovering it.
Main Findings
-
Task formulation: SpaRRTa is a four-way classification problem with discrete labels {left, right, front, back}, predicting the direction from a source object to a target object relative to a chosen viewpoint.
-
Ambiguity is procedurally eliminated: Scenarios where the target falls within ±15° angular regions centered along the diagonals (45°, 135°, 225°, 315°) relative to the viewpoint's forward direction are discarded via rejection sampling, guaranteeing indisputable ground-truth labels.
-
Self-supervised models beat supervised and vision-language models: The paper reports that self-supervised VFMs consistently outperform supervised and vision-language models on SpaRRTa.
-
Efficient probing substantially outperforms linear probing: With linear probing, most models achieve "reasonable" performance, but Efficient Probing yields large gains — particularly for MIM models such as MAE, which retain rich spatial information in patch tokens that global average pooling discards. For example, MAE reaches 93.10 (Forest), 93.71 (Desert), 93.83 (Bridge), and 93.26 (Winter Town) with Efficient Probing in the egocentric task.
-
VGGT under Efficient Probing is state-of-the-art on the egocentric task: VGGT (ViT-L/14, trained with direct 3D supervision on cameras, depth and tracks) achieves the best Efficient Probing result in every environment: 96.18 (Forest), 98.78 (Desert), 95.53 (Winter Town), 96.11 (Bridge), and 94.47 (City). Its mean rank under Efficient Probing is 1.00.
-
DINO-v2 reg (L/14) is the strongest non-3D-supervised model in several settings: It reaches 93.92 (Forest), 96.70 (Desert), 94.92 (Winter Town), 96.11 (Bridge) and 94.27 (City) with Efficient Probing, and holds mean rank 1.00 for linear probing and 1.20 for allocentric linear probing.
-
CLIP performs worst overall: CLIP's Efficient Probing accuracies are 56.33 (Forest), 68.44 (Desert), 63.21 (Winter Town), 64.71 (Bridge) and 64.67 (City) in the egocentric task, giving it mean rank 13.00 (linear, egocentric) and 12.80 (efficient, allocentric). DeiT is also consistently weak, with Efficient Probing allocentric mean rank 12.00.
-
Allocentric perspective-taking causes a large drop: The allocentric task requires implicit perspective transformation, and performance falls sharply versus egocentric. Best allocentric Efficient Probing results are VGGT at 76.91 (Forest), 77.74 (Desert), 79.27 (Winter Town), 78.15 (Bridge) and 71.19 (City), with mean rank 1.20; DINO-v2 reg (L/14) reaches 77.60 on Bridge and is the best allocentric Efficient Probing model on City at 73.61.
-
Linear probing collapses on the allocentric task: Linear probing fails for most models, hovering near the random baseline, indicating that a naive global average pooling of patch tokens is insufficient for perspective-dependent reasoning.
-
Spatial information is patch-level, not global: Comparing probing mechanisms reveals that spatial information is primarily stored at the patch level and is largely obscured by global pooling, highlighting the importance of structured aggregation.
-
Even the strongest models have limits: The paper states that even the best models exhibit limitations in perspective-dependent and cluttered settings.
-
VGGT vs. DINO-v2 reg (L/14) isolates the effect of 3D supervision: Because VGGT was initialized from DINO-v2 reg parameters, comparing them directly analyzes the representational shift caused by adaptation to 3D perception tasks.
-
Not reported in the available content: Sections 4.3 (comparison of SpaRRTa performance against other spatial-awareness and image-recognition tasks) and 4.4 (analysis of inner representations) are described in the text but their results are not included in the provided content. Detailed dataset size statistics are referenced to Section A.1 but are likewise not present in the provided content.
Methodology in Plain English
The authors built a pipeline with two halves: generating images, and probing models on them.
Generation. An evaluation controller uses Unreal Engine 5.5 through its Python API to set up a scene — selecting an environment and 3D assets, then randomly placing a source object, a target object, and (for the allocentric variant) a viewpoint object on the ground, with positions sampled from a Gaussian distribution. The camera is sampled uniformly over an area surrounding the map center and oriented toward the objects. The geometric relationship is validated, and any configuration where the target lands within ±15° of a diagonal decision boundary is discarded. Unreal then renders a ray-traced RGB image, captured alongside a segmentation mask using the UnrealCV plugin. The objects are chosen from a curated asset library aligned with common ImageNet super-categories — Animals, Everyday Objects, Nature and Humans — so that failures reflect geometry rather than a failure to recognize the object.
Five environments. Forest (based on the Electric Dreams Environment demo), Desert, Winter Town, Bridge, and City (derived from the City Sample demo), spanning sparse organic landscapes to dense structured urban geometry, with different textures, lighting and layouts.
Probing. The rendered image goes into a frozen VFM backbone; only a lightweight head is trained, so the evaluation measures intrinsic pre-trained representations rather than adaptability. Three heads are compared: linear probing on Global Average Pooling of patch tokens; AbMILP, which learns a scalar attention weight per patch via a 2-layer MLP to filter background noise; and Efficient Probing, which uses multi-query cross-attention with 4 learnable queries to produce diverse attention maps. Features are taken from the final transformer block, with intermediate layers analyzed in Appendix D. All models use a ViT-Base backbone at 224×224 resolution, except VGGT (ViT-L/14) and DINO-v2 reg ViT-L/14, which are included for a fair large-backbone comparison.
Protocol. For each (source, target, viewpoint) triple, the authors build a dataset of different object layouts and split it 80/10/10 into train, validation and test folds at the level of rendered samples, so no identical rendered image or metadata instance appears in more than one fold. This tests generalization to unseen spatial configurations within the same environment and object triple, not to held-out environments or unseen object categories; the paper notes that splits may still share environment-specific patterns such as background appearance, lighting or scene structure. Probes are trained with AdamW, cosine decay, weight decay 0.001, learning rates of 10⁻², 10⁻³ or 10⁻⁴, dropout of 0.2, 0.4 or 0.6, batch size 256, and linear warmup of 200 (linear probing) or 100 (AbMILP, Efficient Probing) steps. Training runs 1000 epochs for linear probing and 500 epochs for the other two heads. The procedure is repeated with two random seeds and three distinct object triples from each of the five environments.
Hardware. Rendering runs on a Windows workstation with two NVIDIA RTX 2080 Ti GPUs (11 GB VRAM each); probing runs on Linux with a single NVIDIA RTX 4090 (24 GB VRAM).
Why This Matters
Impact on research. The paper reframes how spatial awareness in VFMs should be measured. Instead of precise metric prediction (depth, pose, surface normals), it targets relational understanding — a capability the authors argue underpins more advanced human-like spatial reasoning. This matters because prior work (Probe3D, the Mid-Level Vision Probe) showed that higher ImageNet accuracy does not guarantee stronger 3D reasoning and that performance correlates only moderately across probes, motivating task-agnostic diagnostics. SpaRRTa also provides a controllable synthetic alternative to real-world 3D data collection, with exact ground truth and arbitrarily large scale.
Real-world applications (as framed by the paper):
- Embodied agents and robotics: navigation, planning and object manipulation, where an agent must know where things are relative to itself or others.
- Perspective-taking interaction: allocentric reasoning matters when an agent must interpret a scene from a human's viewpoint, for example in assistive or collaborative settings.
- Semantic perception at scale: the paper cites existing VFM deployment in retail and medical imaging, domains where understanding layout may complement recognition.
- Policy transfer for navigation and object search: geometric and temporal cues are described as important inductive biases for next-generation vision encoders, which task-driven suites such as VC-1 and SPA link to downstream action success.
Industry relevance. The results suggest that adding a 3D objective such as depth estimation is not automatically sufficient — VGGT, trained with direct multi-task 3D supervision, does best, but the gap depends heavily on how representations are read out. For practitioners building embodied or interactive systems, the practical takeaway is that naive global pooling throws away measurable spatial signal that a modest attention-based head can recover, and that vision-language models such as CLIP are a poor choice when relational spatial reasoning is required.
Future Directions
- Closing the perspective-taking gap: Allocentric performance is far below egocentric performance, and linear probing collapses toward the random baseline there; what training signal or architecture would let models implicitly transform viewpoints remains open.
- Improving cluttered and perspective-dependent settings: The authors state that even the strongest models show limitations in these conditions, without fully characterizing them in the available content.
- Better aggregation as a design principle: Since spatial information is patch-level and obscured by global pooling, designing structured aggregation into VFM architectures — not just probes — is a natural next step.
- Linking SpaRRTa to other spatial and semantic capabilities: Section 4.3 is described as comparing SpaRRTa performance with other spatial-awareness and image-recognition tasks; those results are not reported in the provided content, leaving the correlation structure between SpaRRTa and other benchmarks as an open item.
Target Audience
Researchers and engineers working on visual foundation models, representation learning, and embodied AI who need a diagnostic for whether a backbone actually encodes spatial structure. It is also relevant to practitioners deciding between DINO-family, MAE-family, 3D-supervised and vision-language models for robotics or interactive applications, and to benchmark designers interested in synthetic, photorealistic evaluation environments built on Unreal Engine 5. The paper is most accessible to readers already comfortable with Vision Transformers and probing protocols.
Authors’ abstract
Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work incorporates some 3D tasks (such as depth estimation) into VFM training. However, VFM performance remains inconsistent across other spatial tasks, raising the question of whether these models truly have spatial awareness or overfit to specific 3D objectives. To address this question, we introduce the Spatial Relation Recognition Task (SpaRRTa) benchmark, which evaluates the ability of VFMs to identify relative positions of objects in the image. Unlike traditional 3D objectives that focus on precise metric prediction (e.g., surface normal estimation), SpaRRTa probes a fundamental capability underpinning more advanced forms of human-like spatial understanding. SpaRRTa generates an arbitrary number of photorealistic images with diverse scenes and fully controllable object arrangements, along with freely accessible spatial annotations. Evaluating a range of state-of-the-art VFMs, we reveal significant disparities between their spatial reasoning abilities. Through our analysis, we provide insights into the mechanisms that support or hinder spatial awareness in modern VFMs. We hope that SpaRRTa will serve as a useful tool for guiding the development of future spatially aware visual models.