Skip to content
AI.info

Research

ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments

Overview Research area: Computer vision and human visual attention prediction, specifically object-based attention modeling in interactive virtual-reality street-crossing environments. Technical level

ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments
arXiv
2601.13218
Published
2026-01-19
Authors
Igor Vozniak, Philipp Mueller, Nils Lipp, Janis Sprenger, Konstantin Poddubnyy, Davit Hovhannisyan, Christian Mueller, Andreas Bulling, Philipp Slusallek

AI summary

Overview

Research area: Computer vision and human visual attention prediction, specifically object-based attention modeling in interactive virtual-reality street-crossing environments.

Technical level: Advanced. The paper assumes familiarity with saliency prediction metrics, state-space sequence models (Mamba), U-Net-style encoder-decoder architectures, and graph neural networks.

Scope: The paper introduces a 120-participant VR street-crossing gaze dataset with panoptic, depth, and skeleton annotations, a new object-level evaluation metric (oSIM), and a graph-augmented attention prediction model (SUMGraph).

What This Paper Is About

Cognitive science has long held that human attention operates on objects, not just spatial locations, but computational attention models are still evaluated almost entirely on pixel-level spatial agreement. The reason is practical: object-based evaluation needs accurate per-object segmentations, which most existing gaze datasets lack. This paper builds a VR dataset where segmentation is available by construction, defines a metric that scores predictions per object rather than per pixel, and trains a model that uses explicit object information (vehicle skeletons plus speed, distance, and direction) as part of its input.

Key Contributions

  1. ObjectVisA-120 dataset: A 120-participant VR street-crossing attention dataset containing 7,200 videos and 6.14M frames (6,142,167 frames sampled from a 90Hz recording at every 3rd frame, i.e. 30Hz), with per-participant gaze fixations (2D/3D), derived saliency maps, panoptic segmentation following the CityScapes labeling policy, depth maps, and vehicle skeleton labels (24 keypoints, matching the sparse ApolloCar3D variation used in OpenPifPaf).

  2. Object-based Similarity (oSIM): A metric that compares predicted and ground truth attention by aggregating saliency within each panoptic object mask and taking the minimum per mask, so that every object contributes equally regardless of its pixel size. It is also added as a training loss term.

  3. SUMGraph model: A Mamba U-Net-based attention prediction model that extends the SUM architecture with Graph VSS and Graph C-VSS blocks, encoding vehicles in the field of view as graphs (up to 24 keypoints as nodes, plus global attributes of speed, distance, and direction) and fusing that context into image features.

  4. Empirical demonstration: Fine-tuned models, the oSIM loss, and the graph structure are evaluated together, with the authors reporting that SUMGraph outperforms prior methods in 28 out of 30 metrics and that adding the oSIM loss improves results in 25 out of 30 metrics.

Main Findings

  • Graph model leads on most metrics: The middle block of Table II shows SUMGraph achieving CC 0.4564, KLD 1.6747, AUC 0.9683, SIM 0.3568, NSS 6.4357, and oSIM 0.6086, which the authors describe as the best performance among the compared fine-tuned models on KLD, AUC, SIM, and oSIM.

  • Baseline SUM remains strongest on CC and NSS among fine-tuned models: SUM reports CC 0.4722 and NSS 6.2607, versus SUMGraph's CC 0.4564 and NSS 6.4357; ContextSalNet with the SUM loss reports the highest NSS at 6.6447 and the second-highest oSIM at 0.6042.

  • Explicit graph scaling changes the trade-off: "SUMGraph (Ours) & Scale" reports CC 0.4643, KLD 1.6581, AUC 0.9683, SIM 0.3481, NSS 6.5241, and oSIM 0.6025, improving CC, KLD, and NSS relative to the default SUMGraph while lowering SIM and oSIM.

  • Global attributes matter: Removing the global attributes from the graph ("SUMGraph (Ours) (no global attr.)") gives CC 0.4508, KLD 1.6729, AUC 0.9681, SIM 0.3482, NSS 6.3898, and oSIM 0.6019, which the authors describe as worse than the version with global attributes.

  • Fine-tuning is essential: Without fine-tuning on ObjectVisA-120, models perform poorly — SUM scores CC 0.2961, KLD 2.6820, AUC 0.9160, SIM 0.2003, NSS 3.0513, oSIM 0.4304; ContextSalNet scores CC 0.0093, KLD 3.9165, AUC 0.6177, SIM 0.0406, NSS 0.0760, oSIM 0.3484. The authors attribute this to the gap between classical saliency estimation and sparse attention prediction in this setting.

  • oSIM loss improves most metrics: Reported deltas from adding the oSIM loss include SUM (+0.0019 AUC, +0.0001 SIM, +0.1112 NSS, +0.0097 oSIM, −0.0034 CC, −0.0424 KLD), TranSalNet (+0.0135 CC, −0.0234 KLD, +0.0002 AUC, +0.0085 SIM, +0.1648 NSS, +0.0072 oSIM), and SUMGraph (−0.0023 CC, −0.0130 KLD, +0.0001 AUC, +0.0062 SIM, +0.0148 NSS, +0.0045 oSIM).

  • Object-level scoring differs from pixel-level scoring: The paper's Figure 2 illustrates that predictions staying inside the correct object receive a much higher oSIM than SIM, since oSIM ignores within-object spatial misalignment.

  • Dataset scale compared to prior VR work: Table I lists ContextSalNet at 11 participants and 35K frames and HOT3D at 19 participants and 3.7M frames, against ObjectVisA-120's 120 participants and 6.14M frames; the authors note that none of the previous datasets provide panoptic segmentation.

Methodology in Plain English

The researchers first ran an immersive street-crossing study in the Unity 3D engine with bidirectional traffic, varying traffic density, presence or absence of crosswalks, and non-playable characters behaving either riskily or cautiously. Participants wore HTC Vive Pro Eye head-mounted goggles with a wireless adapter inside a 9 × 8 meter tracking footprint, and each of the 120 participants completed 60 street-crossing trials after a warm-up phase and eye-tracking calibration. Demographic spread is reported as an even German/Japanese split, ages 20–50 (mean 30.56, SD 9.02), 60 male and 60 female, height mean 172.46 cm (SD 9.47 cm), weight mean 66.18 kg (SD 14.57 kg), driving experience 0–32 years (mean 9.42, SD 8.85), and VR familiarity of 81 yes versus 39 no.

For the dataset, they sampled every third frame, and generated ground truth attention maps from the last three fixation points using Gaussians with a standard deviation of 3 × dva, citing an approximate 2° foveal eccentricity and VR eye-tracking accuracy of 0.5°–1.1°. Objects occluded in the view are excluded from all labels, matching real-world perception.

The oSIM metric works by taking, for each object mask, the smaller of the total predicted saliency and the total ground truth saliency inside that mask, then summing over masks. For the model, each visible vehicle becomes a graph: its keypoints are nodes, and speed, distance, and direction are global attributes. Two graph convolutions refine the node features, a mean pool combines them with a projected attribute vector into one embedding per object, and these are averaged into a scene context vector that is fused with image features inside novel Graph VSS and Graph C-VSS blocks. If no relevant objects are in the field of view, the system falls back to the baseline SUM model.

Training used PyTorch on 10 × A100 (80vGB) GPUs for 15 epochs with early stopping after 4 epochs, Adam with initial learning rate 1 × 10⁻⁴ and a scheduler that reduces it tenfold after four epochs, distributed data-parallel training with overall batch size 750, and 256 × 256 resolution. Data was split 70/10/20 into 84, 12, and 24 participants, balanced for gender and nationality and checked with a Kolmogorov–Smirnov test on age and height. The loss combines KLD, CC, SIM, NSS, MSE, and oSIM with weights λ₁ = 10, λ₂ = −2, λ₃ = −1, λ₄ = −1, λ₅ = 1, and λ₆ = −1. SUMGraph initializes from SUM weights pre-trained on six datasets.

Why This Matters

Impact on research: The paper argues that object-based attention has been held back by the absence of suitable datasets, metrics, and loss functions, and it supplies all three in one package — a synthetic dataset with instance-accurate panoptic labels, a metric that scores per object, and a model that consumes explicit object representations. It also challenges the field's default assumption that better pixel-level saliency automatically means better modeling of human attention.

Real-world applications:

  • Pedestrian safety systems that need to know whether a person crossing a street is likely looking at an approaching vehicle, rather than where their gaze lands on a pixel grid.
  • Autonomous driving and intelligent vehicle research, which is the venue this work was accepted to (IEEE Intelligent Vehicles Symposium, IV 2026).
  • Human-robot collaboration, which the authors explicitly name as a domain oSIM can be extended to.
  • Driver attention monitoring, also named by the authors as an extension target.
  • Training and simulation of street-crossing behavior, aided by the 2D/3D bounding boxes and additional modalities the saved state space allows the authors to generate.

Industry relevance: The safety-critical framing (vehicle detection while crossing) maps directly onto ADAS and autonomous vehicle perception stacks, and the authors note that the same skeleton representations could be obtained for real-world use by fine-tuning on the ApolloCar3D dataset, or by generating segmentation labels on demand with state-of-the-art methods.

Future Directions

  • Generalization beyond urban daytime scenes: The authors state that the training data is predominantly urban-centric and may not generalize to rural areas, extreme weather, or other daylight conditions.
  • Age generalization: The participants were aged 20–50, so behavior for other age groups remains uninvestigated.
  • Cultural generalization: Only German and Japanese participants were included, and the authors call for a wider scope of backgrounds.
  • Extending oSIM beyond panoptic-level objects: The paper notes that oSIM supports hierarchical, part-level segmentation, allowing objects to be decomposed into semantically meaningful parts such as pedestrian body parts.
  • Extending graph representations: The authors state that incorporating pedestrian skeletons is beneficial because of their relevance for behavioral and intent inference, while extending graphs to static scene elements offers limited value and unnecessary computation — leaving open where the boundary should be drawn.
  • Cross-domain deployment: The paper leaves open how the approach transfers from synthetic VR data to real-world recordings, mentioning on-demand segmentation labeling and fine-tuning on ApolloCar3D as possible paths.

Target Audience

Researchers and engineers working on visual attention prediction, saliency modeling, and gaze estimation; autonomous driving and intelligent vehicle safety teams interested in pedestrian intent; and practitioners of VR-based human behavior studies who need annotated gaze datasets. It is also relevant to anyone evaluating attention models who wants to move beyond pixel-level metrics, and to graph-learning researchers curious about applying explicit object graphs inside state-space model architectures.

Authors’ abstract

The object-based nature of human visual attention is well-known in cognitive science, but has only played a minor role in computational visual attention models so far. This is mainly due to a lack of suitable datasets and evaluation metrics for object-based attention. To address these limitations, we present ObjectVisA-120 -- a novel 120-participant dataset of spatial street-crossing navigation in virtual reality specifically geared to object-based attention evaluations. The uniqueness of the presented dataset lies in the ethical and safety affiliated challenges that make collecting comparable data in real-world environments highly difficult. ObjectVisA-120 not only features accurate gaze data and a complete state-space representation of objects in the virtual environment, but it also offers variable scenario complexities and rich annotations, including panoptic segmentation, depth information, and vehicle keypoints. We further propose object-based similarity (oSIM) as a novel metric to evaluate the performance of object-based visual attention models, a previously unexplored performance characteristic. Our evaluations show that explicitly optimising for object-based attention not only improves oSIM performance but also leads to an improved model performance on common metrics. In addition, we present SUMGraph, a Mamba U-Net-based model, which explicitly encodes critical scene objects (vehicles) in a graph representation, leading to further performance improvements over several state-of-the-art visual attention prediction methods. The dataset, code and models will be publicly released.

Read the original paper