Skip to content
AI.info

Research

Spot The Ball: A Benchmark for Visual Social Inference

Spot the Ball: A Benchmark for Visual Social Inference Overview Research area: Computer vision / vision-language models; specifically visual social inference and theory-of-mind-style reasoning from pe

Spot The Ball: A Benchmark for Visual Social Inference
arXiv
2511.00261
Published
2025-10-31
Authors
Neha Balamurugan, Sarah Wu, Adam Chun, Gabe Gaw, Cristobal Eyzaguirre, Tobias Gerstenberg

AI summary

Spot the Ball: A Benchmark for Visual Social Inference

Overview

Research area: Computer vision / vision-language models; specifically visual social inference and theory-of-mind-style reasoning from perceptual cues.

Technical level: Intermediate. The benchmark design and metrics are accessible without deep math, though the evaluation section uses distributional distance measures (Wasserstein distance, normalized entropy) that assume some familiarity with statistics.

Scope: The paper introduces a 150-image, three-sport benchmark in which the ball has been removed from a photograph and both humans and vision-language models must localize where it was, testing whether models can infer hidden objects from social cues such as gaze, pose, and spatial positioning.

What This Paper Is About

The core problem is that vision-language models (VLMs) are increasingly deployed in visually rich social settings, yet it is unclear whether they can extract and combine subtle behavioral cues the way humans do. The paper's goal is to isolate one component of that ability — visual social inference, defined here as inferring hidden elements of a scene from cues like gaze, pose, and orientation — and to measure it cleanly. The authors do this by adapting a classic newspaper puzzle: given a sports photograph with the ball removed, predict which grid cell the ball occupied.

Key Contributions

  1. A curated evaluation set with human baselines: 150 sports images (soccer, basketball, volleyball) drawn from public YouTube footage, procedurally selected and manually verified, with the ball removed by inpainting and its ground-truth location recorded. Human baselines were collected from a final sample of N=150 participants (50 per sport).

  2. A systematic VLM evaluation: Four state-of-the-art models — Gemini-2.0-flash-001, GPT-4.1-mini, LLaMA-3.2-11B-Vision-Instruct, and Qwen-2.5-VL-7B-Instruct — tested under three prompting strategies (Base, Cue-Directed, Chain-of-Thought) across three sports.

  3. A scalable generation pipeline: A modular YouTube-to-dataset pipeline (video retrieval, CLIP-based frame filtering, YOLOv8 ball and player detection, Stable Diffusion inpainting, 6×10 grid overlay) that produced 3,000 additional soccer images beyond the curated evaluation set.

  4. A structured behavioral analysis: Beyond accuracy, the paper measures Euclidean error, Wasserstein distance to human predictions, near-player rate, overlap rate, center ratio, and normalized entropy, plus an embedding-based analysis of whether model rationales discuss pose versus gaze.

Main Findings

  • Humans substantially outperform all models. Across sports, human accuracy ranged from 19–34% (the abstract states 20–34%), while all four models stayed at or below 17%. The authors describe humans as two to three times more accurate.

  • The gap is not just close misses. Euclidean errors show model predictions are often far from the true location. In volleyball, where humans were most precise (72.0 ± 40.1 pixels), model error was roughly twice as large on average.

  • Sports difficulty differs between humans and models. Humans performed best in basketball, worse in volleyball, and worst in soccer. Models performed similarly in basketball and soccer but struggled most in volleyball.

  • Richer prompting does not reliably help. Cue-Directed prompting (explicitly directing attention to gaze and pose) yielded some improvements in some cases, but gains were inconsistent. Chain-of-Thought prompting sometimes degraded performance — for example, LLaMA reached 272.6 ± 50.7 pixels of error in volleyball under CoT, and GPT performed worse under CoT than under Base or Cue in soccer and basketball. Gemini performed best with CoT in soccer but worst with it in basketball.

  • Models over-rely on a "guess near a player" heuristic. Models (90%) were more likely than humans (65–75%) to predict the ball lies near a player. This fails in volleyball, where the ball is struck rather than held and often travels away from players. For reference, 52.2% of ground-truth balls were near players, and 20.9% were near players by overlap.

  • Both humans and models are center biased, and humans are more dispersed. Ground-truth balls fell in the center window 36.3% (soccer), 46.2% (volleyball), and 56.3% (basketball) of the time. Humans showed high center bias (R = 1.63) and higher normalized entropy (0.855) than all models, which ranged from 0.698 to 0.808. The authors interpret this as humans considering more possibilities than models.

  • Distributional alignment differs by model type. Open-source models (Qwen, LLaMA) produced guess distributions less similar to human responses than proprietary models (GPT, Gemini). All models still outperformed a uniform-distribution baseline. Entropy analysis helps distinguish under-dispersion from misplaced probability mass.

  • Models talk about pose more than gaze; CoT shifts this but does not improve accuracy. Embedding analysis using Google's Gemini embedding model (models/embedding-001) against pose and gaze reasoning templates showed models disproportionately rely on pose cues across all sports, while humans use gaze and pose relatively evenly. CoT prompting increased mentions of gaze in reasoning text but did not produce consistent accuracy gains — models described more human-like reasoning without predicting more accurately.

  • Three recurring failure modes were identified: neglect of gaze cues despite strong evidence of ball location; role confusion, where models misidentify which player has possession or is about to act; and a default-to-center heuristic, such as placing a volleyball prediction at the net, an unlikely location since play would have terminated.

Methodology in Plain English

Building the data. The authors started from publicly available sports broadcast footage on YouTube, retrieved with sport-specific queries using action-focused keywords. Videos were decoded with OpenCV and sampled at roughly 1 FPS. Each frame was scored with CLIP against prompts like "picture of volleyball players in action with ball," and only frames above a similarity threshold were kept. Frames then went through YOLOv8 to detect balls and players, filtered by confidence and spatial plausibility — requiring exactly one ball per frame, non-overlapping with and proximal to players. The ball region was removed and filled using Stable Diffusion inpainting, with player masks preserving body posture and gaze. Images were manually checked to remove leftover ball shadows or artifacts. Finally, a 6×10 alphanumeric grid (rows A–F, columns 1–10, so 60 cells total) was overlaid, and ground-truth labels were the one or more cells covering the original ball location.

The task. Every participant — human or model — selects one grid cell and provides text reasoning. Multiple adjacent cells can count as correct if they overlap the ball region.

Human study. 176 participants were recruited from Prolific and paid $12/hour base plus up to $1 for accuracy. After excluding 26 for failed attention checks, the final sample was N=150 (50 per sport). Each participant saw 52 images (50 test images plus 2 attention checks with visible balls) from a single sport in randomized order, making three guesses per test image by clicking grid cells. Average completion time was 17.3 minutes (SD=7.4). The experiments were pre-registered on the Open Science Framework, and the volleyball and basketball conditions were identical to the preregistered soccer experiment. Humans received only the Base Prompt.

Model runs. Models were tested under three prompts: a Base Prompt asking for the cell location; a Cue-Directed Prompt adding an instruction to consider player gaze, pose, and positions; and a Chain-of-Thought Prompt that first asks three one-shot questions about where players are, where they are looking, and how they are positioned, then feeds those responses back as context for the final prediction. To estimate distributional behavior, the authors sampled n=50 predictions per image for Base and Cue-Directed and n=20 for Chain-of-Thought, all at temperature T=0.6.

Measuring behavior. Performance was measured by accuracy (whether the predicted cell is in the ground-truth set) and Euclidean error (pixels from predicted cell center to the nearest ground-truth cell center). Alignment with humans was measured with Wasserstein distance (Earth Mover's Distance) between model and human prediction distributions. Behavioral strategies were captured by Near Player Rate (fraction of predictions within τ=0.08 of the image diagonal from any player box), Overlap Rate (fraction of predictions overlapping a player box by at least θ=0.02 of the cell area), Center Ratio (prediction mass versus ground-truth mass in a 3×5 central region), and normalized entropy over the 60 cells. The thresholds τ=0.08 and θ=0.02 were fit via grid search over τ ∈ [0.04, 0.20] and θ ∈ [0.01, 0.20] using a balanced objective equally weighting NR and OR; the objective showed a broad plateau, with neighboring configurations within 1% of the maximum.

Why This Matters

Impact on research. The paper isolates a capability — inferring hidden objects from social cues — that existing benchmarks miss, because prior visual social-reasoning benchmarks tend to use fully observable scenes or focus on physical occlusion without social cues. It provides the first structured evaluation of whether VLMs can use pose, gaze, and spatial structure to infer hidden objects in real-world scenes, along with released data, pipeline, and evaluation code.

Real-world applications:

  • Embodied AI and robotics: A robot nurse that misinterprets a patient's gesture, or an autonomous car that fails to anticipate a pedestrian's intention, may behave unsafely — the paper explicitly frames social-cue interpretation as safety-critical.
  • Human-robot collaboration: Systems that must infer intent from gaze and body orientation rather than explicit instructions.
  • Sports analytics and broadcast augmentation: Automatically reasoning about player intent and likely ball trajectories.
  • Surveillance and scene understanding: Inferring what agents are attending to or reaching for when the target object is not directly visible.

Industry relevance. The finding that proprietary models outperform open-weight models on distributional similarity to humans, and that prompting strategies (including Chain-of-Thought) fail to close the gap, is directly relevant to anyone deploying VLMs in interactive or visually rich environments. The result that models use shallow heuristics — near-center and near-player guessing — suggests the limitation is architectural rather than a matter of prompt engineering or missing world knowledge, given that ball sports are heavily represented in pretraining data.

Future Directions

  1. Architectures that explicitly encode behavioral cues. The authors argue progress may require integrating perceptual priors or architectures explicitly designed to capture agentive and relational dynamics, rather than relying on prompting.

  2. Temporal information. The current benchmark uses static images specifically to isolate social reasoning from motion dynamics; adding video could supply additional cues that models currently lack.

  3. Pipeline extension and difficulty control. The modular pipeline allows extension to other ball sports and to difficulty controls such as varying player density or occlusion severity. The paper demonstrates this by producing 3,000 additional soccer images for training and analysis.

  4. Understanding why richer descriptions do not produce better predictions. CoT prompting made models mention gaze more often and describe more human-like reasoning, yet did not improve accuracy — an open question about the gap between stated reasoning and actual inference.

Target Audience

This paper is most useful for computer vision and multimodal machine learning researchers working on VLM evaluation, benchmark design, and social or theory-of-mind reasoning in AI. It is also relevant to cognitive scientists studying visual social inference and human behavior, to robotics and embodied-AI practitioners who need models to interpret human pose and gaze in safety-critical settings, and to dataset or evaluation engineers interested in the paper's scalable generation pipeline and behavioral metrics (center ratio, entropy, near-player rate) as reusable evaluation tools.

Authors’ abstract

Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This ability drives everyday social reasoning in humans and is critical for developing more human-like AI agents. We introduce Spot The Ball, a challenging benchmark for evaluating visual social inference in vision-language models (VLMs) using sports as a test domain. The task is to localize a removed sports ball from soccer, basketball, and volleyball images. We present a curated evaluation set with human baselines and a scalable pipeline for generating additional test items. We evaluate four state-of-the-art VLMs (Gemini, GPT, LLaMA, Qwen) using three prompting strategies, finding that humans are consistently two to three times more accurate (20-34%) than models ($\leq$ 17%) across all sports. Our analyses show that models rely on superficial spatial heuristics--such as guessing near the image center or nearby players--while humans leverage social cues like gaze direction and body pose. These findings reveal a persistent human-model gap in visual social reasoning and underscore the need for architectures that explicitly encode structured behavioral cues to achieve robust, human-like inference.

Read the original paper