Research
EgoCampus: Egocentric Pedestrian Eye Gaze Model and Dataset
Overview Research area: Computer vision / egocentric vision, specifically eye-gaze prediction for pedestrians navigating outdoor environments, with connections to saliency modeling, embodied AI, and r

- arXiv
- 2512.07668
- Published
- 2025-12-08
- Authors
- Ronan John, Aditya Kesari, Vincenzo DiMatteo, Kristin Dana
AI summary
Overview
Research area: Computer vision / egocentric vision, specifically eye-gaze prediction for pedestrians navigating outdoor environments, with connections to saliency modeling, embodied AI, and robotics.
Technical level: Intermediate. The paper is readable without deep math, but it assumes familiarity with standard gaze/saliency evaluation metrics (AUC-Judd, CC, KLD, SIM, F1) and with pretrained video backbones.
Scope: The paper introduces an outdoor egocentric pedestrian gaze dataset (EgoCampus) collected with Meta's Project Aria glasses and a baseline transformer-free fusion model (EgoCampusNet, "ECN") for predicting where a walking pedestrian will look next.
What This Paper Is About
Most eye-gaze and saliency research studies people looking at static images, watching pre-recorded video, or performing structured indoor tasks such as cooking. The authors argue that gaze during real-world outdoor locomotion — walking around a campus — is largely unstudied, and building embodied agents that share space with humans requires knowing where pedestrians actually look.
The paper's goal is twofold: release a multimodal dataset of pedestrians walking the same outdoor routes while their eye gaze is recorded, and provide a baseline model that predicts gaze heatmaps from egocentric video so future work has something to compare against.
Key Contributions
- EgoCampus, an egocentric video dataset with synchronized eye gaze, RGB video, and auxiliary sensors (IMU, GPS, Wi-Fi), captured during pedestrian locomotion on outdoor campus paths — 82 unique participants, 25 distinct paths spanning 6 km, totaling approximately 32 hours of video (≈3.5 million frames).
- EgoCampusNet (ECN), a gaze prediction model that fuses spatio-temporal features from a pretrained video backbone (primarily X3D) with features learned from a single query frame through light-weight ResNet blocks and a CNN decoder.
- An evaluation of ECN against six existing methods from image saliency, video saliency, and eye gaze prediction, showing that off-the-shelf state-of-the-art models do not generalize well to this dataset without fine-tuning.
- A prior-relative weighting strategy for the evaluation metrics, which uses Jensen-Shannon divergence between a prediction and the dataset gaze prior to down-weight frames whose ground truth is already explained by the strong center bias in the data.
Main Findings
- Strong center bias exists in egocentric locomotion. Under unweighted metrics, both a Dataset Prior (averaged per-frame gaze heatmaps over the test set) and a Center Prior are remarkably competitive, often outperforming every method except ECN and GLC.
- ECN is a competitive but not top-performing baseline. In the unweighted comparison (Table 2), ECN reaches AUC-J 0.972, CC 0.696, KLD 0.816, SIM 0.555, F1 0.614, Recall 0.544, and Precision 0.705 with 42.5M parameters — the smallest parameter count of the learned models compared (GLC 70.2M, EML-NET 47.2M, SUM 57.5M, CV_MM 420.5M, VistaHL 187.7M, DeepGazeIIE 104M). GLC performs best across most metrics.
- Weighted metrics punish center-bias reliance. After prior-relative weighting (Table 3), performance generally drops, particularly for the priors (up to an 18% drop) and for DeepGazeIIE, revealing their heavy reliance on center bias. ECN shows a more modest reduction.
- GLC is the most robust model. Under weighted metrics GLC actually improves on metrics such as F1, by 9%.
- Training cost. ECN trained for 10 epochs in approximately 8 hours on a single NVIDIA RTX 3090 GPU.
- Head-turn attention targets. A manual behavioral analysis using a rotational-velocity threshold of 1.8 rad/s found that high-velocity head movements make up approximately 12.5% of the dataset, and that attention during those turns is dominated by structural landmarks (buildings, trees, lampposts) and navigational cues (pedestrians, paths).
- Qualitative failure modes. EML-Net and DeepGazeIIE tend to overestimate the likely gaze area. Other models often predict that other pedestrians' faces are the most likely place to look, which the authors note is often not the case during egocentric locomotion. ECN learns a strong center-bias contribution, often predicting the image center, which tends to correspond to the direction of travel.
Methodology in Plain English
Data collection. Participants wore Meta's Project Aria glasses, which capture a 1408 × 1408 RGB camera stream at 30Hz, inward-facing eye-tracking cameras, and motion and localization sensors (IMU, magnetometer, barometer, Wi-Fi/Bluetooth, GPS). Each of 82 subjects walked several of 25 predefined campus paths individually, after being shown a bird's-eye map of the start and end points. Paths ranged from 100 to 200 meters, averaging 150 meters. Recordings spanned different times of day, seasons, and weather. Passersby who did not consent were blurred using EgoBlur, following an IRB protocol.
Processing. Raw VRS recordings were temporally aligned across sensor streams. The 1408 × 1408 30Hz RGB video was downscaled to 224 × 224 JPEG frames (both raw and downscaled video are released). The 30Hz eye-tracking coordinates were saved to a NumPy array with one entry per nearest-timestamped RGB frame, and the high-frequency 1000Hz IMU readings were averaged to align with each frame's timestamp.
Model. Given an egocentric video of length T with spatial dimensions H and W, the goal is to predict the most likely gaze point per frame. A pretrained video encoder produces spatio-temporal features; the feature at the time index corresponding to the query frame is passed through ResNet blocks to give a H/4 × W/4 × 96 map. In parallel, the query frame itself is encoded into a same-shaped map, and the two are concatenated along the channel dimension into a H/4 × W/4 × 192 encoded gaze map. A learned CNN decoder upscales this to a full-resolution heatmap; the output is then blurred, the center prior is added, and normalization produces the final predicted gaze heatmap.
Training and evaluation protocol. Models were trained in PyTorch with the ADAM optimizer, MSE loss, and a learning rate of 0.002. Experiments sample 16 frames at a time, uniformly from a 64-frame window (≈ 2.1 seconds); temporal models see the full clip and predict gaze in the last frame, while image-based models see only the last frame. The train/test split is roughly 70/30, with 16 randomly selected paths for training and the remaining 9 for testing.
Why This Matters
Impact on research. The paper fills a reported gap: existing egocentric datasets either focus on structured indoor tasks (EPIC-Kitchens, HD-EPIC, HoloAssist, EgoExo4D, EgoExoLearn, EGTEA Gaze+) or capture outdoor navigation with far fewer subjects and much shorter continuous clips (GEETUP has 43 subjects on two paths with a median clip length of 14.7 seconds, versus EgoCampus's 82 pedestrians and clips averaging 108.2 seconds). Because multiple subjects traverse the same routes, the dataset also supports cross-subject comparisons of gaze in shared spatio-temporal contexts. The authors also release a companion dataset, YOPO-Campus, in which the same 6 km of campus paths were traversed by a tele-operated Clearpath Jackal robot, enabling human-versus-robot comparisons.
Real-world applications:
- Training embodied agents and robots that share space with people, where anticipating pedestrian attention could support navigation and cooperative behavior.
- Developing spatio-temporal environment-sampling methods that mimic human attention, so agents spend computation where humans would look.
- Human-robot interaction research, particularly in outdoor settings where the YOPO-Campus robot-view data can be paired with the pedestrian-view EgoCampus data.
- Benchmarking and diagnosing saliency and gaze models, since the paper shows which model families over-rely on center bias.
Industry relevance. Companies and labs working on AR glasses, autonomous navigation, delivery robots, and assistive vision systems can use the dataset and the ECN baseline as an in-domain reference point. The finding that state-of-the-art saliency and gaze models fail to generalize to outdoor pedestrian locomotion without fine-tuning is directly relevant to anyone deploying attention-prediction models outside the lab.
Future Directions
- Fine-tuning or redesigning saliency and gaze models specifically for outdoor pedestrian navigation, since the paper shows existing state-of-the-art methods do not generalize well to this setting without in-domain training.
- Reducing reliance on the center bias that dominates egocentric locomotion gaze, which the proposed prior-relative weighting exposes but does not solve.
- Better modeling of the specific targets that attract attention during head turns — structural landmarks and navigational cues — rather than the faces of other pedestrians that generic models favor.
- Extending the human-robot comparison by leveraging the companion YOPO-Campus robot-view dataset over the same 6 km of campus paths.
- The paper does not report cross-subject generalization experiments, subject-demographic breakdowns, or performance per path, leaving those as open empirical questions.
Target Audience
Researchers and graduate students in computer vision, egocentric video, saliency and gaze prediction, and human-robot interaction who need an outdoor pedestrian gaze benchmark. It is also useful for practitioners building AR glasses or navigation and robotics systems that must anticipate human visual attention, and for anyone studying the limitations of current saliency models when applied outside indoor, task-oriented settings.
Authors’ abstract
We address the challenge of predicting human visual attention during real-world navigation by measuring and modeling egocentric pedestrian eye gaze in an outdoor campus setting. We introduce the EgoCampus dataset, which spans 25 unique outdoor paths over 6 km across a university campus with recordings from more than 80 distinct human pedestrians, resulting in a diverse set of gaze-annotated videos. The system used for collection, Meta's Project Aria glasses, integrates eye tracking, front-facing RGB cameras, inertial sensors, and GPS to provide rich data from the human perspective. Unlike many prior egocentric datasets that focus on indoor tasks or exclude eye gaze information, our work emphasizes visual attention while subjects walk in outdoor campus paths. Using this data, we develop EgoCampusNet, a novel method to predict eye gaze of navigating pedestrians as they move through outdoor environments. Our contributions provide both a new resource for studying real-world attention and a resource for future work in gaze prediction models for navigation. Dataset and code will be made publicly available at a later date at https://github.com/ComputerVisionRutgers/EgoCampus .