Skip to content
AI.info

Research

TAPVid-360: Tracking Any Point in 360 from Narrow Field of View Video

TAPVid-360: Tracking Any Point in 360 from Narrow Field of View Video Overview Research area: Computer vision — long-term point tracking (Track Any Point / TAP), panoramic (360°) video, and allocentri

TAPVid-360: Tracking Any Point in 360 from Narrow Field of View Video
arXiv
2511.21946
Published
2025-11-26
Authors
Finlay G. C. Hudson, James A. D. Gardner, William A. P. Smith

AI summary

TAPVid-360: Tracking Any Point in 360 from Narrow Field of View Video

Overview

Research area: Computer vision — long-term point tracking (Track Any Point / TAP), panoramic (360°) video, and allocentric scene representation.

Technical level: Advanced. The paper assumes familiarity with the TAP task, camera intrinsics/extrinsics, SO(3) rotations, and existing trackers such as CoTracker3, TAPIR, and SpatialTracker.

Scope: The paper defines a new task (TAPVid-360), constructs a 10k-sample dataset and benchmark from 360° video without 4D ground truth, and proposes a fine-tuned CoTracker3 baseline called CoTracker360.

What This Paper Is About

Existing Track Any Point methods predict 2D pixel tracks and effectively stop working when a point leaves the camera's field of view, turning re-entry into a re-identification problem. TAPVid-360 instead asks a model to predict the 3D direction (a unit vector in the camera coordinate frame) to each queried scene point in every frame, including points far outside the visible narrow field of view. The authors show this supervision can be generated by resampling 360° videos into perspective clips while tracking points across the full panorama with a 2D pipeline, avoiding the need for dynamic 4D ground-truth scene models.

Key Contributions

  1. A new task, TAPVid-360: given query points as pixel coordinates in the first frame of a narrow field-of-view video, predict unit-vector directions in the camera coordinate frame for all points across the sequence, even when points have left the field of view.
  2. A scalable data-generation pipeline that needs no 3D ground truth: starting from 360° video, the authors segment dynamic objects with Lang-SAM and SAM2, render object-centred perspective clips, track points with CoTracker3, convert those 2D tracks to direction vectors using camera intrinsics, then resample new perspective camera trajectories.
  3. A new dataset and benchmark, TAPVid360-10k: 10k perspective videos with ground-truth directional point tracking, generated from around 130k filtered 10-second clips, plus a comparison against existing 2D and 3D TAP datasets.
  4. A baseline, CoTracker360: CoTracker3 modified to predict per-point, per-frame rotation matrices (3×3, projected onto SO(3)) that update a point's direction from frame to frame, fine-tuned on 5k additional perspective clips.

Main Findings

  • Out-of-frame accuracy: CoTracker360 achieves the best <δ^x_avg out-of-frame score at 0.1160 ± 0.1094, which the paper states is 1.3x higher than the next-best baseline, TAPIP3D (0.0850 ± 0.1227).
  • Out-of-frame angular error: CoTracker360 attains an AD^x_avg out-of-frame value of 10.9829 ± 9.7290, which the paper describes as a greater than 4-fold reduction over TAPIP3D (45.7951 ± 31.5080).
  • In-frame precision trade-off: CoTracker3 (offline), SpatialTracker, and TAPIP3D show high in-frame <δ^x_avg (0.5588 ± 0.1574, 0.4893 ± 0.1946, and 0.4698 ± 0.2391 respectively) but large angular error magnitudes; CoTracker360 trades some in-frame thresholded accuracy (0.4060 ± 0.1355) for much lower overall angular error (8.2749 ± 7.4466 across all points).
  • In-frame accuracy is preserved on angular distance: CoTracker360's in-frame AD^x_avg of 3.9496 ± 4.4395 is far lower than CoTracker3's 17.6352 ± 21.3299 and SpatialTracker's 22.1635 ± 21.4754.
  • 2D trackers fail out of frame: TAPNext (0.0004 ± 0.0009), TAPIR (0.0003 ± 0.0006), and BootsTAPIR (0.0005 ± 0.0009) produce near-zero out-of-frame <δ^x_avg.
  • Dataset breadth: TAPVid360-10k contains 36.28M points within frame and 45.64M points out of frame. Its largest category is Travel and Events at 27.8%.
  • Filtering scale: From the 360-1M dataset of approximately 1 million YouTube links, keeping only videos with greater than 15 likes leaves around 100k videos, which after coarse and fine filtering yields around 130k valid 10-second clips.
  • Training data split: The authors generate 5k training and 10k test samples, with the training clips kept distinct from TAPVid360-10k.

Methodology in Plain English

The authors begin with a large collection of YouTube 360° videos in equirectangular format and clean it aggressively. They discard videos without 360° metadata, videos that are really flat perspective footage or posters projected onto a sphere, static videos, and videos with visible seams or watermarks. They cut the survivors into 10-second clips.

For each clip, they take 32 frames. On the first frame they run Lang-SAM with a list of typically dynamic object classes (such as person, dog, car) to find moving things, then use SAM2 to carry those masks through the whole clip. Any object whose mask cannot be maintained is dropped. For each surviving object, they render a virtual perspective camera that keeps the object centred, producing a narrow field-of-view clip and a matching sequence of object masks. Query pixels sampled inside the first-frame mask are fed to CoTracker3, which produces 2D tracks; tracks with too little cumulative motion are discarded as static.

Those 2D tracks are converted into unit direction vectors in camera coordinates by multiplying homogeneous pixel coordinates by the inverse of the intrinsic matrix and normalising. Finally, the authors resample a new perspective camera trajectory from the same 360° video — using motion strategies named static, spin along x/y/z, spiral, simulated human, random, and back-to-front. Because the ground-truth directions come from tracking in the full panorama, they remain valid even when the object is behind the resampled camera.

The baseline, CoTracker360, keeps CoTracker3's architecture but replaces the last decoder layer with a linear layer producing 9 outputs, reshaped into a 3×3 matrix and orthonormalised to the nearest rotation using special orthogonal Procrustes. Starting from the pretrained CoTracker3 offline weights, the model applies these rotations to each point's initial direction, supervised with Huber loss. Training uses 120 epochs, the Adam optimizer, a learning rate of 1e-4, a single NVIDIA A40 GPU, batch size 8, and 32 query points across 32 frames.

Evaluation uses 256 query points and 32-frame clips. Accuracy is measured with the standard TAP <δ^x_avg metric, but thresholds are angular rather than pixel-based: because the dataset's field of view gives 0.2755° per pixel, thresholds are set to 1, 2, 4, 8, and 16 px°. Mean angular distance (AD^x_avg) is also reported, and both metrics are split into in-frame and out-of-frame subsets.

Why This Matters

Impact on research: The paper shows a way to obtain allocentric, panorama-persistent supervision at scale without dynamic 4D scene models — a major obstacle for 3D tracking benchmarks such as TAPVid-3D, whose 2.5D representation still breaks when points leave the view frustum. It also makes out-of-frame tracking an explicit, separately measured evaluation axis rather than something padded away.

Real-world applications the paper identifies:

  • Robotics and active vision: a robot with a persistent panoramic model can point its camera in the predicted direction to reacquire an object that left its field of view.
  • Re-identification priors: predicted directions can reject candidate reappearances that would require implausible motion, improving re-ID robustness when objects return to view.
  • Video generation conditioning: directional tracks can be used to enforce temporal consistency and plausible object motion, reducing the tendency of generative models to forget or hallucinate objects after they leave and re-enter view.
  • Pretraining for 3D tasks: because 4D ground truth is hard to acquire, the authors propose using this pipeline to pretrain models before fine-tuning on smaller 3D datasets.

Industry relevance: The work is directly relevant to robotics, AR/VR, and video synthesis pipelines. The authors also state that the same advances could be misused for more sophisticated surveillance that erodes privacy, or for autonomous weaponry with reduced human oversight.

Future Directions

  • Handling zoom and varying fields of view. The dataset uses a fixed field of view, so it cannot test sensitivity to field of view or dynamically changing field of view (zoom). The authors note this also means the baseline can use a fixed positional encoding per patch.
  • Per-patch directional encoding. To support zoom correctly, the model would need a directional encoding based on the direction through each patch centre rather than a fixed positional encoding.
  • Uncertainty modelling. The baseline would benefit from representing a directional distribution, allowing uncertainty to grow as points leave the field of view.
  • Scaling and transfer. The pipeline is described as capable of producing significantly larger-scale datasets with similar characteristics, and the authors argue directional tracking can serve as pretraining for 3D-related downstream tasks.

Target Audience

Researchers and engineers working on point tracking, 3D/scene reconstruction, 360° video, and embodied or robotic perception — particularly those who need long-horizon tracking that survives occlusion and view-frustum exit. It also suits practitioners building video generation or re-identification systems who want geometric motion priors, and benchmark designers interested in evaluation protocols that score out-of-frame behaviour explicitly.

Authors’ abstract

Humans excel at constructing panoramic mental models of their surroundings, maintaining object permanence and inferring scene structure beyond visible regions. In contrast, current artificial vision systems struggle with persistent, panoramic understanding, often processing scenes egocentrically on a frame-by-frame basis. This limitation is pronounced in the Track Any Point (TAP) task, where existing methods fail to track 2D points outside the field of view. To address this, we introduce TAPVid-360, a novel task that requires predicting the 3D direction to queried scene points across a video sequence, even when far outside the narrow field of view of the observed video. This task fosters learning allocentric scene representations without needing dynamic 4D ground truth scene models for training. Instead, we exploit 360 videos as a source of supervision, resampling them into narrow field-of-view perspectives while computing ground truth directions by tracking points across the full panorama using a 2D pipeline. We introduce a new dataset and benchmark, TAPVid360-10k comprising 10k perspective videos with ground truth directional point tracking. Our baseline adapts CoTracker v3 to predict per-point rotations for direction updates, outperforming existing TAP and TAPVid 3D methods. Project page: https://finlay-hudson.github.io/tapvid360

Read the original paper