Skip to content
AI.info

Research

Argos: Adapt Rich Geometric Priors for Generalizable Online Scene-Change-Detection

Overview Research area: Computer vision for robotics — scene change detection, 3D reconstruction, and long-term SLAM/mapping. Technical level: Advanced. The paper assumes familiarity with transformer

Argos: Adapt Rich Geometric Priors for Generalizable Online Scene-Change-Detection
arXiv
2610.10181
Published
2026-10-07
Authors
Ruihan Xu, Jiae Yoon, Kaichen Zhou, Ue-Hwan Kim, Luca Carlone

AI summary

Overview

Research area: Computer vision for robotics — scene change detection, 3D reconstruction, and long-term SLAM/mapping.

Technical level: Advanced. The paper assumes familiarity with transformer architectures, cross-attention, foundation models, camera pose/depth estimation, and SLAM systems.

Scope: The paper introduces Argos, a feedforward model that adapts frozen Geometric Foundation Model features for joint scene change detection and 3D reconstruction, plus a large-scale benchmark (Argos-CD) and an online robotics system (Argos-SLAM) for change-aware 4D mapping.

What This Paper Is About

Robots that revisit a place need to know what changed — a moved tool, a new obstacle, a removed chair — and they need to keep an accurate map of the current state of the world. Existing learning-based change detectors mostly compare 2D image features, which break down under large viewpoint changes and occlusions, are sensitive to noise, and transfer poorly across domains, while explicit 3D methods usually require costly offline optimization. Argos instead reuses the implicit 3D knowledge inside a Geometric Foundation Model to compare two unposed RGB image collections from two visits, jointly predicting change masks, camera poses, and depth.

Key Contributions

  1. Geometrically grounded change detection. Argos adapts implicit 3D priors from a frozen VGGT-Ω backbone through a Co-visibility-aware Spatio-Temporal Comparator (Camera Reader, Inter-session Cross-Attention, DPT Fusion Head) to jointly detect changes and reconstruct scene geometry from a pair of RGB sequences.
  2. Large-scale dataset and generalization. The authors introduce the Argos-CD benchmark, comprising two synthetic datasets (Argos-CD-HSSD and Argos-CD-SceneSmith) and one real-world evaluation set (Argos-CD-real), and show that pretrained GFM features plus multi-dataset training enable strong benchmark performance and synthetic-to-real transfer.
  3. Online change-aware mapping. Argos-SLAM extends VGGT-SLAM 2.0 so that the same Argos network serves as both geometry backbone and change model, producing online change-aware point-cloud maps with asynchronous change-detection jobs.

Main Findings

  • Large gains on video scene change detection. On Argos-CD-SceneSmith, Argos reports increases of 42.01% in change IoU and 27.91% in F1 over baselines. On Argos-CD-HSSD it reaches 78.64 change IoU, 88.51 mIoU, and 62.16 F1; on Argos-CD-SceneSmith 82.43 change IoU, 90.75 mIoU, 69.08 F1; on VSCD 68.69 change IoU, 83.46 mIoU, 53.14 F1.
  • Strong real-world generalization without real fine-tuning. Ours-synthetic-only reaches 39.84 change IoU on the SceneDiff Dataset versus 20.70 for the best fine-tuned learned baseline (VSCDNet-finetune), and 20.87 versus 11.72 on Argos-CD-real, also surpassing the zero-shot SceneDiff baseline.
  • Fine-tuning and foundation training help further. Ours-finetune reaches 59.91 change IoU and 44.73 F1 on the SceneDiff Dataset and 36.29 change IoU and 39.38 F1 on Argos-CD-real. On the SceneDiff Dataset, Ours-foundation scores 62.54 change IoU (80.66 mIoU, 45.98 F1), higher than Ours-finetune; on Argos-CD-real, Ours-foundation scores 29.50 change IoU versus 36.29 for Ours-finetune.
  • Architecture matters more than training data alone. Training the closest baseline, VSCDNet, with the same strategy as Ours-foundation still yields 19.27 change IoU on the SceneDiff Dataset versus 62.54 for Ours-foundation.
  • Advantage extends to paired-image change detection. With only one image per session, Argos achieves the highest change IoU and F1 on PSCD (62.83 change IoU, 78.72 mIoU, 69.65 F1), ChangeSim (50.03, 72.52, 61.73), and VL-CMU-CD (77.80, 88.03, 86.3).
  • Ablation confirms complementary components. On VSCD, replacing the VGGT-Ω backbone with SAM features drops change IoU to 27.89 (from 47.97 with the GFM backbone and change head alone). Adding Inter-session Cross-Attention raises VSCD change IoU to 63.51; adding the Camera Reader alone lowers it to 25.96; combining both gives the best result at 68.69 change IoU and 53.14 F1.
  • Online system keeps pace with reconstruction. With inputs capped at 64 images, each change-detection forward pass takes 0.59 s. During online deployment the SLAM loop produces a new submap every 8.2 s on average, and the asynchronous change-detection job completes in 6.9 s.
  • Dataset scale. Argos-CD-HSSD has 501 pairs, 644K views, and 167 scenes; Argos-CD-SceneSmith has 612 pairs, 138K views, and 204 scenes; Argos-CD-real has 8 pairs, 5,514 views, and 8 scenes. Argos-CD-HSSD covers 4× more visually distinct places on average and up to 2× longer maximum video duration than VSCD and SceneDiff.

Methodology in Plain English

The authors start from the observation that a Geometric Foundation Model trained to match the same 3D location across views already encodes a useful prior for spotting content that does not match across visits. Probing VGGT-Ω attention, they find that tokens respond strongly to corresponding object locations across viewpoints within a visit but weakly when the object is absent in the other visit.

Argos builds on a frozen VGGT-Ω backbone that processes all images from both visits jointly. Three new trainable modules sit on top. A Camera Reader injects the final camera token into intermediate patch tokens at layers {4, 11, 17, 23}, making patch features aware of the camera. Inter-session Cross-Attention treats each visit's tokens as queries and the other visit's tokens as context, with shared weights in both directions since change is symmetric. A lightweight DPT Fusion Head decodes multi-layer features into per-pixel change logits, thresholded at 0.5, alongside the backbone's existing depth and pose predictions. Unprojecting predicted depth with estimated cameras and masking out changed regions yields a change-aware point cloud per queried visit.

Training uses a weighted binary cross-entropy (positive class up-weighted by 5) applied only to labeled pixels, with the backbone frozen and only the three new modules trained. The change head is warm-started from the pretrained depth head. Batches draw image counts from {2, 8, 16, 32, 48, 64} and aspect ratios from [0.5, 2.0], with paired-image and video streams interleaved by deficit scheduling; training uses AdamW (weight decay 0.05, peak learning rate 10⁻⁵, warm-up over the first 5% of steps, cosine decay, gradient clipping at norm 1.0, bfloat16) for 50 epochs on two RTX 4090 GPUs, taking about three days. Training data spans VL-CMU-CD, PSCD, ChangeSim, VSCD, Argos-CD-HSSD, Argos-CD-SceneSmith, and SceneDiff.

Synthetic data is generated in five steps: scene-state sampling with random object removal, coverage planning with a TSP over viewpoints, rendering at 512×512 and 30 fps, mask derivation from object identities rather than pixel differences, and cross-visit frame matching by scene overlap and pose similarity. Argos-SLAM then uses the same Argos network both to reconstruct submaps and to detect changes asynchronously, retrieving only the most relevant prior keyframes via cached SALAD descriptor matching, strided geometric reprojection filtering, and scene-centroid coverage search.

Why This Matters

Impact on research. The paper argues that implicit geometric representations inside foundation models are a stronger substrate for change detection than handcrafted 2D comparison or explicit offline 3D optimization, and it backs that claim with a new benchmark and cross-domain experiments. It also bridges two previously separate literatures — learned RGB change detection and long-term robot mapping — through Argos-SLAM.

Real-world applications:

  • Long-term autonomous mobile robots that must keep maps current as furniture, tools, and objects move.
  • Mobile manipulators operating in cluttered indoor or workshop spaces, as demonstrated with an Agilex mobile manipulator and Intel RealSense D455.
  • Warehouse or facility monitoring, where detecting additions and removals supports inventory-style audits.
  • Construction or industrial inspection, where a change-aware map distinguishes altered geometry from viewpoint or occlusion artifacts.

Industry relevance. The system runs online: change updates typically finish (6.9 s) before the next submap is produced every 8.2 s, and sub-second inference (0.59 s per forward pass) supports deployment on a desktop GPU. The finding that a synthetic-only model can beat real-data fine-tuned baselines reduces the cost of collecting and annotating change data, which is scarce for long robot traversals.

Future Directions

  • Whether the "foundation model for scene change detection" goal holds under more diverse domains, given that Ours-foundation scores lower than Ours-finetune on Argos-CD-real (29.50 vs 36.29 change IoU) while matching or exceeding it on the SceneDiff Dataset (62.54 vs 59.91).
  • Extending the approach from the two-visit formulation to longer revisits and continuous change streams, building on the asynchronous submap-level updates in Argos-SLAM.
  • Understanding why the Camera Reader helps when combined with Inter-session Cross-Attention but hurts in isolation (25.96 change IoU on VSCD, below the 47.97 of the GFM backbone with change head alone).
  • Applying the implicit geometric prior to other spatial reasoning tasks beyond change detection, as the conclusion explicitly encourages.

Target Audience

Robotics and computer vision researchers working on long-term autonomy, SLAM, and dynamic-environment mapping; practitioners building change detection or map-maintenance systems; and machine learning engineers interested in adapting geometric foundation models to downstream 3D reasoning tasks. Readers without background in multi-view geometry, attention mechanisms, or SLAM will find the architecture and system sections demanding.

Authors’ abstract

Robots operating in dynamic environments require reliable detection of how their surroundings change over time. Existing learning-based methods largely rely on pairwise 2D image features, which struggle under large viewpoint changes and occlusions, are sensitive to noise, and show limited generalization across domains, while explicit 3D approaches typically require costly offline optimization. We show that the implicit 3D knowledge of Geometric Foundation Models (GFMs) provides a strong basis for addressing these limitations. We introduce Argos, which adapts GFM features for joint scene change detection and 3D reconstruction. To address data scarcity and take a step toward a foundation model for scene change detection, we introduce a large-scale benchmark comprising two synthetic datasets and one real-world dataset, and train jointly across diverse datasets to improve cross-domain generalization. We further introduce Argos-SLAM, a real-time system designed for robotics, which performs online change detection and change-aware 4D mapping. Across benchmarks, our framework substantially outperforms existing baselines, with gains of up to 42.01% in change IoU and 27.91% in F1, while supporting scalable deployment in changing real-world environments.

Read the original paper