Skip to content
AI.info

Research

A Study of Finetuning Video Transformers for Multi-view Geometry Tasks

Overview Research area: Computer vision — multi-view geometry (optical flow estimation, stereo matching, unrectified two-view depth estimation) using fine-tuned video foundation models. Technical leve

A Study of Finetuning Video Transformers for Multi-view Geometry Tasks
arXiv
2512.18684
Published
2025-12-21
Authors
Huimin Wu, Kwang-Ting Cheng, Stephen Lin, Zhirong Wu

AI summary

Overview

Research area: Computer vision — multi-view geometry (optical flow estimation, stereo matching, unrectified two-view depth estimation) using fine-tuned video foundation models.

Technical level: Advanced. The paper assumes familiarity with Vision Transformers, self-attention, cost volumes, iterative refinement (RAFT-style), and standard geometry benchmarks.

Scope: A single paper investigating whether general-purpose, video-pretrained vision transformers can be adapted with minimal modification to outperform heavily engineered, task-specific multi-view geometry pipelines.

What This Paper Is About

State-of-the-art multi-view geometry systems such as optical flow estimators use complicated pipelines: convolutional feature stems, transformer encoders that build explicit cost volumes, recurrent refinement networks, and domain-specific pretraining. The authors ask whether a general-purpose video transformer — pretrained only on video data with no geometric or task-specific pretraining — can be repurposed for these geometry tasks by simply changing how the input patches and position encodings are handled, and whether that simple adaptation can match or beat the engineered alternatives. Their answer is yes, in the form of a model called GeoViT.

Key Contributions

  1. A minimal adaptation recipe for 3D video ViTs to two-frame tasks. The pretrained 2D spatial positional encodings are interpolated to the target input size, the temporal positional encodings (which account for 8 frames in pretraining) are split into two halves and averaged to produce source-image and target-image temporal embeddings, and the 3D patch embedding is summed along the temporal dimension to become a 2D patch embedding. The final token is the sum of 2D patch embedding, 2D positional encoding, and temporal encoding.

  2. A demonstration that a plain linear decoder is sufficient. Appending a linear layer to each output patch representation and regressing the geometric quantity directly already yields strong results — an EPE of 2.0 on Sintel (final), surpassing SAMFlow (2.11).

  3. A cost-volume-free iterative refinement decoder. Instead of querying a cost volume, the framework dynamically warps the target image using the previous prediction and re-feeds the image pair through the pretrained encoder, with a ConvGRU predicting the residual, aggregated over iterations and supervised with an exponentially weighted L1 loss (γ = 0.9 by default).

  4. A single unified architecture spanning three geometry tasks — optical flow, stereo matching, and unrectified two-view depth estimation — evaluated with submitted results against state-of-the-art baselines and multiple ablations over pretraining scheme, model size, iteration count, and decoder design.

Main Findings

  • State-of-the-art optical flow results on cross-dataset generalization (Table 1). GeoViT achieves EPE of 0.69 (Sintel clean), 1.78 (Sintel final), and F1-epe / F1-all of 3.15 / 11.45 on KITTI-15 train, trained on Chairs then Things ("C+T"). The prior SAMFlow reports 0.87, 2.11, 3.44 and 12.28 respectively.

  • Reported relative improvements over SAMFlow on the training splits. The paper states a 20.7% and 15.6% error reduction on Sintel clean and Sintel final, and 8.4% (EPE) and 6.8% (F1) on KITTI. The introduction phrases the headline reductions as 21%, 9.6%, and 11.2% for EPE (Sintel clean, Sintel final) and F1 respectively.

  • New records on the online benchmarks. On the Sintel test server, GeoViT reaches EPE of 0.79 (clean) and 1.88 (final), stated as improving on prior state-of-the-art SAM-Flow by 21% on clean and 9.6% on final. On the KITTI benchmark it achieves an F1-all of 3.79, described as an 11.2% error reduction over prior SOTA.

  • Video pretraining matters more than semantics (Table 2a). MAE_st initialization gives the best flow results (0.69 / 1.78 / 3.15 / 11.45), beating MVD (0.84 / 1.91 / 3.61 / 14.28), InternVideo (0.70 / 1.87 / 3.22 / 12.27), UMT (0.78 / 1.89 / 3.73 / 14.13), and the image-only spatial MAE (0.78 / 1.99 / 4.18 / 14.85). Random initialization collapses to 13.49 / 13.49 / 36.74 / 85.35. The authors attribute this to semantics-oriented video objectives not transferring well to low-level representations.

  • Refinement iterations help up to a point (Table 2b). Results improve from linear decoding, through 1 and 2 iterations, to 6 iterations (0.69 / 1.78 / 3.15 / 11.45); 12 iterations is only marginally better (0.70 / 1.78 / 3.13 / 11.29), so 6 is chosen as the default for efficiency.

  • Larger models perform better (Table 2c). ViT-Small gives 0.99 / 2.32 / 6.33 / 19.81, ViT-Base 0.87 / 2.22 / 4.72 / 15.81, and ViT-Large 0.75 / 1.82 / 3.70 / 13.78 — though this table uses reproduced pretraining with fewer steps (800 epochs instead of the default 1600).

  • Image warping beats cost-volume lookup (Table 2d). RAFT with CNN features and queried cost volume gives 1.43 / 2.71 on Sintel clean/final; substituting the video ViT as context encoder improves it to 0.85 / 1.99; GeoViT's warping-based refinement reaches 0.69 / 1.78.

  • Stereo matching on ETH3D (Table 5). GeoViT obtains bad 1.0 / 2.0 / 4.0 of 1.16 / 0.19 / 0.03, best on two of the three metrics. It is comparable to CREStereo on bad 1.0 (0.98 for CREStereo) and beats the prior unified solution GMStereo (1.16 vs. 1.83).

  • Depth estimation on the DeMoN test datasets (Table 6). On RGBD-SLAM GeoViT reports Abs Rel 0.106, Sq Rel 0.204, RMSE 0.508, RMSE log 0.171 — with the largest shortfall versus the best prior result being 0.027 in Sq Rel, and an improvement of 0.048 in RMSE. On SUN3D it is best on 3 of 4 metrics (0.095, 0.068, 0.552, 0.124). On Scenes11 it is best on 2 of 4 metrics (0.118, 0.059, 0.318, 0.146).

  • Qualitative advantages. On Sintel (clean) the authors report more detailed estimates (separating arm and back motion in case #1), higher recall of small objects (all bird motions in case #2), and sharper motion boundaries. A failure case analysis finds the model tends to produce overly smooth depth predictions, which affects complex scenes.

  • Computational cost is manageable. The largest model can be trained on 8 V100 GPUs for at most 2 weeks using techniques such as gradient checkpointing. The paper explicitly notes that controlling factors such as data size and model size for fair comparison is challenging.

Methodology in Plain English

The authors start from a transformer that has already learned from large amounts of video. Such a model splits video into 3D patches (space plus time) and uses self-attention to relate all patches to each other, which the authors argue implicitly encodes correspondence information useful for geometry.

To use it on a pair of images, they make small plumbing changes: rescale the spatial position codes to the new image size, split the temporal position codes in half and average each half so one half represents the source image and the other the target image, and collapse the 3D patch embedding into 2D by summing over time. Everything else in the encoder stays pretrained.

For a first pass, they attach one linear layer to each output patch and directly predict the geometry — surprisingly effective on its own. For the full system, they borrow the "predict a residual, add it back, repeat" idea from RAFT, but throw away RAFT's cost volume. Instead, at each step they warp the target image using the current prediction so the pair of images already accounts for most of the motion; the encoder processes that warped pair, and a ConvGRU decoder outputs a correction Δg to add to the previous estimate. This is trained with an L1 loss over all iterations, weighting later iterations more heavily (γ = 0.9).

The same formulation covers all three tasks: for flow and stereo the warp is straightforward (depth is converted into pixel displacement using camera parameters for the depth task, then converted back). Training follows a standard curriculum — FlyingChairs first (40K steps, batch size 8, 368×496, peak learning rate 1×10⁻⁴), then FlyingThings (400K iterations, batch size 8, 384×768 crops, peak learning rate 2.5×10⁻⁵) — denoted "C+T", with a tiling evaluation scheme using a stride of 224.

Why This Matters

Impact on research. The paper argues that the connection between video understanding and geometric computer vision has been underleveraged. It shows that a general-purpose video model plus a linear or recurrent decoder can beat systems built with custom architectures, cost volumes, and domain-specific pretraining, and that semantics-heavy video pretraining is actually worse for low-level geometry than spatiotemporal masked autoencoding. That reframes what kind of pretraining the geometry community should pursue.

Real-world applications:

  • Optical flow for video editing, frame interpolation, and motion-based visual effects, where dense and accurate per-pixel motion is required.
  • Stereo matching for depth sensing in robotics, autonomous driving, and 3D reconstruction from camera pairs.
  • Unrectified multi-view depth estimation for casual photography and AR/VR, where cameras are not aligned to a common plane and more natural, diverse camera configurations are needed.
  • Action recognition, tracking, video segmentation, and self-supervised learning pipelines that consume flow estimates as auxiliary supervisory signals.

Industry relevance. The approach requires no bespoke data collection or custom pretraining pipeline, which lowers the engineering barrier for companies that want geometry capabilities from an off-the-shelf video backbone. Training cost is bounded to 8 V100 GPUs for at most 2 weeks for the largest model, and the authors provide a project website (geovit-aaai26.github.io), which matters for teams deciding whether to replace hand-built geometry stacks with a finetuned foundation model.

Future Directions

  • Scaling further. Table 2c shows consistent gains from ViT-Small to ViT-Large; the authors note this underscores potential improvement with even larger model sizes.
  • Fixing overly smooth depth. The failure case analysis attributes weaker results on complex scenes like RGBD-SLAM and Scenes11 to smooth depth predictions, raising the question of how to sharpen output on cluttered geometry.
  • More geometry tasks. The paper demonstrates three tasks within one architecture; the same two-frame formulation is not reported as being extended to other geometric quantities.
  • Cleaner controlled comparisons. The authors state directly that controlling data size and model size for fair comparison is challenging, leaving open a more rigorous study of how much of the gain comes from pretraining data, model scale, versus the refinement design.
  • Better pretraining objectives for geometry. Since MAE_st beat semantics-rich models like InternVideo, MVD, and UMT, there is an open question of what video pretraining objective would transfer best to low-level geometric reasoning.

Target Audience

Researchers and engineers working on optical flow, stereo matching, multi-view depth estimation, and 3D reconstruction; practitioners exploring whether video foundation models can replace task-specific architectures; and students with a background in Vision Transformers and iterative refinement methods who want a concrete case study in adapting large pretrained models to dense per-pixel geometric prediction.

Authors’ abstract

This paper presents an investigation of vision transformer learning for multi-view geometry tasks, such as optical flow estimation, by fine-tuning video foundation models. Unlike previous methods that involve custom architectural designs and task-specific pretraining, our research finds that general-purpose models pretrained on videos can be readily transferred to multi-view problems with minimal adaptation. The core insight is that general-purpose attention between patches learns temporal and spatial information for geometric reasoning. We demonstrate that appending a linear decoder to the Transformer backbone produces satisfactory results, and iterative refinement can further elevate performance to stateof-the-art levels. This conceptually simple approach achieves top cross-dataset generalization results for optical flow estimation with end-point error (EPE) of 0.69, 1.78, and 3.15 on the Sintel clean, Sintel final, and KITTI datasets, respectively. Our method additionally establishes a new record on the online test benchmark with EPE values of 0.79, 1.88, and F1 value of 3.79. Applications to 3D depth estimation and stereo matching also show strong performance, illustrating the versatility of video-pretrained models in addressing geometric vision tasks.

Read the original paper