Skip to content
AI.info

Research

RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo

Overview Research area: Computer vision — test-time adaptation for deep stereo depth estimation, specifically continual test-time adaptation (CTTA) under evolving domain shifts. Technical level: Inter

RobIA: Robust Instance-aware Continual Test-time Adaptation for Deep Stereo
arXiv
2511.10107
Published
2025-11-13
Authors
Jueun Ko, Hyewon Park, Hyesong Choi, Dongbo Min

AI summary

Overview

Research area: Computer vision — test-time adaptation for deep stereo depth estimation, specifically continual test-time adaptation (CTTA) under evolving domain shifts.

Technical level: Intermediate to Advanced. The paper assumes familiarity with stereo disparity estimation, parameter-efficient fine-tuning (PEFT), mixture-of-experts routing, and the stability/plasticity trade-off in continual adaptation.

Scope: The paper proposes RobIA, a CTTA framework for stereo depth estimation that combines an instance-aware, PEFT-based mixture-of-experts module (AttEx-MoE) with a batch-normalization-based teacher that supplies dense pseudo-labels, and evaluates it on DrivingStereo, DSEC, and KITTI RAW over 10 adaptation rounds.

What This Paper Is About

Deep stereo networks are usually trained on synthetic data and degrade when deployed in the real world, where conditions such as weather, lighting, and scene type keep changing. Existing test-time adaptation methods for stereo assume a single, static target domain and apply the same adaptation to every input, so they struggle when the domain shifts repeatedly and when the only available supervision is sparse. RobIA targets this harder setting, adapting continuously and per-instance while keeping the pre-trained backbone frozen.

Key Contributions

  1. A CTTA framework for stereo (RobIA). The authors formulate stereo depth estimation as a continual test-time adaptation problem under dynamic domain shifts, rather than the single-domain or long-static-sequence settings used in prior stereo TTA work.

  2. Attend-and-Excite Mixture-of-Experts (AttEx-MoE). A parameter-efficient module that routes each input through frozen convolutional channel experts using a lightweight self-attention router constrained to operate row-wise along epipolar lines, plus a Sigmoid-based soft gating function instead of the usual ReLU.

  3. A Robust AdaptBN Teacher. A PEFT-based teacher that adapts only the affine parameters (scale and shift) of batch normalization layers, used to generate predictions in regions where handcrafted pseudo-labels are unavailable.

  4. A dual-source dense pseudo-label scheme. Reliable handcrafted pseudo-labels supervise high-confidence regions while teacher predictions supervise low-confidence regions, combining a proxy loss and a teacher loss weighted by a factor λ.

Main Findings

  • DrivingStereo (10 rounds, mean over all domains). AttEx-MoE + AT reaches 2.77 D1-all / 0.91 EPE, versus 2.98 / 0.97 for AttEx-MoE alone, 3.04 / 0.92 for full tuning, 3.08 / 0.94 for AdaptBN, 3.07 / 0.95 for the EcoTTA MetaNet baseline, and 5.56 / 1.26 for the unadapted CoEx model. MADNet2 without adaptation scores 10.44 / 1.70.

  • DSEC (nighttime, 4 night conditions, 10 rounds). AttEx-MoE + AT reaches 4.46 D1-all / 1.11 EPE on average; AttEx-MoE alone 4.47 / 1.12; AdaptBN 4.54 / 1.13; full tuning 4.59 / 1.13; MetaNet 5.13 / 1.21; unadapted CoEx 8.68 / 1.60. On Night#4 specifically, dense teacher supervision drops D1-all from 4.57 to 4.15 for full tuning and from 4.61 to 4.34 for AttEx-MoE.

  • KITTI RAW (city, residential, campus, road; 10 rounds). AttEx-MoE scores 1.06 D1-all / 0.84 EPE and AttEx-MoE + AT 1.09 / 0.84, beating MADNet2 without adaptation (3.92 / 1.09), full tuning on MADNet2 (1.13 / 0.85), MAD++ (1.24 / 0.88), and the MetaNet baseline (1.33 / 0.88). The lowest mean D1-all on this benchmark is 0.82 for FT + AT on CoEx, with AdaptBN at 0.91 and full tuning at 0.86.

  • Dense supervision helps most in later rounds. The paper reports that models trained with dense pseudo-labels show notably lower error after long-term adaptation than sparse-only counterparts, indicating dual-source supervision limits over-reliance on confident regions.

  • Sparse-only supervision degrades in invalid regions. Using a higher learning rate of 2e-6, the authors show (Figure 2) that models trained only on sparse handcrafted labels steadily improve in valid regions but sharply degrade in invalid regions, even when previously seen domains return. Adding the AdaptBN teacher reduces this degradation.

  • Computational cost. On DrivingStereo, AttEx-MoE uses 1.19M trainable parameters, 1392 MB memory, and 97.34 ms runtime, versus 2.73M / 2744 MB / 255 ms for full tuning and 2.73M / 2835 MB / 267 ms for full tuning with an EMA teacher. AttEx-MoE + AT records the best mean D1-all (2.77) and EPE (0.91) at 1.22M trainable parameters, 4096 MB, and 204 ms. Measurements were taken on an NVIDIA RTX 3090.

  • Architecture ablation. Row-wise self-attention with Sigmoid gating gives the best result (2.98 D1-all / 0.97 EPE), ahead of self-attention + Sigmoid (3.09 / 0.94), shallow embedding + Sigmoid (3.11 / 0.96), GAP + Sigmoid (3.27 / 0.95), and column-wise self-attention + Sigmoid (3.20 / 0.97). ReLU gating underperforms in every router configuration, e.g. row-wise self-attention with ReLU reaches 3.77 / 1.09.

  • Loss weight λ. On DrivingStereo with AttEx-MoE as student, λ = 0.1 performs best overall (2.77 D1-all / 0.91 EPE) compared with 0.05 (2.97 / 0.91), 0.2 (2.84 / 0.90), and 0.3 (2.89 / 0.91). The paper notes larger λ limits early-round adaptability but helps in later rounds as teacher predictions improve.

  • Label density varies sharply by dataset. The effective pseudo-label density is roughly 92% for KITTI RAW, 72% for DrivingStereo, and 45% for DSEC, which the authors use to explain why dense teacher supervision matters most on the sparser benchmarks.

Methodology in Plain English

The authors start from CoEx, a compact real-time stereo network with a MobileNetV2 backbone, retrained on the synthetic source data (Flyingthings3D from SceneFlow) with strong augmentation. Before deployment, they add the AttEx-MoE module to the deepest encoder block at 1/32 resolution and train only that module on labeled source data, keeping the backbone frozen.

AttEx-MoE treats each convolutional output channel as an individual expert. A router decides how strongly to activate each expert for the current input. Instead of the usual shallow embedding or global average pooling, the router uses self-attention, but computes it separately for each row of the feature map — rows correspond to epipolar lines in rectified stereo, which is where matching evidence is expected to lie. Attention outputs along each row are averaged over the height dimension to produce the gating input, which is then passed through a per-layer gating network. The gate uses a Sigmoid rather than a ReLU so that all experts can be softly activated rather than some being zeroed out — the authors argue this matters when the backbone is frozen and learning capacity is limited.

For supervision, ground truth is unavailable at test time. The authors run Semi-Global Matching (SGM) to produce a handcrafted disparity proxy and filter it with a left-right consistency confidence threshold, giving a valid mask and its complement, an invalid mask. In valid regions the student is trained on the proxy labels with a smooth L1 loss. In invalid regions the student is instead trained on predictions from the AdaptBN teacher, which is a model that adapts by updating only the batch normalization affine parameters. The total loss is the proxy loss plus λ times the teacher loss.

At test time, the backbone stays frozen; only the AttEx-MoE router and gating network plus the decoder's regression parameters are updated. The model predicts each frame first, then adapts on that frame before moving to the next.

For evaluation, the authors sample 500 frames per domain from existing TTA sequences to build short sequences with frequent domain shifts: dusky → cloudy → rainy for DrivingStereo, night1 through night4 for DSEC, and city → residential → campus → road for KITTI RAW. Each cycle of 3–4 domains is repeated over 10 rounds. They note this differs from prior work that simulates long-term shifts over roughly 44K frames. Metrics are End-Point Error (EPE) and D1-all, the percentage of pixels whose absolute disparity error exceeds 3 pixels and 5% of the ground truth.

Why This Matters

Impact on research. The paper argues that prior stereo TTA work assumes a single static target domain or long sequences within one domain (typically more than 2K frames per domain), and that standard PEFT modules are input-invariant, applying the same transformation to every input. It also challenges the standard CTTA practice of using fixed or EMA-updated teachers, arguing that in stereo, handcrafted proxy labels already provide stability, so the teacher's job should be generalization into regions the proxy cannot cover. The proposed CTTA benchmark — 500 frames per domain, 3–4 domains per cycle, 10 rounds — is offered as a more demanding alternative.

Real-world applications

  • Autonomous driving, where weather, lighting, and road type change continuously during a single drive and the DSEC benchmark specifically covers nighttime conditions.
  • Robotics and mobile platforms that must estimate depth on the fly without ground-truth labels or the ability to store source data.
  • 3D scene understanding pipelines that depend on dense disparity, where the paper notes handcrafted proxy labels cover only 45% of pixels on DSEC and 72% on DrivingStereo.
  • Deployment on embedded or resource-limited hardware, where the paper's 1.19M trainable parameters and 1392 MB memory for AttEx-MoE are roughly half the trainable parameters and clearly less memory than full-tuning alternatives such as 2.73M / 2744 MB.

Industry relevance. Real-time stereo is already used in automotive and robotics products, and the authors deliberately chose CoEx, described as a compact and real-time stereo network, as the base architecture, with runtime reported in milliseconds. The efficiency framing — parameter-efficient tuning that preserves the pre-trained backbone — is directly relevant to teams who cannot fine-tune large models on device. The code is released at https://github.com/0ju-un/RobIA.

Future Directions

  • Generalizing the adaptive module. AttEx-MoE is applied only to the single deepest encoder block at 1/32 resolution. The paper does not report results for applying it to additional encoder blocks or at other resolutions.

  • Backbone generalization. The paper notes that most MoE research is conducted on Transformer-based architectures and that MoE within CNNs remains relatively unexplored. All reported experiments use CoEx with a MobileNetV2 backbone; whether the design transfers to other stereo architectures is not reported.

  • Scheduling the teacher weight. The λ ablation shows a tension between early rounds, where larger λ limits adaptability, and later rounds, where larger λ helps as the teacher improves. The paper does not report an adaptive or time-varying λ schedule.

  • Extending beyond stereo depth. The paper motivates the work partly by noting that PEFT-based CTTA methods lose effectiveness on dense, spatially structured tasks such as semantic segmentation. Whether the AttEx-MoE and dual-source supervision ideas carry over to other dense prediction tasks is not reported.

  • Additional material. The paper states that supplementary material contains additional TTA results, further CTTA results, and additional pseudo-supervision ablation studies; those results are not included in the content available here.

Target Audience

Researchers and graduate students working on test-time adaptation, continual learning, or self-supervised depth estimation will get the most from this paper, particularly those interested in parameter-efficient adaptation and mixture-of-experts routing for dense prediction. Practitioners building real-time stereo systems for autonomous driving, robotics, or 3D perception will find the resource-efficiency results and the continual-shift benchmark relevant. Readers without a background in stereo geometry, disparity estimation, or PEFT may need to consult the cited prior work (CoEx, DeepMoE, EcoTTA, AdaptBN) first, since the paper assumes that context.

Authors’ abstract

Stereo Depth Estimation in real-world environments poses significant challenges due to dynamic domain shifts, sparse or unreliable supervision, and the high cost of acquiring dense ground-truth labels. While recent Test-Time Adaptation (TTA) methods offer promising solutions, most rely on static target domain assumptions and input-invariant adaptation strategies, limiting their effectiveness under continual shifts. In this paper, we propose RobIA, a novel Robust, Instance-Aware framework for Continual Test-Time Adaptation (CTTA) in stereo depth estimation. RobIA integrates two key components: (1) Attend-and-Excite Mixture-of-Experts (AttEx-MoE), a parameter-efficient module that dynamically routes input to frozen experts via lightweight self-attention mechanism tailored to epipolar geometry, and (2) Robust AdaptBN Teacher, a PEFT-based teacher model that provides dense pseudo-supervision by complementing sparse handcrafted labels. This strategy enables input-specific flexibility, broad supervision coverage, improving generalization under domain shift. Extensive experiments demonstrate that RobIA achieves superior adaptation performance across dynamic target domains while maintaining computational efficiency.

Read the original paper