Research
I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
Overview Research area: Self-supervised visual representation learning from continuous video streams (computer vision / video representation learning). Technical level: Advanced. The paper assumes fam

- arXiv
- 2609.40333
- Published
- 2026-09-30
- Authors
- Ivan Martinović, Lukas Knobel, Yuki M. Asano
AI summary
Overview
Research area: Self-supervised visual representation learning from continuous video streams (computer vision / video representation learning).
Technical level: Advanced. The paper assumes familiarity with masked autoencoders, contrastive and self-distillation objectives, Vision Transformers, and dense-prediction benchmarks, though its core ideas are described in plain terms here.
Scope: The paper benchmarks and then improves self-supervised pretraining when video frames arrive in temporal order as a sliding-window stream, rather than as shuffled i.i.d. batches.
What This Paper Is About
Standard self-supervised learning pipelines shuffle images globally and revisit them across epochs, producing diverse batches. Real visual experience — from a camera, robot, or wearable — instead arrives as an ordered stream in which consecutive frames look nearly identical. The paper asks whether self-supervised learning can work well, and scale, when trained from scratch on continuous video without global shuffling, multi-epoch replay, or a long-term replay buffer, and it identifies high intra-batch similarity (near-duplicate frames inside a batch) as the main obstacle.
Key Contributions
- WT++ dataset. The authors extend the released WalkingTours (WT) dataset, which contains approximately 13 hours of video, into WT++, a 95-hour collection of 58 public walking-tour videos, and define nested streams of roughly 12, 25, 50, and 95 hours (WT++12h, WT++25h, WT++50h, WT++95h) for streaming pretraining.
- A streaming benchmark of SSL objectives. They compare MoCo v3 (contrastive), DINO (self-distillation), and MAE (masked reconstruction), along with streaming-tailored baselines (MAE with Orthogonal-AdamW and a reproduction of MemoryStoryboard), under an identical sliding-window streaming protocol with end-to-end fine-tuning on classification and dense-prediction tasks.
- A diagnostic decomposition of the streaming gap. Using a controlled experiment in which ImageNet-1K is pre-shuffled once and consumed as a fixed stream, they separate inter-batch similarity (overlap between consecutive sliding-window batches) from intra-batch similarity (redundancy within a batch), and show that intra-batch similarity, not inter-batch similarity, explains the degradation.
- StreamMAE. They propose a method that keeps the MAE reconstruction objective unchanged but adapts the input pipeline with stream-aware regularization (color jitter, a higher drop-path rate, and DataDrop) and stream-aware cropping (two-stage cropping plus motion-biased crop selection based on patch-level frame differences).
Main Findings
- Contrastive and distillation methods struggle under streaming. In the ViT-S WT++12h comparison, MoCo v3 reaches 68.7 IN-1K Acc@1, 52.0 Cityscapes mIoU, 19.4 ADE20K mIoU, 0.760 NYUv2 RMSE, and 4.882 KITTI RMSE — falling below random initialization on ImageNet-1K (71.9 Acc@1). DINO reaches 72.7, 53.2, 23.9, 0.734, and 4.454.
- MAE is the strongest streaming baseline but still lags i.i.d. training. Streaming MAE (ViT-S, WT++12h) reaches 77.1 IN-1K Acc@1, 61.3 Cityscapes mIoU, 23.9 ADE20K mIoU, 0.742 NYUv2 RMSE, and 4.211 KITTI RMSE, versus 77.0, 63.5, 25.9, 0.701, and 4.070 for standard i.i.d. MAE on the same WT++12h data. The gap grows with capacity: the paper reports streaming MAE underperforming by nearly 10 mIoU points on Cityscapes for ViT-B/16 (58.2 versus 67.8 for same-data i.i.d. MAE).
- Inter-batch similarity is not the culprit. Batch-similarity statistics show i.i.d. ImageNet-1K at mean intra-batch similarity 0.004 and mean inter-batch similarity 0.325; a pre-shuffled ImageNet-1K stream at 0.004 and 0.989; and the WT++12h stream at 0.665 and 1.000. Training MAE on the pre-shuffled ImageNet-1K stream with stride s = 8 matches standard i.i.d. MAE (ViT-S: 77.4 IN-1K, 63.6 Cityscapes, 27.2 ADE20K; ViT-B: 81.6, 73.6, 36.5), despite near-maximal inter-batch similarity.
- Intra-batch similarity is the main challenge. Because each streaming batch is drawn from a short temporal window, frames within a batch are near-duplicates, and this — not batch overlap — coincides with the degradation relative to i.i.d. pretraining.
- Gradient behavior tracks downstream quality. Measuring cosine similarity between gradients of consecutive batches (last-block MLP parameters, B = 512), vanilla streaming MAE can show weakly negative consecutive-gradient similarity, MAE with Orthogonal-AdamW shows substantially positive similarity, and i.i.d. MAE stays near zero and more stable. Within the MAE family, methods closer to the i.i.d. MAE reference in mean gradient similarity achieve better dense downstream performance.
- StreamMAE closes most of the gap. With ViT-S on WT++12h, StreamMAE reaches 77.5 IN-1K Acc@1, 63.8 Cityscapes mIoU, 26.1 ADE20K mIoU, 0.694 NYUv2 RMSE, and 3.976 KITTI RMSE — comparable to same-data i.i.d. MAE (77.0, 63.5, 25.9, 0.701, 4.070) and better than every streaming baseline, including Orthogonal-MAE (75.9, 58.3, 22.5, 0.749, 4.209) and MemoryStoryboard (71.0, 49.2, 21.6, 0.792, 4.647).
- Scaling the stream helps. With ViT-S, StreamMAE improves from 77.5/63.8/26.1/0.694/3.976 at WT++12h to 78.4/68.0/29.5/0.646/3.765 at WT++95h. With ViT-B it improves from 81.1/69.0/32.7/0.655/3.797 at WT++12h to 82.0/74.0/36.5/0.584/3.496 at WT++95h.
- Scaling capacity helps StreamMAE specifically. At ViT-B on WT++12h, StreamMAE reaches 81.1 IN-1K, 69.0 Cityscapes, 32.7 ADE20K, 0.655 NYUv2 RMSE, and 3.797 KITTI RMSE, versus 80.0, 58.2, 26.2, 0.753, and 4.543 for streaming MAE. StreamMAE on WT++95h exceeds the ImageNet-1K i.i.d. MAE reference on Cityscapes (74.0 versus 72.3) and ADE20K (36.5 versus 35.9), while remaining slightly below on IN-1K (82.0 versus 81.5) and much better on both depth benchmarks (0.584/3.496 versus 0.588/3.771).
- Ablation: crop selection and regularization do most of the work. Cumulative ablations on ViT-S/16 with WT++12h show DataDrop gives modest gains, while stronger regularization and crop selection account for most of the dense-transfer improvement; motion-biased crop selection gives the largest gains on the close-to-domain Cityscapes and KITTI benchmarks.
- DataDrop stops mattering at scale. DataDrop's contribution becomes negligible for ViT-B/16 at WT++50h and WT++95h (e.g., WT++95h: 82.0/74.0/36.5 with DataDrop versus 81.9/74.3/36.3 without).
- Denser frame sampling does not help. For ViT-B on WT++12h, subsampling factor k = 4 (15 FPS) gives 81.0 IN-1K, 65.2 Cityscapes, 30.1 ADE20K; k = 8 (7.5 FPS) gives 81.0, 66.8, 31.4; k = 16 (3.75 FPS) gives 81.1, 69.0, 32.7 — so more updates from higher FPS do not improve transfer.
- New content beats repeated passes. Repeating WT++12h twice gives 81.3/70.3/33.3 and four times gives 81.5/71.7/33.9, while WT++25h gives 81.6/72.5/35.
Authors’ abstract
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.