I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
Organizations: Faculty of Electrical Engineering and Computing, University of Zagreb · Fundamental AI Lab, University of Technology Nuremberg
Abstract
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretraining. Combined with a comprehensive evaluation suite we find that contrastive and distillation-based methods struggle in this setting, while MAE is more robust but still falls short of standard i.i.d. pretraining. We find that high inter-batch similarity, caused by sliding-window consumption across consecutive batches, does not explain this gap. The main challenge is high intra-batch similarity, where frames within each batch are near-duplicates. To mitigate this, we propose StreamMAE, which preserves the core MAE reconstruction objective while adapting the input pipeline with stream-aware regularization and motion-biased crop selection. StreamMAE outperforms streaming baselines, matches i.i.d. MAE trained on the same video data, remains competitive with ImageNet-pretrained MAE, and scales positively as the pretraining stream grows from 12 to 95 hours.
Figures & tables
| i.i.d. IN-1K | Streaming | ||
| Metric | IN-1K | WT++12h | |
| 0.004 | 0.004 | 0.665 | |
| 0.325 | 0.989 | 1.000 | |
| Backbone | Training regime | IN-1K Acc@1 | CS mIoU | ADE20K mIoU |
| ViT-S | Standard i.i.d. | 77.4 | 64.0 | 26.9 |
| Pre-shuffled stream | 77.4 | 63.6 | 27.2 | |
| ViT-B | Standard i.i.d. | 81.5 | 72.3 | 35.9 |
| Pre-shuffled stream | 81.6 | 73.6 | 36.5 |
| Pretraining Method | Pretraining Data | IN-1K [ 15 ] Acc @ 1 | City [ 12 ] mIoU | ADE [ 57 ] mIoU | NYUv2 [ 46 ] RMSE | KITTI [ 19 ] RMSE |
| Random Init | – | 71.9 | 47.7 | 16.7 | ||
| Standard I.I.D. Setup | ||||||
| MAE – standard IID | WT++12h | 77.0 | 63.5 | 25.9 | ||
| MAE – standard IID | ImageNet-1K | 77.4 | 64.0 | 26.9 | ||
| Streaming Setup | ||||||
| MOCO-v3 [ 10 ] | WT++12h | 68.7 | 52.0 | 19.4 | ||
| Pretraining Method | Encoder | Pretraining Data | IN-1K Acc @ 1 | Cityscapes mIoU | ADE20K mIoU | NYUv2 RMSE | KITTI RMSE |
| Random Init | ViT-B | – | 78.5 | 51.0 | 17.5 | ||
| Standard I.I.D. Setup | |||||||
| MAE – standard IID | ViT-B | WT++12h | 81.2 | 67.8 | 32.9 | ||
| MAE – standard IID | ViT-B | ImageNet-1K | 81.5 | 72.3 | 35.9 | ||
| Streaming Setup | |||||||
| StreamMAE (ours) | ViT-S | WT++12h | 77.5 | 63.8 | 26.1 | ||
| Method | WT duration | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| StreamMAE | WT++95h | 82.0 | 74.0 | 36.5 |
| Avg. {40,60,80,100}% | WT++95h | 81.9 | 75.5 | 37.0 |
| Avg. {60,100}% | WT++95h | 81.9 | 75.3 | 37.3 |
| Method | ViT-S/16 | ViT-B/16 | ||
| Acc@1 | Acc@5 | Acc@1 | Acc@5 | |
| MAE – standard i.i.d. | 40.0 | 63.0 | 46.4 | 69.0 |
| MAE (streaming) | 33.4 | 55.8 | 37.3 | 59.6 |
| StreamMAE | 40.2 | 63.6 | 47.1 | 69.9 |
| Pretraining Data | Method | CS (mIoU) | ADE (mIoU) | NYUv2 (RMSE) | KITTI (RMSE) |
| HD-EPIC [ 38 ] | MAE – standard i.i.d. | 65.2 | 30.7 | ||
| MAE (streaming) | 63.2 | 29.3 | |||
| StreamMAE (ours) | 65.6 | 32.0 | |||
| CROWD [ 3 ] | MAE – standard i.i.d. | 68.1 | 32.3 | ||
| MAE (streaming) | 65.7 | 29.1 | |||
| StreamMAE (ours) | 71.5 | 32.6 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Training regime | ViT-S/16 | ViT-B/16 | ||||
| ADE20K (mIoU) | Cityscapes (mIoU) | ADE20K (mIoU) | Cityscapes (mIoU) | |||
| i.i.d. | 0.171 | 0.738 | 25.9 | 63.5 | 32.9 | 67.8 |
| Pre-shuffled stream | 0.171 | 0.999 | 26.0 | 63.9 | 32.3 | 68.8 |
| Chronological stream | 0.665 | 1.000 | 23.9 | 61.3 | 26.2 | 58.2 |
| Method | ViT-S/16 | ViT-B/16 | ||||
| IN-1K Acc@1 | ADE20K mIoU | Cityscapes mIoU | IN-1K Acc@1 | ADE20K mIoU | Cityscapes mIoU | |
| MAE – standard i.i.d. | ||||||
| MAE (streaming) | ||||||
| StreamMAE | ||||||
| Method | Batch size | DataDrop | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| StreamMAE (ViT-B) | ✗ | 81.7 | 73.6 | 36.4 | |
| ✗ | 81.8 | 73.3 | 35.4 | ||
| ✗ | 81.4 | 72.2 | 34.2 |
| Method | Batch size | DataDrop (75%) | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| MOCOv3 [ 10 ] baseline (ViT-S) | ✗ | 66.3 | 51.7 | 18.6 | |
| ✓ | 68.7 | 52.0 | 19.4 | ||
| MAE [ 26 ] baseline w/ Orthogonal-AdamW [ 22 ] (ViT-S) | ✗ | 75.2 | 55.3 | 20.6 | |
| ✗ | 75.5 | 58.6 | 21.5 | ||
| ✓ | 75.9 | 58.3 | 22.5 | ||
| MAE [ 26 ] baseline (ViT-S) | ✗ | 76.6 | 57.9 | 23.2 |
| Method | WT duration | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| StreamMAE (ViT-B) | WT++12h +Budapest | 81.1 | 69.4 | 32.5 |
| WT++25h +Budapest | 81.5 | 72.2 | 35.2 | |
| WT++50h | 81.9 | 73.8 | 35.6 |
| Method | Masking Strategy | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| StreamMAE (ViT-B) | Block Masking [ 55 ] (75%) | 80.9 | 68.1 | 32.8 |
| I-JEPA [ 4 ] / Bootleg [ 33 ] | 80.4 | 66.9 | 31.7 | |
| Random Uniform [ 26 ] (75%) | 81.1 | 69.0 | 32.7 |
| Method | AdamW | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| StreamMAE (ViT-B) | 81.0 | 67.1 | 31.8 | |
| 81.1 | 69.0 | 32.7 |
| Components | Downstream performance | |||||
| DataDrop | Color Jit. & Drop Path | Two-stage Cropping | Motion-biased Crop Selection | Cityscapes | ADE20K | KITTI |
| – | – | – | – | 60.8 | 23.2 | |
| ✓ | – | – | – | 61.3 | 23.9 | |
| ✓ | ✓ | – | – | 62.6 | 25.1 | |
| ✓ | ✓ | ✓ | – | 62.8 | 26.1 | |
| ✓ | – | ✓ | ✓ | 62.4 | 24.7 | |
| Method | ADE (mIoU) | CS (mIoU) | NYUv2 (RMSE) | KITTI (RMSE) |
| StreamMAE w/o motion-biased selection | 33.9 | 71.5 | ||
| StreamMAE (ours) | 34.5 | 72.2 |
| Method | Pretraining data | IN-1K Acc@1 | CS (mIoU) | ADE (mIoU) |
| Standard i.i.d. MAE (ViT-S) | ImageNet-1K | 78.5 | 65.9 | 28.5 |
| StreamMAE (ViT-S) | WT++95h | 78.4 | 68.0 | 29.5 |
| Standard i.i.d. MAE (ViT-B) | ImageNet-1K | 82.5 | 75.2 | 39.8 |
| StreamMAE (ViT-B) | WT++95h | 82.0 | 74.0 | 36.5 |
| StreamMAE (ViT-B, avg. {60,100}%; see Table 7 ) | WT++95h | 81.9 | 75.3 | 37.3 |
| Stream | # Videos | Duration | |
| WT++12h | 1 | 12.0h | London stream |
| WT++25h | 10 | 24.8h | WT++12h + 9 videos from original WalkingTours [ 49 ] |
| WT++50h | 26 | 50.0h | WT++25h + 16 appended videos |
| WT++95h | 58 | 94.5h | Full WT++ collection |
| Stream stage | # New videos | Cities / videos in training order |
| WT++12h | 1 | London |
| WT++25h | 9 | Venice, Amsterdam, Singapore, Istanbul, Bangkok, Stockholm, Kuala Lumpur, Zurich, Chiang Mai |
| WT++50h | 16 | Frankfurt, Phnom Penh, Sevilla, Hoi An, Prague, Chemnitz, Marmaris, Dubrovnik, Mostar, Gibraltar, Barcelona, Valencia, Rhodes, Helsinki, Copenhagen, Budapest |
| WT++95h | 32 | Vilnius, Danang, Klaipeda, Liege, Marbella, Monschau, Naples, Nessebar, Port de Soller, Sarajevo, Arcadia, Timisoara, Warsaw, Malacca, Florence, Cardiff, Georgetown, Luxembourg, Alicante, Ipoh, Belfast, Kaliningrad, Maastricht, Valletta, Vienna, Saigon, Cadiz, Kyiv, Oxford, Sihanoukville, Chongqing, Tokyo |
| Experiment | Backbone | GPUs | Wall-clock time |
| StreamMAE pretraining on WT++12h | ViT-S/16 | H100 | 7h |
| StreamMAE pretraining on WT++12h | ViT-B/16 | H100 | 8h |
| MoCo v3 pretraining on WT++12h | ViT-S/16 | H100 | 18h |
| DINO pretraining on WT++12h | ViT-S/16 | H100 | 25h |
| ImageNet-1K fine-tuning | ViT-S/16 | H100 | 4h |
| ImageNet-1K fine-tuning | ViT-B/16 | H100 | 5h |