Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched 4×6=24 architecture-objective study at roughly 170M ~ 190M encoder scale on ∼1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Figures & tables
Figure 1: Overview of TT-VidT’s decoupled pretraining design. A wide first-frame spatial feature supplies the appearance anchor, while a compact temporal transfer path emits frame-specific motion tokens for Diff Compression.
Figure 2: TT-VidT pretraining overview. The encoder interleaves DINOv3 2D ViT processing with Temporal Transfer. The Diff Compression decoder performs diffusion-style denoising conditioned on the first-frame anchor and motion tokens, with loss computed against the target frame.
Objective
Dataset
Arch
MAE
AdaAR
tjAR
AR
MAE-Diff
DiffComp
Jester
ViT3D
39.47
12.85
13.01
12.85
31.68
19.82
DisMo
29.71
16.46 / 46.95
15.94
13.54
21.66
12.98
TT1D
26.65
11.70
10.45
11.86
15.68
24.32
TT3D
11.07
12.85
10.42
11.83
10.32
53.89
SSv2
ViT3D
14.95
3.99
4.86
3.93
11.73
6.57
Table 1: Pretraining sweep: top-1 classification accuracy (%) via frozen attentive probing on Jester, Something-Something V2, and ARID. Each cell is one architecture-objective configuration under the matched recipe (§ 4.1 ) with a common decS_imgnet decoder, single run. Bold = best in block; underline = second-best. Blue = VideoMAE, green = DisMo (second number with dual augmentation), orange = ours.
Model
FLOPs (GF)
TT-VidT TT1D
400.5
TT-VidT TT3D
456.1
DisMo
874.7
V-JEPA 2
1009.9
VideoMAE
1012.6
Table 3: FLOPs at 2562 For each model
Frozen
FT
Jester
ViT3D+MAE
39.47
51.50
DisMo dual-aug
46.95
51.76
TT-VidT
73.25
72.79
SSv2
ViT3D+MAE
14.95
13.40
Table 6: Frozen probe vs. 30-epoch end-to-end finetuning.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Each point is a dataset. X-axis: DINOv3 ViT-B single-frame attentive accuracy (appearance baseline). Y-axis: best motion-model attentive across our entries. Region above the diagonal: motion features add value over appearance.
Method
k NN@20
DINOv3 1f (appearance reference)
100.00
V-JEPA 2, 16f
98.79
DisMo with dual augmentation
99.56
VideoMAE
98.90
TT-VidT
70.99
random baseline
20.00
Appendix
Table 7: IARD identity classification (5 actors, random split). k NN@20 over frozen attentive features. Lower is better: a motion-prioritized representation should retain less per-frame identity. Random baseline is 20.00% .
Figure 4: The canonical configurations under the six frozen probes, from the weakest ( k NN) to the strongest (attentive) readout.
Figure 5: Attentive probe (finetune for Diving48) of the 24 architecture-objective cells on every task. Colour is scaled per panel, the best cell is bold.
Figure 6: Decoder size and initialization on TT3D + Diff Compression, attentive probe (finetune for Diving48). Crosses are size-S decoders from random initialization.
Figure 7: From left: the motion-inversion probe on SSv2 (flip above the axis, stay below), frozen probe against end-to-end finetuning on Jester (dashed: DisMo without augmentation), the canonical cells over three pretraining and three probe seeds, and the cumulative SSv2 gain over a 15-epoch continuation (the dotted line marks the 8-epoch budget).
We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of relying on heavy 3D architectures or reconstruction-based autoencoders for learning with unlabeled data, we repurpose standard image Vision Transformers by representing videos as super images which are grids composed of frames sampled from videos. From each super image, we construct two views: one with spatial patch masking and the other with temporal frame masking, ensuring no information leakage across frames. A shared Vision Transformer (ViT) encoder aligns their embeddings using a masked Siamese loss, capturing both motion and appearance cues without reconstruction. Our decoder-free formulation leverages an image foundation model towards efficient video representation learning. Starting from pretrained DINO-v3 and DeiT-v3 image encoders, VideoMSN achieves state-of-the-art performance on Kinetics-400, UCF101, and HMDB51 while requiring up to 32× fewer and 160× fewer video pretraining epochs compared to prior video self-supervised learning methods. Our proposed approach also shows strong performance in low-shot classification, confirming the transferability of the learned representations in a label-scarce scenario. Project Page: https://cvir.github.io/projects/videomsn.
Owais Iqbal, Sudipta Sarkar, Shyam Marjit +3
Indian Institute of Technology Kharagpur, India · Indian Institute of Science Bangalore, India · École de technologie supérieure Montreal, Canada
Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with weaker inductive biases generally outperform those with stronger assumptions. This is particularly characteristic of the field of Visual Representation Learning, where approaches have gone from being dominated by Supervised Learning, to Weakly Supervised Learning, to the now widespread success of Self-Supervised Learning without human labels. Yet, even modern Self-Supervised Learning approaches still depend on strong inductive biases such as augmentations, masking, or cropping. If this trend holds, even these remaining biases should become bottlenecks at scale -- and our experiments confirm this: the optimal strength of inductive biases decreases as data grows. This motivates the search for approaches that rely on fewer assumptions. To this end, we introduce Temporal Difference in Vision (TDV), a new paradigm for self-supervised learning from video that avoids existing inductive biases, relying instead on a causal assumption that the past causes the future. TDV functions by jointly training an image encoder and a motion encoder so that the current frame's representation plus the encoded motion equals the next frame's representation. Despite not leveraging any strong inductive biases, TDV matches state-of-the-art recipes on dense spatial tasks, laying the foundation for representation learning without strong assumptions.
Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations: first, scaling video models can reach prohibitive costs and second, learning from language restricts the range of concepts that can be learned to those in captions. As a result, video models still struggle with temporal understanding. In this paper we propose a novel approach that uses motion as the central modality for video representation. In particular, given the motion in a video in the form of point-tracks, we use a masked-autoencoder to mask some of the tracks and train the autoencoder to reconstruct the missing tracks. This allows us to learn a representation in a self-supervised manner. We show that using motion to represent videos actually addresses both of the core limitations of video technology. First, it allows us to massively reduce the scale of training data, as motion is inherently appearance-independent and hence needs fewer examples to generalize well. Second, motion allows us to bypass the language-dependent training paradigm, learning better fine-grained concepts. The result is an embedding that we call TIME (Temporally Informed Motion Embedding), a representation trained exclusively on synthetic motion data. We test this embedding on a wide set of tasks in a zero-shot manner. We observe that without bells and whistles, performance is on par with state-of-the-art models using up to 4 orders of magnitude less training data. This is a stepping stone towards a new paradigm of video models that are both more temporally aware as well as more scalable.
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara