Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized representations, whose gains concentrate on frame-to-frame change while retaining useful appearance. We address this with a matched 4×6=24 architecture-objective study at roughly 170M ~ 190M encoder scale on ∼1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. TT-VidT combines a DINOv3-initialized ViT-B/16 per-frame spatial path with a compact Temporal Transfer Layer, trained by Diff Compression to reconstruct target frames from a first-frame appearance anchor and frame-specific motion tokens. The sweep shows that TT3D with Diff Compression, not either component alone, enters the strongest motion-sensitive regime, and decoder ablations favor a compact video-pretrained decoder. In final comparison, TT-VidT leads Jester, Something-Something V2, ARID, and Diving48 fine-tuning simultaneously, improving over the strongest non-TT row by 54% ~ 121%, while using 48% fewer encoder FLOPs than DisMo and 55% fewer than VideoMAE or V-JEPA2. HMDB51, IARD, and EPIC-Kitchens bound the claim.
Figures & tables
Figure 1: Overview of TT-VidT’s decoupled pretraining design. A wide first-frame spatial feature supplies the appearance anchor, while a compact temporal transfer path emits frame-specific motion tokens for Diff Compression.
Figure 2: TT-VidT pretraining overview. The encoder interleaves DINOv3 2D ViT processing with Temporal Transfer. The Diff Compression decoder performs diffusion-style denoising conditioned on the first-frame anchor and motion tokens, with loss computed against the target frame.
Objective
Dataset
Arch
MAE
AdaAR
tjAR
AR
MAE-Diff
DiffComp
Jester
ViT3D
39.47
12.85
13.01
12.85
31.68
19.82
DisMo
29.71
16.46 / 46.95
15.94
13.54
21.66
12.98
TT1D
26.65
11.70
10.45
11.86
15.68
24.32
TT3D
11.07
12.85
10.42
11.83
10.32
53.89
SSv2
ViT3D
14.95
3.99
4.86
3.93
11.73
6.57
Table 1: Pretraining sweep: top-1 classification accuracy (%) via frozen attentive probing on Jester, Something-Something V2, and ARID. Each cell is one architecture-objective configuration under the matched recipe (§ 4.1 ) with a common decS_imgnet decoder, single run. Bold = best in block; underline = second-best. Blue = VideoMAE, green = DisMo (second number with dual augmentation), orange = ours.
Model
FLOPs (GF)
TT-VidT TT1D
400.5
TT-VidT TT3D
456.1
DisMo
874.7
V-JEPA 2
1009.9
VideoMAE
1012.6
Table 3: FLOPs at 2562 For each model
Frozen
FT
Jester
ViT3D+MAE
39.47
51.50
DisMo dual-aug
46.95
51.76
TT-VidT
73.25
72.79
SSv2
ViT3D+MAE
14.95
13.40
Table 6: Frozen probe vs. 30-epoch end-to-end finetuning.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Each point is a dataset. X-axis: DINOv3 ViT-B single-frame attentive accuracy (appearance baseline). Y-axis: best motion-model attentive across our entries. Region above the diagonal: motion features add value over appearance.
Method
k NN@20
DINOv3 1f (appearance reference)
100.00
V-JEPA 2, 16f
98.79
DisMo with dual augmentation
99.56
VideoMAE
98.90
TT-VidT
70.99
random baseline
20.00
Appendix
Table 7: IARD identity classification (5 actors, random split). k NN@20 over frozen attentive features. Lower is better: a motion-prioritized representation should retain less per-frame identity. Random baseline is 20.00% .
Figure 4: The canonical configurations under the six frozen probes, from the weakest ( k NN) to the strongest (attentive) readout.
Figure 5: Attentive probe (finetune for Diving48) of the 24 architecture-objective cells on every task. Colour is scaled per panel, the best cell is bold.
Figure 6: Decoder size and initialization on TT3D + Diff Compression, attentive probe (finetune for Diving48). Crosses are size-S decoders from random initialization.
Figure 7: From left: the motion-inversion probe on SSv2 (flip above the axis, stay below), frozen probe against end-to-end finetuning on Jester (dashed: DisMo without augmentation), the canonical cells over three pretraining and three probe seeds, and the cumulative SSv2 gain over a 15-epoch continuation (the dotted line marks the 8-epoch budget).