W2Rep: Learning Visual Representations by Watching the World Change
Authors: Wen Huang, Hang Guo, Jiarui Yang, Zheng Liu, Tao Dai, Shu-tao Xia
Organizations: Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · Nankai University, Tianjin, China · Shenzhen University, Shenzhen, China
Images capture the world at one moment, whereas video reveals how it changes. Image self-supervision learns spatial structure from a single moment, while video methods commonly learn temporal relationships inside a representation computed jointly from several frames. We ask whether watching a scene change can instead improve features available from one image without sacrificing the ability to represent video. We introduce W2Rep, a masked feature-prediction framework in which an independently encoded source image participates in prediction at the same or another moment. The predictor is conditioned on visible video context, the queried location, and the signed time interval between source and target. This gives the cross-frame objective two complementary roles: the image path learns features that remain useful across time, while the video path must gather evidence that is missing from the source image. Across model scales and downstream tasks, W2Rep improves frozen and fine-tuned recognition under our comparison protocol, while joint video encoding provides further gains over frame-wise aggregation. Controlled experiments show that these gains depend on directly updating the source-image features and on using both video context and temporal displacement. Overall, change across a video can supervise a visual encoder whose representations remain useful at either image or video granularity. Code is available at~\href{https://wenooi.github.io/W2Rep}{https://wenooi.github.io/W2Rep}.
Figures & tables
Figure 1: Three training interfaces for visual self-supervision. Image methods learn from different views or masked regions of one image. Video methods commonly encode several frames together. W2Rep uses change across a video to train image features. Its video-context path supplies the complementary multi-frame evidence needed for cross-frame prediction, so the same objective also trains the encoder to use video inputs. The retained encoder can later process either an image or a video. The thumbnails show four frames from a single Something-Something V2 video; the illustration is schematic rather than an exhaustive taxonomy.
Figure 2: Overview of W2Rep . The same student visual encoder separately processes a masked source image and a masked video. A predictor combines the resulting source-image features with visible video context, masked spatial queries, and a signed temporal offset to predict features produced by a slowly updated target encoder. After pretraining, only the visual encoder is retained for image or video inference.
Method
ImageNet-1K
ADE20K
ViT-B/16
I-JEPA ( Assran et al., 2023 )
28.50±0.12
20.55
VideoMAE ( Tong et al., 2022 )
26.70±0.10
25.79
V-JEPA ( Bardes et al., 2024 )
24.90±0.09
21.97
TDV ( Daithankar et al., 2026 )
7.29±0.15
14.05
RSP ( Jang et al., 2024 )
24.76±0.07
17.84
Table 1: Frozen transfer at two backbone scales. ImageNet-1K and action columns report top-1 accuracy (%); ADE20K reports mean intersection-over-union (mIoU). Ind. averages eight independently encoded frames, whereas joint applies space–time attention to the same eight frames. Values with ± average three probe seeds. Bold denotes the best result within each backbone and readout. A dash indicates that the readout is not applicable or was not evaluated.
Figure 3: Qualitative cross-frame patch similarity for four SSv2 examples, arranged as (a–b) on the left and (c–d) on the right. Query and target frames are encoded independently, without access to temporal context. The green box marks a 16×16 query patch in the source frame. Each heatmap shows the mean-centered cosine similarity between that query and the final-layer target-frame patch features. Colors are normalized within each map using its 5th and 95th similarity percentiles and therefore indicate spatial structure, not similarity magnitudes across models.
Group
Configuration
IN1K
ADE20K
SSv2 joint-8
UCF101 joint-8
Diving48 joint-8
Reference
Final W2Rep
34.60±0.05
22.21
25.38±0.10
56.60±0.38
11.29±0.25
Objective
Cross-frame only
32.26±0.05
21.61
29.77±0.18
58.15±0.21
11.54±0.16
Same-frame only (bundled)
28.08±0.10
21.89
11.21±0.09
47.83±0.21
7.99±0.21
No temporal offset ( Δ=0 )
30.68±0.06
20.50
15.43±0.11
52.85±0.26
8.80±0.33
Gradient path
Source stop-gradient (cross-frame)
29.75±0.12
19.05
13.99±0.12
50.59±0.10
10.41±0.36
Condition
Zero z(V)
26.82±0.10
20.92
9.33±0.08
44.94±0.12
7.92±0.05
Table 2: Controlled ViT-B/16 ablations using a common pretraining seed (42). Recognition columns report frozen top-1 accuracy (%), averaged over three probe seeds; ADE20K reports frozen-backbone mIoU. IN1K denotes ImageNet-1K. A dash indicates that the transfer task was not evaluated.
Figure 4: Temporal identification from predicted features on 512 SSv2 validation videos. Each row requests one target time; each column compares the prediction with EMA features from one actual frame at the same masked query positions. Cells show cosine similarity relative to the mean of their row, in percentage points, and source-time requests are excluded from the retrieval statistics. Correct temporal conditioning produces a pronounced diagonal and retrieves the requested frame well above the 12.5% random baseline. Perturbing the signed offset, clip order, or clip-dependent latents removes this structure.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Encoding
Top-1
Top-5
Macro Acc.
I-JEPA
Ind.-8
12.23
28.43
8.19
V-JEPA
Joint-8
45.48
70.44
37.25
VideoMAE
Joint-8
55.16
79.13
48.30
W2Rep
Joint-8
58.77
81.84
51.79
Appendix
Table 5: SSv2 full fine-tuning with ViT-B/16. All methods use the same data, augmentations, optimization schedule, and fixed epoch-50 reporting rule. Results are validation accuracy (%) from each method’s designated primary checkpoint. I-JEPA encodes frames separately; all other methods encode them together.
Encoder input
Top-1 (%)
Ordered drop
Ordered
25.38±0.10
–
Reversed
13.51±0.04
11.88
Fixed shuffled
8.98±0.06
16.40
Static repeat
4.24±0.09
21.14
Appendix
Table 8: Final-checkpoint SSv2 order sensitivity. Top-1 values are mean and standard deviation over the same three frozen heads. Drops are paired against ordered input at the video level.
Figure 5: Content and functional diagnostics of z(V) at the final checkpoint. (a) An SSv2 linear head trained on ordered clip latents is evaluated without refitting; accuracy falls when the same frames are reversed or shuffled, and further when one frame is repeated. Error bars show standard deviation over three head seeds. (b) We hold the source, targets, masks, and offsets fixed and change only z(V) . Prediction is strongest with context from the matched video; a donor from another video remains much less useful even when it has the same action label.
Figure 6: Qualitative nearest-neighbor retrieval on EgoDex with frozen W2Rep representations. Each item shows four chronological frames from a real episode. Green borders denote the same directed action as the query, while orange borders denote its inverse. The upper example preserves both action direction and visual context. In the lower example, the near-identical object and recording setup outweigh action direction for both the retained encoder representation and the auxiliary video context z(V) .
Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.
Lukas Knobel, Andrew Zisserman, Yuki M. Asano
Fundamental AI Lab, UTN · VGG, University of Oxford
Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations: first, scaling video models can reach prohibitive costs and second, learning from language restricts the range of concepts that can be learned to those in captions. As a result, video models still struggle with temporal understanding. In this paper we propose a novel approach that uses motion as the central modality for video representation. In particular, given the motion in a video in the form of point-tracks, we use a masked-autoencoder to mask some of the tracks and train the autoencoder to reconstruct the missing tracks. This allows us to learn a representation in a self-supervised manner. We show that using motion to represent videos actually addresses both of the core limitations of video technology. First, it allows us to massively reduce the scale of training data, as motion is inherently appearance-independent and hence needs fewer examples to generalize well. Second, motion allows us to bypass the language-dependent training paradigm, learning better fine-grained concepts. The result is an embedding that we call TIME (Temporally Informed Motion Embedding), a representation trained exclusively on synthetic motion data. We test this embedding on a wide set of tasks in a zero-shot manner. We observe that without bells and whistles, performance is on par with state-of-the-art models using up to 4 orders of magnitude less training data. This is a stepping stone towards a new paradigm of video models that are both more temporally aware as well as more scalable.
Mantas Skackauskas, Xinyue Hao, Laura Sevilla-Lara
Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with weaker inductive biases generally outperform those with stronger assumptions. This is particularly characteristic of the field of Visual Representation Learning, where approaches have gone from being dominated by Supervised Learning, to Weakly Supervised Learning, to the now widespread success of Self-Supervised Learning without human labels. Yet, even modern Self-Supervised Learning approaches still depend on strong inductive biases such as augmentations, masking, or cropping. If this trend holds, even these remaining biases should become bottlenecks at scale -- and our experiments confirm this: the optimal strength of inductive biases decreases as data grows. This motivates the search for approaches that rely on fewer assumptions. To this end, we introduce Temporal Difference in Vision (TDV), a new paradigm for self-supervised learning from video that avoids existing inductive biases, relying instead on a causal assumption that the past causes the future. TDV functions by jointly training an image encoder and a motion encoder so that the current frame's representation plus the encoded motion equals the next frame's representation. Despite not leveraging any strong inductive biases, TDV matches state-of-the-art recipes on dense spatial tasks, laying the foundation for representation learning without strong assumptions.