W2Rep: Learning Visual Representations by Watching the World Change
Authors: Wen Huang, Hang Guo, Jiarui Yang, Zheng Liu, Tao Dai, Shu-tao Xia
Organizations: Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · Nankai University, Tianjin, China · Shenzhen University, Shenzhen, China
Images capture the world at one moment, whereas video reveals how it changes. Image self-supervision learns spatial structure from a single moment, while video methods commonly learn temporal relationships inside a representation computed jointly from several frames. We ask whether watching a scene change can instead improve features available from one image without sacrificing the ability to represent video. We introduce W2Rep, a masked feature-prediction framework in which an independently encoded source image participates in prediction at the same or another moment. The predictor is conditioned on visible video context, the queried location, and the signed time interval between source and target. This gives the cross-frame objective two complementary roles: the image path learns features that remain useful across time, while the video path must gather evidence that is missing from the source image. Across model scales and downstream tasks, W2Rep improves frozen and fine-tuned recognition under our comparison protocol, while joint video encoding provides further gains over frame-wise aggregation. Controlled experiments show that these gains depend on directly updating the source-image features and on using both video context and temporal displacement. Overall, change across a video can supervise a visual encoder whose representations remain useful at either image or video granularity. Code is available at~\href{https://wenooi.github.io/W2Rep}{https://wenooi.github.io/W2Rep}.
Figures & tables
Figure 1: Three training interfaces for visual self-supervision. Image methods learn from different views or masked regions of one image. Video methods commonly encode several frames together. W2Rep uses change across a video to train image features. Its video-context path supplies the complementary multi-frame evidence needed for cross-frame prediction, so the same objective also trains the encoder to use video inputs. The retained encoder can later process either an image or a video. The thumbnails show four frames from a single Something-Something V2 video; the illustration is schematic rather than an exhaustive taxonomy.
Figure 2: Overview of W2Rep . The same student visual encoder separately processes a masked source image and a masked video. A predictor combines the resulting source-image features with visible video context, masked spatial queries, and a signed temporal offset to predict features produced by a slowly updated target encoder. After pretraining, only the visual encoder is retained for image or video inference.
Method
ImageNet-1K
ADE20K
ViT-B/16
I-JEPA ( Assran et al., 2023 )
28.50±0.12
20.55
VideoMAE ( Tong et al., 2022 )
26.70±0.10
25.79
V-JEPA ( Bardes et al., 2024 )
24.90±0.09
21.97
TDV ( Daithankar et al., 2026 )
7.29±0.15
14.05
RSP ( Jang et al., 2024 )
24.76±0.07
17.84
Table 1: Frozen transfer at two backbone scales. ImageNet-1K and action columns report top-1 accuracy (%); ADE20K reports mean intersection-over-union (mIoU). Ind. averages eight independently encoded frames, whereas joint applies space–time attention to the same eight frames. Values with ± average three probe seeds. Bold denotes the best result within each backbone and readout. A dash indicates that the readout is not applicable or was not evaluated.
Figure 3: Qualitative cross-frame patch similarity for four SSv2 examples, arranged as (a–b) on the left and (c–d) on the right. Query and target frames are encoded independently, without access to temporal context. The green box marks a 16×16 query patch in the source frame. Each heatmap shows the mean-centered cosine similarity between that query and the final-layer target-frame patch features. Colors are normalized within each map using its 5th and 95th similarity percentiles and therefore indicate spatial structure, not similarity magnitudes across models.
Group
Configuration
IN1K
ADE20K
SSv2 joint-8
UCF101 joint-8
Diving48 joint-8
Reference
Final W2Rep
34.60±0.05
22.21
25.38±0.10
56.60±0.38
11.29±0.25
Objective
Cross-frame only
32.26±0.05
21.61
29.77±0.18
58.15±0.21
11.54±0.16
Same-frame only (bundled)
28.08±0.10
21.89
11.21±0.09
47.83±0.21
7.99±0.21
No temporal offset ( Δ=0 )
30.68±0.06
20.50
15.43±0.11
52.85±0.26
8.80±0.33
Gradient path
Source stop-gradient (cross-frame)
29.75±0.12
19.05
13.99±0.12
50.59±0.10
10.41±0.36
Condition
Zero z(V)
26.82±0.10
20.92
9.33±0.08
44.94±0.12
7.92±0.05
Table 2: Controlled ViT-B/16 ablations using a common pretraining seed (42). Recognition columns report frozen top-1 accuracy (%), averaged over three probe seeds; ADE20K reports frozen-backbone mIoU. IN1K denotes ImageNet-1K. A dash indicates that the transfer task was not evaluated.
Figure 4: Temporal identification from predicted features on 512 SSv2 validation videos. Each row requests one target time; each column compares the prediction with EMA features from one actual frame at the same masked query positions. Cells show cosine similarity relative to the mean of their row, in percentage points, and source-time requests are excluded from the retrieval statistics. Correct temporal conditioning produces a pronounced diagonal and retrieves the requested frame well above the 12.5% random baseline. Perturbing the signed offset, clip order, or clip-dependent latents removes this structure.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Encoding
Top-1
Top-5
Macro Acc.
I-JEPA
Ind.-8
12.23
28.43
8.19
V-JEPA
Joint-8
45.48
70.44
37.25
VideoMAE
Joint-8
55.16
79.13
48.30
W2Rep
Joint-8
58.77
81.84
51.79
Appendix
Table 5: SSv2 full fine-tuning with ViT-B/16. All methods use the same data, augmentations, optimization schedule, and fixed epoch-50 reporting rule. Results are validation accuracy (%) from each method’s designated primary checkpoint. I-JEPA encodes frames separately; all other methods encode them together.
Encoder input
Top-1 (%)
Ordered drop
Ordered
25.38±0.10
–
Reversed
13.51±0.04
11.88
Fixed shuffled
8.98±0.06
16.40
Static repeat
4.24±0.09
21.14
Appendix
Table 8: Final-checkpoint SSv2 order sensitivity. Top-1 values are mean and standard deviation over the same three frozen heads. Drops are paired against ordered input at the video level.
Figure 5: Content and functional diagnostics of z(V) at the final checkpoint. (a) An SSv2 linear head trained on ordered clip latents is evaluated without refitting; accuracy falls when the same frames are reversed or shuffled, and further when one frame is repeated. Error bars show standard deviation over three head seeds. (b) We hold the source, targets, masks, and offsets fixed and change only z(V) . Prediction is strongest with context from the matched video; a donor from another video remains much less useful even when it has the same action label.
Figure 6: Qualitative nearest-neighbor retrieval on EgoDex with frozen W2Rep representations. Each item shows four chronological frames from a real episode. Green borders denote the same directed action as the query, while orange borders denote its inverse. The upper example preserves both action direction and visual context. In the lower example, the near-identical object and recording setup outweigh action direction for both the retained encoder representation and the auxiliary video context z(V) .