Earth observation (EO) data provide rich temporal supervision, yet existing remote sensing foundation models mainly exploit sequential observations through imposing predefined pairwise relations or aggregating holistic reconstruction context. We seek to further exploit the sparse and nonuniform temporal sampling inherent in EO sequences as supervisory signals. To this end, we propose T-JEPA, a temporal joint-embedding predictive architecture that learns time-gap-conditioned latent transitions. A shared single-frame encoder processes each observation, while a temporal predictor estimates the complete target latent field from a masked source latent representation and the actual elapsed time. Across multiple temporal intervals, these predictive constraints organize observed states into structured latent trajectories. Asymmetric metadata injection mitigates shortcut learning, and direct supervision across multiple temporal scales proves more effective than recursively rolling out intermediate states. In parallel, masked pixel reconstruction provides complementary supervision for preserving spatial details. Under matched pre-training data and throughput, T-JEPA achieves leading transfer performance on both static and temporal tasks. Analyses further reveal that T-JEPA learns representations with time-gap-dependent transition predictability and coherent latent dynamics, while maintaining strong cross-period consistency, representation diversity, and semantic discriminability.
Figures & tables
Figure 1: Temporal supervision paradigms for remote sensing representation learning. (a) Relation-based methods impose predefined pairwise constraints between observations. (b) Context-based methods jointly encode multiple observations to form holistic reconstruction context. (c) T-JEPA independently encodes each observation and learns time-gap-conditioned transitions between their representations.
Figure 2: Overview of T-JEPA. A shared online encoder independently processes masked observations with optional spatiotemporal metadata. The reconstruction branch predicts masked image patches, while the time-gap-conditioned joint-embedding branch predicts the complete latent field of a target observation produced by a metadata-free EMA encoder. Gradients are stopped through the target pathway, and all auxiliary modules are discarded after pre-training.
Linear classification
Semantic segmentation
Change detection
BigEarthNet
PASTIS
DFC20 (no metadata)
PASTIS ( T=8 )
OSCD
Method
1%
10%
T=1
T=8
OA
κ
mIoU
OA
κ
mIoU
Prec.
Rec.
F1
Random init.
37.8
41.5
22.7
29.2
58.4
49.4
31.8
79.5
74.4
46.6
47.1
55.9
51.1
MAE
58.3
63.8
43.0
45.9
64.8
57.3
37.6
84.0
80.1
54.7
66.0
50.5
57.3
I-JEPA
59.1
65.1
40.3
44.9
65.5
58.3
38.9
84.6
80.9
56.3
59.2
54.2
56.6
V-JEPA
52.6
58.0
33.7
41.6
63.2
55.6
36.9
84.4
80.5
56.8
56.4
57.2
56.8
Table 1: Downstream task performance. Metrics include macro mean average precision (mAP) for linear classification; overall accuracy (OA), Cohen’s kappa ( κ ), and mean intersection over union (mIoU) for semantic segmentation; and precision, recall, and F1 score for change detection. The best results are in bold .
Figure 3: t-SNE visualization of raw observations and learned representations. We randomly select 10 CACo locations, each with five observations acquired during 2017–2018. Panel (a) shows raw observations, and the remaining panels show representations extracted by the compared pre-trained models. T-JEPA preserves the natural location-wise organization while forming more compact clusters with structured intra-location variation.
Method
CKA
R @ 1
R @ 5
ER
mAP ↑
MAE
0.52
48.31
54.75
134.04
49.14
I-JEPA
0.21
88.78
93.42
233.06
50.96
V-JEPA
0.19
43.82
58.75
156.68
44.38
SiamMAE
0.60
37.21
47.71
132.02
55.15
SeCo
0.34
31.77
49.35
25.69
45.72
CACo
0.40
31.35
47.16
26.28
46.18
Table 2: CKA measures global alignment with the raw-observation geometry; R @ 1 and R @ 5 evaluate local cross-period consistency through retrieval among 2,000 locations; Effective Rank (ER) diagnoses representation diversity and potential rank collapse; and mAP reports semantic discriminability using k NN classification ( k=20 ) on BigEarthNet ( 1% ). These metrics are interpreted jointly, since neither higher CKA nor higher ER alone necessarily indicates a better representation.
Figure 4: Visualization of temporal dynamics learned by T-JEPA. Starting from the anchor observation in the first column, the model predicts and decodes a full annual trajectory. Ground-truth observations are indicated by dates shown in bold red. The predicted states generally follow the seasonal vegetation trends in the available observations and reveal plausible unobserved states, such as continuous winter snow cover. Predictions at cloud-contaminated dates often preserve a clear underlying land-surface state.
Method
PredSim
Gain
Target R @ 1 ↑
Shuffled R @ 1
MAE
0.94
0.05
24.25
18.75
I-JEPA
0.77
0.02
7.80
7.40
V-JEPA
0.91
0.12
63.95
33.50
SiamMAE
0.99
0.01
30.90
19.95
SeCo
0.82
0.07
11.90
9.85
CACo
0.86
0.05
14.75
12.10
Table 3: Time-conditioned predictability of frozen representations. PredSim measures prediction–target similarity, Gain measures improvement over directly using the source representation, and Target R @ 1 retrieves the correct target among five observations of the same location. Shuffled R @ 1 replaces the true time interval with a shuffled one at inference.
Meta.
Pred.
BEN
PASTIS
DFC20
MAE baseline
58.3
45.9
37.6
✓
62.4
47.6
38.8
✓
60.5
46.8
38.3
✓
✓
66.0
50.7
39.4
Table 4: Ablation studies of the temporal prediction branch. We report linear classification mAP on BigEarthNet with 1% labels and PASTIS with T=8 , and mIoU of DFC20 semantic segmentation.
Figure S1: Qualitative comparison on PASTIS ( T=8 ) segmentation.
Figure S2: Qualitative comparison on DFC20 segmentation.
Figure S3: Qualitative comparison on OSCD change detection.
λ
BEN
PASTIS
DFC20
60.5
46.8
38.3
0.01
62.6
48.0
38.0
0.1
64.9
50.1
37.7
0.5
66.0
50.7
39.4
1.0
65.5
50.3
39.4
5.0
65.4
49.5
38.2
Table S1: Ablation study on prediction loss weight.
Method
Params
FLOPs/seq.
GPU time
Throughput
Peak memory
Wall-clock
(M)
(G)
(ms/iter)
(seq./s)
(GB)
(s/epoch)
MAE
49.04
45.24
142.24
224.98
9.57
50.15
T-JEPA
64.06
145.16
267.52
119.62
14.14
60.18
Relative change
+30.64%
+220.90%
+88.08%
−46.83%
+47.83%
+20.00%
Table S2: Pre-training cost of T-JEPA compared with MAE. FLOPs are reported per four-observation sequence. GPU time and throughput exclude data loading, whereas wall-clock time covers the complete training pipeline.
Component
Added FLOPs/seq. (G)
Share of added FLOPs
Forward CUDA time (ms/iter)
EMA target-encoder forward
72.35
72.40%
49.80
Prediction Transformer blocks
27.14
27.16%
18.16
Prediction head
0.35
0.35%
0.60
Prediction preparation
0.09
0.09%
1.03
Time-gap embedder
<0.01
<0.01%
0.24
EMA parameter update
–
–
9.25
Table S3: Component-wise breakdown of the additional pre-training computation introduced by T-JEPA. CUDA times measure the corresponding forward operations, except for the separately reported EMA parameter update.
Space Applications Centre, Indian Space Research Organisation, Ahmedabad, India · Centre of Studies in Resources Engineering, Indian Institute of Technology Bombay, Mumbai, India
aGraduate School of Artificial Intelligence Convergence Engineering, Changwon National University, Changwon, South Korea · bDepartment of Artificial Intelligence Engineering, Changwon National University, Changwon, South Korea