RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction--diffusion equations
Organizations: School of Data Science, The Chinese University of Hong Kong, Shenzhen, Shenzhen, China
Abstract
Learning surrogates for time-dependent partial differential equations often requires a new simulation corpus when the governing operator changes. We introduce RD-JEPA, a joint-embedding predictive architecture for self-supervised pretraining on reaction-diffusion trajectories. A single model is pretrained on five parameterized systems and then adapted to three held-out systems whose reaction operators and trajectories are excluded from pretraining. Using one, five, or ten complete trajectories from a held-out system, RD-JEPA achieves lower mean relative discrete field error and mean absolute spatial first-difference error than five supervised surrogate baselines, an independently trained control that removes the trajectory-dependent predictive latent pathway, and an architecture-matched model trained from scratch. Within the evaluated equations, output resolution, forecast horizons, and choices of adaptation trajectories, the results indicate that prediction of future-state representations can support data-efficient adaptation across related reaction-diffusion systems.
Figures & tables
| Relative | Gradient | ||||||
|---|---|---|---|---|---|---|---|
| PDE | Model | ||||||
| Lambda–Omega | RD-JEPA | ||||||
| FNO | |||||||
| LNO | |||||||
| ReViT | |||||||
| RieszNO | |||||||
| System | Domain and grid | Integrator and spatial discretization | Internal | Saved window |
|---|---|---|---|---|
| Gray–Scott | , | Explicit Euler; second-order periodic central differences | ||
| FitzHugh–Nagumo | Periodic lattice, | RK4; fourth-order periodic finite difference Laplacian | ||
| Brusselator | , | Explicit Euler; second-order periodic central differences | ||
| Complex Ginzburg–Landau | , | Explicit Euler; second-order periodic central differences | ||
| Schnakenberg | , | Explicit Euler; second-order periodic central differences | ||
| Lambda–Omega | , | Explicit Euler; second-order periodic central differences |
| Target type | Probability | Selection of time offsets |
|---|---|---|
| Future block | 0.50 | , where is sampled uniformly from . |
| Future tube | 0.40 | contains four distinct offsets sampled uniformly without replacement from . The same spatial region is used at each selected offset. |
| Same time | 0.10 | , targeting the latest context frame . The patches in are masked in the online encoder input. |
| Component | Input | Output |
|---|---|---|
| Online encoder | ||
| Target encoder | ||
| Latent predictor | context tokens and query pairs | predicted tokens |
| Forecast decoder | predicted tokens and the latest context field |
| Setting | Value |
|---|---|
| Shared settings | |
| Optimizer | AdamW, |
| Batch size | 4 |
| Gradient clipping | Global norm 1.0 |
| Predictive pretraining | |
| Trainable modules | Online encoder and predictor |
| Setting | RD-JEPA | No predictive latent | External baselines |
|---|---|---|---|
| Training updates | 5,000 | 5,000 | 5,000 |
| Batch size | 4 | 4 | 4 |
| Optimizer | AdamW | AdamW | AdamW |
| Learning rate | predictor; decoder | conditioner; decoder | |
| Weight decay | |||
| Adam betas |
| PDE | Model | Avg. | |||||
|---|---|---|---|---|---|---|---|
| Gray–Scott | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| FitzHugh–Nagumo | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO |
| PDE | Model | Avg. | |||||
|---|---|---|---|---|---|---|---|
| Gray–Scott | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| FitzHugh–Nagumo | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO |
| PDE | Model | Avg. | |||||
|---|---|---|---|---|---|---|---|
| Gray–Scott | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| FitzHugh–Nagumo | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO |
| PDE | Model | Avg. | |||||
|---|---|---|---|---|---|---|---|
| Gray–Scott | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| FitzHugh–Nagumo | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO |
| PDE | Model | Avg. | |||||
|---|---|---|---|---|---|---|---|
| Gray–Scott | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| FitzHugh–Nagumo | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO |
| PDE | Model | Avg. | |||||
|---|---|---|---|---|---|---|---|
| Gray–Scott | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| FitzHugh–Nagumo | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO |
| Relative | Gradient | ||||||
|---|---|---|---|---|---|---|---|
| Model | ID | Coeff.-OOD | IC-OOD | ID | Coeff.-OOD | IC-OOD | |
| RD-JEPA | 5 | ||||||
| 10 | |||||||
| 20 | |||||||
| No predictive latent | 5 | ||||||
| 10 | |||||||
| Relative | Gradient | ||||||
|---|---|---|---|---|---|---|---|
| PDE | RD-JEPA | No predictive latent | Reduction (%) | RD-JEPA | No predictive latent | Reduction (%) | |
| Lambda–Omega | 1 | ||||||
| 5 | |||||||
| 10 | |||||||
| Barkley | 1 | ||||||
| 5 | |||||||
| Relative | Gradient | ||||||
|---|---|---|---|---|---|---|---|
| PDE | Model | ||||||
| Lambda–Omega | RD-JEPA | ||||||
| Architecture-matched scratch | |||||||
| Barkley | RD-JEPA | ||||||
| Architecture-matched scratch | |||||||
| Oregonator | RD-JEPA | ||||||
| Relative | Gradient | ||||||
|---|---|---|---|---|---|---|---|
| Boundary condition | Model | ||||||
| Neumann | RD-JEPA | ||||||
| No predictive latent | |||||||
| FNO | |||||||
| Robin | RD-JEPA | ||||||
| No predictive latent | |||||||