Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle dynamics. Visual compression reduces representation dimension, while Bayesian variable selection reduces active regression dimension. Posterior prediction combines forecasts across predictor subsets using their posterior probabilities as weights and accounts for parameter uncertainty and future disturbances. The model forecasts joint visual--physical states recursively and physical targets directly. Across four forecasting tasks spanning object motion, vegetation greenness and solar power, ViBR-WM achieves lower mean overall physical-target error than Temporal Straightening, ConvLSTM, PredRNN and SimVP on every task. Repeated fitting and resampling support these overall gains.
Figures & tables
Figure 1: Predictive architectures and Bayesian inference. (a) TS predicts future features from encoded images, actions and observed measurements. During training, prediction loss compares predicted features with encoded future observations; the curvature penalty discourages changes in direction between consecutive displacements along encoded trajectories of observed images. (b) ViBR-WM’s trend, seasonal, cycle and regression contributions add before posterior prediction. Dashed arrows expand the mechanisms below. (c) Principal component analysis (PCA) compresses visual features, which join other covariates as regression inputs. Bayesian selection assigns posterior probabilities to predictor subsets. Posterior prediction combines the corresponding predictive distributions using these probabilities as weights. (d) Parameter uncertainty and future disturbances yield distributions over possible futures.
Figure 2: Wall accuracy and forecasting. (a) Overall NRMSE: mean ± sample standard deviation (SD) across ten fits, with the lowest mean highlighted; percentages show how much lower ViBR-WM’s mean error is, expressed as a percentage of each comparator’s mean error. (b) Forecasting workflow for the test trajectory in (c–d). DINOv2 extracts features from images at frames 5 and 10. These pass through a projector trained with the TS objective. PCA then yields four visual principal-component (PC) coordinates per frame. Each state-history column contains these four visual coordinates and the observed horizontal and vertical positions, x and y , forming a six-dimensional state. The joint-transition model uses the states at frames 5 and 10 and the known actions to predict the state at frame 15. Each predicted state updates the history for the next five-frame transition. Each simulated trajectory uses one parameter set sampled from the posterior, held fixed through all its transitions. (c) Filled points show the PC values at frames 5 and 10 used in the history; open points show the observed PC values after frame 10. (d) Observed positions and ViBR-WM posterior means; labels identify frames.
Figure 3: Forecasting accuracy on four tasks (columns A–D). Top: overall mean NRMSE across ten repetitions; error bars span one sample standard deviation on either side of the mean. Overall scores summarize four forecast horizons and, for PhenoCam, both camera sites. Bottom: mean NRMSE by forecast horizon. PointMaze and Wall horizons give the number of frames after forecasting starts at frame 10. NRMSE normalizes errors by the corresponding training-target standard deviations. Appendices C – E report 95% bootstrap intervals for differences between methods’ mean errors.
Figure 4: Thirteen-week-ahead PhenoCam forecasts from repetition 1 at the Mead 1 (top) and Mead 3 (bottom) camera sites. Columns show weeks 15–26 and 30–52 of 2023. Forecasts for weeks 27–29 are unavailable because the required input data are incomplete. Vertical axes show green chromatic coordinate (Gcc). Black curves show observed Gcc; colored curves show predictions from the five methods. Blue shading denotes ViBR-WM’s 95% predictive interval for each week.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Forecast
Candidate design
Inclusion prior
PointMaze
Recursive, m=8
P=107 : two states, actions, state–action products and regional terms
All π=1
Wall
Recursive, m=6
P=42 : two states and actions
All π=1
PhenoCam
Direct, m=1
P=14 : 4 visual, 4 Gcc-history and 6 calendar terms
Visual π=0.5 ; others π=1
SKIPP’D
Direct, m=1
P=58 : 30 visual, 16 power-history and 12 calendar terms
Visual π=0.5 ; others π=1
Appendix
Table S1: Selected ViBR-WM specifications. P counts candidate predictors per response. All rows use regression only, κ=0.01 and R2=0.8 . Coefficients are estimated for all included terms.
Method
Frame 15
Frame 25
Frame 35
Frame 45
ViBR-WM
0.1197
0.2880
0.4045
0.4834
Temporal Straightening
0.2965
0.3767
0.4751
0.5552
ConvLSTM
0.7135
0.8363
1.0050
1.1157
PredRNN
0.6338
0.9056
0.9459
1.0108
SimVP
0.6271
0.8024
0.9946
1.0977
Appendix
Table S2: Wall per-target NRMSE averaged equally over the same 50 trajectories and ten fits. Within each target and trajectory, squared standardized error is averaged over two coordinates before taking its square root. The overall score takes RMS jointly over horizons and coordinates within each trajectory.
Repetition
ViBR-WM
Temporal Straightening
ConvLSTM
PredRNN
SimVP
1
0.379430
0.477133
1.054874
0.968377
0.902873
2
0.369477
0.431232
0.972504
0.899301
0.848144
3
0.345590
0.444623
0.948536
0.944810
0.962504
4
0.364094
0.405766
0.964438
0.947825
0.961308
5
0.357922
0.458497
1.127643
0.905311
1.007133
6
0.358408
0.459781
0.909451
0.937824
1.086151
Appendix
Table S3: All ten Wall fit aggregates. Each cell averages whole-trajectory RMS scores; sample SD uses denominator nine. ViBR-WM and Temporal Straightening share the trained representation within each repetition; RGB models are initialized and trained independently.
Comparator
Target
Difference
95% interval
Temporal Straightening
Overall
-0.0951
[−0.1554,−0.0350]
Frame 15
-0.1767
[−0.2183,−0.1309]
Frame 25
-0.0887
[−0.1443,−0.0276]
Frame 35
-0.0706
[−0.1510,0.0084]
Frame 45
-0.0717
[−0.1624,0.0214]
ConvLSTM
Overall
-0.6141
[−0.7168,−0.5119]
Appendix
Table S4: Wall overall and per-target ViBR-WM-minus-comparator error differences and 95% percentile intervals from 2,000 bootstrap resamples, without adjustment for multiple comparisons. Each resample uses the same selected trajectories across methods and horizons; ViBR-WM and Temporal Straightening resample paired fits that share a trained visual representation, while RGB fits are resampled independently. Conditioning assumptions are given in Section B .
Design
q
L
P
κ
NRMSE
Additive
4
2
46
0.01
0.640590
Additive
8
2
54
0.01
0.658151
Additive
16
2
70
0.01
0.670158
Additive
4
3
54
0.01
0.659372
Additive
8
3
66
0.01
0.674185
Additive
16
3
90
0.01
0.691281
Appendix
Table S5: PointMaze validation candidates. q is visual dimension, L history length and P predictors per response.
Method
Frame 15
Frame 25
Frame 35
Frame 60
ViBR-WM
0.3342
0.4936
0.6138
0.6362
Temporal Straightening
0.4293
0.5921
0.6706
0.8003
ConvLSTM
1.4612
3.1474
2.8551
4.1492
PredRNN
1.3558
1.0120
1.0448
0.9708
SimVP
1.2258
1.1962
1.3989
1.6190
Appendix
Table S6: PointMaze per-target NRMSE averaged equally over the same 50 trajectories and ten fits. Within each target and trajectory, squared standardized error is averaged over four coordinates before taking its square root. The overall score takes RMS jointly over horizons and coordinates within each trajectory.
Repetition
ViBR-WM
Temporal Straightening
ConvLSTM
PredRNN
SimVP
1
0.538330
0.621356
3.470498
1.132026
2.131528
2
0.548061
0.661925
4.247897
1.140820
1.134285
3
0.588117
0.659380
3.140727
1.141516
1.161737
4
0.566431
0.692336
2.428698
1.136431
1.781546
5
0.571354
0.690617
2.307833
1.141411
1.319354
6
0.535430
0.652987
3.720713
1.141336
1.147300
Appendix
Table S7: All ten PointMaze fit aggregates. Each cell averages whole-trajectory RMS scores; sample SD uses denominator nine. ViBR-WM and Temporal Straightening share the trained representation within each repetition; RGB models are initialized and trained independently.
Comparator
ViBR-WM minus comparator
95% interval
Temporal Straightening
-0.1082
[−0.1544,−0.0532]
ConvLSTM
-2.5604
[−2.9868,−2.1363]
PredRNN
-0.5805
[−0.6817,−0.4778]
SimVP
-0.8754
[−1.0927,−0.6840]
Appendix
Table S8: PointMaze overall ViBR-WM-minus-comparator error differences and 95% percentile intervals from 2,000 bootstrap resamples, without adjustment for multiple comparisons. Each resample uses the same selected trajectories across methods and horizons; ViBR-WM and Temporal Straightening resample paired fits that share a trained visual representation, while RGB fits are resampled independently. Conditioning assumptions are given in Section B .
Structural states
K
Inclusion
Pass
NRMSE
Level + seasonal
16
All fixed
8/8
0.506103
Level
16
All fixed
8/8
0.456370
None
16
All fixed
8/8
0.385197
None
16
Visual selection
8/8
0.360727
None
16
All selectable
7/8
Unqualified
None
4
All fixed
8/8
0.363566
Appendix
Table S9: PhenoCam validation sensitivity within the declared search. All candidates use the same 2018–2020 training data, 2021 validation targets, κ=0.01 and R2=0.8 . NRMSE averages two sites and four horizons. “Pass” counts the site–horizon combinations passing predictive-sampling checks out of eight; no aggregate is reported for an unqualified candidate.
Repetition
ViBR-WM
Temporal Straightening
ConvLSTM
PredRNN
SimVP
1
0.436729
1.140654
1.006600
1.044145
0.857985
2
0.436722
1.354659
0.986479
1.049004
0.793107
3
0.436683
0.998471
0.991100
0.976366
1.151777
4
0.436687
3.403910
0.952713
0.976920
0.829342
5
0.436793
0.781621
0.977743
1.104369
0.757989
6
0.436661
4.080768
0.982536
1.081207
1.011120
Appendix
Table S10: All ten PhenoCam fit aggregates. Each entry averages two fixed sites and four horizons; sample SD uses the ten aggregates with denominator nine. Repetition indices pair sites within a method; methods use independent training randomness. ViBR-WM repeats vary MCMC seeds on fixed DINO features, PCA and design; neural fits vary initialization.
Comparator
13-week primary
26-week sensitivity
Temporal Straightening
[−1.9527,−0.6770]
[−1.9873,−0.7291]
ConvLSTM
[−0.6770,−0.3550]
[−0.6854,−0.4865]
PredRNN
[−0.7012,−0.4316]
[−0.6451,−0.4588]
SimVP
[−0.7025,−0.2146]
[−0.7276,−0.3794]
Appendix
Table S11: Unadjusted 95% bootstrap percentile intervals for overall ViBR-WM-minus-comparator NRMSE; negative values favor ViBR-WM. Each block length uses 2,000 draws. Shared calendar weights preserve joint dependence across sites and horizons; training repeats are sampled independently between methods.
Method
Overall NRMSE ± SD
ViBR-WM reduction (%)
ViBR-WM
0.2180±0.000011
—
Temporal Straightening
0.5448±0.097869
59.99
ConvLSTM
0.5072±0.047198
57.02
PredRNN
0.4290±0.019511
49.20
SimVP
0.4892±0.024810
55.45
Appendix
Table S12: SKIPP’D 2019 forecasting comparison. Means equally average four horizons and ten fits; SD is the sample standard deviation of ten fit means. Percentages show how much lower ViBR-WM’s mean NRMSE is, relative to each comparator’s mean NRMSE. Training and validation use data from 2017 and 2018, respectively.
Method
5 min
15 min
30 min
60 min
ViBR-WM
0.1583
0.2046
0.2345
0.2744
Temporal Straightening
0.2471
0.4144
0.6198
0.8979
ConvLSTM
0.4737
0.4807
0.4985
0.5758
PredRNN
0.3814
0.3688
0.4405
0.5254
SimVP
0.4435
0.4571
0.5004
0.5557
Test origins
3124
3124
3124
3124
Appendix
Table S13: Mean NRMSE over ten fits on 2019 forecast starting times at each horizon. For each fit and horizon, RMSE across these times is divided by the corresponding 2017 training-target standard deviation. ViBR-WM uses visual, observed-PV and calendar regressors; Temporal Straightening receives RGB images and PV observations through its visual and proprioceptive interfaces, respectively; ConvLSTM, PredRNN and SimVP use RGB only.
Repetition
ViBR-WM
TS
ConvLSTM
PredRNN
SimVP
1
0.217960
0.443109
0.505898
0.426314
0.520961
2
0.217946
0.653381
0.527859
0.418804
0.433620
3
0.217951
0.435658
0.525488
0.401077
0.501159
4
0.217981
0.531924
0.473520
0.454480
0.515727
5
0.217958
0.487822
0.449582
0.423499
0.495842
6
0.217953
0.597544
0.541185
0.451756
0.494600
Appendix
Table S14: All ten fit means, each equally averaging four horizon-specific NRMSEs. Four chains form one ViBR-WM fit; repetitions vary MCMC seeds on fixed training features and design. Neural models are initialized and trained independently. Sample SD uses denominator nine.
Comparator
Difference
95% interval
Temporal Straightening
-0.3268
[−0.3881,−0.2688]
ConvLSTM
-0.2892
[−0.3330,−0.2483]
PredRNN
-0.2111
[−0.2384,−0.1871]
SimVP
-0.2712
[−0.3004,−0.2412]
Appendix
Table S15: Overall ViBR-WM-minus-comparator NRMSE and percentile intervals from 2,000 bootstrap draws. Calendar-week blocks are shared across methods and horizons; training fits are resampled independently within each method. Negative values favor ViBR-WM. Intervals are unadjusted for multiple comparisons.
Dataset
CRPS
Energy
95% cov. (%)
95% width
PhenoCam
0.237
—
78.87
1.096
PointMaze
0.302
0.737
93.66
2.220
Wall
0.209
0.331
89.03
1.159
SKIPP’D
0.111
—
97.43
1.223
Appendix
Table S16: ViBR-WM posterior predictive scores. Each fitted model’s posterior predictive distribution is scored first, then the scores are averaged over fits and the task’s evaluation units. Scores and interval widths are expressed in training-standard-deviation units; coverage (cov.) is the percentage of observed targets within the central 95% predictive intervals. For scalar Gcc and PV targets, the energy score equals CRPS and is omitted.
Task
Method
CRPS
Energy
95% cov. (%)
95% width
PhenoCam
ViBR-WM
0.293
—
0.28
0.001
PhenoCam
Temporal Straightening
0.546
—
85.45
4.811
PhenoCam
ConvLSTM
0.607
—
44.39
0.879
PhenoCam
PredRNN
0.460
—
79.86
1.740
PhenoCam
SimVP
0.432
—
67.57
1.250
PointMaze
ViBR-WM
0.388
0.937
15.25
0.238
Appendix
Table S17: Between-fit dispersion, represented by one equal-weight empirical distribution of ten predicted means per method. CRPS, energy scores and interval widths are expressed in training-standard-deviation units; coverage (cov.) is the percentage of observed targets within the central 95% empirical intervals. PhenoCam averages equally over horizons and sites, PointMaze and Wall over trajectories and horizons, and SKIPP’D over forecast starting times and horizons. The energy score is omitted for scalar Gcc and PV because it equals CRPS. Central 95% intervals use linearly interpolated empirical quantiles.
Method
PointMaze
PhenoCam
ViBR-WM (unchanged)
0.560177
0.436700
Temporal Straightening adaptation
0.669387
0.327527
ConvLSTM adaptation
0.543618
0.533886
PredRNN adaptation
0.505283
0.354508
SimVP adaptation
0.529514
0.342000
Appendix
Table S18: Sensitivity analysis: NRMSE under task-specific input, training or readout adaptations.
Predictive distribution
D5=0
D5=2
D5=4
Regression posterior
3.160±0.017
3.235±0.016
3.445±0.017
Regression mean only
4.465±0.024
4.572±0.023
4.882±0.024
Recentered empirical paths
3.182±0.017
3.257±0.016
3.468±0.017
Moment-matched Gaussian
3.135±0.013
3.210±0.013
3.421±0.013
True two-branch mixture
3.135±0.013
3.210±0.013
3.421±0.013
Appendix
Table S19: Multiple-future diagnostic: mean whole-path energy score for five-step trajectories, ± the standard deviation across 20 data seeds; lower is better. D5 measures separation between the two branch means relative to within-branch variation. The regression component is evaluated without a visual encoder.
Task
ViBR-WM
TS
ConvLSTM
PredRNN
SimVP
PointMaze
80.37
8.30
9.23
68.94
23.76
Wall
30.89
5.07
3.52
23.79
8.65
PhenoCam
5.36
18.84
2.97
11.18
5.05
SKIPP’D
267.44
289.82
182.65
1393.96
119.35
Appendix
Table S20: Mean elapsed fitting-stage time in minutes. PhenoCam is per site, with all four horizons included in each statistical fit.
Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evolution. Deep forecasters can fit temporal correlations, and physics-informed models can regularize predictions with scientific constraints, but these directions are often connected only at the decoded-output level. As a result, the hidden predictive state that generates future trajectories may remain statistically useful but physically unstructured. We introduce Phys-JEPA, a physics-informed joint-embedding predictive architecture for multivariate time-series forecasting. Phys-JEPA learns a latent world model in which predictive states are decomposed into physical and residual components, and physical consistency is imposed directly on latent states and latent transitions rather than only on decoded forecasts. This formulation uses known physical variables to organize the representation space while retaining residual capacity for unresolved dynamics. On Jena Climate 2009--2016, Phys-JEPA reduces aggregate MSE from 0.12482 to 0.12273 and temperature MSE from 0.01892 to 0.01831 at H=24. On Traffic, full Phys-JEPA improves aggregate MSE over the supervised baseline across all tested horizons, reducing H=192 MSE from 0.800784 to 0.773873. On Electricity, the best variant depends on horizon: static latent consistency is strongest at H=24 and H=48, while full Phys-JEPA gives the best aggregate and target-variable MSE at H=192. These initial results suggest that moving physics-informed learning from output space to latent predictive state space is a promising direction for interpretable temporal world models.
Learning predictive world models from visual observations is a core problem in embodied AI, with applications to model-based reinforcement learning and robotic planning. Existing latent world models typically generate future states with unconstrained neural transition functions, while modern video generation systems often prioritize perceptual plausibility or introduce physical structure through auxiliary losses, external guidance, or separate dynamics modules. As a result, long-horizon rollouts can remain weakly grounded in the physical principles that govern real dynamics, leading to compounding error, energy drift, and physically inconsistent futures. We propose Least Action World Models (LaWM), a latent world-modeling framework that operationalizes the Principle of Least Action in learned visual latent space: future rollouts are governed by a learned Lagrangian action functional rather than produced only by an unconstrained transition predictor. Our main technical realization is a latent variational integrator: LaWM encodes observations into learned generalized coordinates, learns a latent discrete Lagrangian over consecutive latent states, constructs a discrete action functional, and advances prediction by solving the corresponding discrete integration condition. Thus, physical structure is not merely used to score, regularize, or constrain a completed trajectory; it defines the latent transition rule itself. Because the transition is induced by a discrete variational principle, LaWM provides a structure-preserving bias for long-horizon visual prediction. Across physics-clean synthetic dynamics and embodied robot interaction benchmarks, LaWM improves physical invariance, background consistency, motion smoothness, and appearance and geometric prediction metrics over video-generation and world-model baselines.
Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing.We introduce EO-WM, a video diffusion transformer for multispectral EO forecasting. EO-WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather-response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity-aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched-Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO-WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63% and improves directional hit rate by a relative 7.80%, while remaining competitive on standard pixel-level metrics. The benchmarks and model will be made open-source at https://github.com/Luo-Z13/EO-WM.