A time series world model (TSWM) predicts a controlled system's state from its observed history and planned actions and exogenous inputs. Current approaches build forecasters with actions as covariates, trained and evaluated on prediction error under the executed plan. Yet world models compare unexecuted plans, but their responses to changed plans remain untested. We ask which design choices matter and whether accurate forecasters respond to changed plans as real systems do. We address both with a formalization and benchmark. The formalization separates state, actions and exogenous inputs, distinguishes continuous, mode and event actions, and introduces mechanism consistency, a metric built on declared action-state relations with known directions, such as a vasopressor raising blood pressure: it checks whether shifting an action moves the forecast in the declared direction. The benchmark consolidates eight public datasets with real actions from engineered infrastructure and clinical care, varying prediction space, plan fusion and plan encoding across seven backbones and five seeds. First, a frozen latent prediction space lowers MAE by 9.9% over observation space and gated output fusion lowers it by 12.7% over input concatenation on average, with both improving all eight datasets; temporal plan encoding changes average MAE by at most 2.2%. Second, prediction error and mechanism consistency diverge: the lowest-error configuration is at or below chance in consistency on four of five datasets with declared mechanisms, and no design choice avoids this. Finally, directional supervision, a loss penalizing the wrong-signed part of the response to a shifted action, significantly raises consistency on penalized mechanisms with no change in MAE. Together they give TSWMs a recipe: a frozen latent space and output-side fusion for accuracy, and a training objective for mechanism consistency.
Figures & tables
Figure 1: Two plans from the same history. We sweep three design choices and score mechanism consistency.
Figure 2: Overview of the formalization and the model. Top: the channel roles, the window of history, plan and outcome, and the two metrics (Section 3.1 ). Bottom: the model and its three design choices (Section 3.2 ); the concatenation baseline conditions the backbone instead (Appendix B ).
scale
series dim.
actions
Dataset
Domain
Δt
systems
windows
state
exog.
cont.
mode ( K )
event
Greenhouse
horticulture
5 min
5/1
163
4
8
7
–
–
PleiaData
building HVAC
10 min
50/12
7,548
2
8
1
2 (2,6)
–
PREDIST
district heating
10 min
46/12
6,109
5
2
2
3 (2,3,2)
–
Wastewater
water treatment
2 min
1 plant
1,311
3
2
3
2 (2,2)
–
VitalDB
anesthesia
2 s
2,754/689
40,193
4
–
2
–
–
Table 1: The eight datasets, four engineered and four clinical. Systems are counted train/held-out; windows are held-out windows at L=256 , H=16 , stride 80 . The action columns are the three factors of U : dc continuous channels, dm mode channels with their cardinalities Kj , and the number of event types ∣A∣ .
Dataset
Action
State
Sign
Greenhouse
heating-pipe temp.
air temp.
+
leeward vent
air temp.
−
windward vent
air temp.
−
CO 2 dosing
CO 2 level
+
VitalDB
propofol
BIS
−
propofol
MAP
−
Table 2: The twenty-one declared mechanisms. A plus sign means that raising the action raises the state, a minus sign that it lowers it; the PleiaData mechanisms hold in the named mode. BIS, bispectral index; MAP, mean arterial pressure; FiO 2 , inspired oxygen fraction.
Option
Greenhouse
PleiaData
PREDIST
Wastewater
VitalDB
CGMacros
Shanghai
MIMIC-Cardio
Rank
Space
Observation
.0821 ±.0269
.0476 ±.0306
.1866 ±.0099
.1718 ±.0116
.0377 ±.0090
.0131 ±.0054
.1330 ±.0132
.0562 ±.0029
2.52
AE
.0761 ±.0200
.0337 ±.0252
.1821 ±.0047
.1695 ±.0123
.0329 ±.0026
.0114 ±.0049
.1249 ±.0096
.0552 ±.0021
1.93
VAE
.0956 ±.0134
.0339 ±.0244
.1836 ±.0054
.1672 ±.0106
.0336 ±.0039
.0105 ±.0025
.1206 ±.0048
.0552 ±.0020
1.95
JEPA
.1041 ±.0271
.0607 ±.0359
.2143 ±.0327
.1824 ±.0266
.0403 ±.0112
.0270 ±.0233
.1418 ±.0208
.0621 ±.0095
3.61
Fusion
None
.0751 ±.0364
.0292 ±.0172
.1858 ±.0219
.1928 ±.0085
.0396 ±.0187
.0174 ±.0187
.1257 ±.0155
.0608 ±.0174
4.89
Concat
.0761 ±.0200
.0337 ±.0252
.1821 ±.0047
.1695 ±.0123
.0329 ±.0026
.0114 ±.0049
.1249 ±.0096
.0552 ±.0021
5.70
Table 3: Validation MAE of the three sweeps. Mean ± std over the 35 backbone–seed runs of one dataset and one option ( 70 for encoding, which pools Gate and FiLM-0 ); the std spans backbones. The fixed choices of each sweep are given in Section 4.2 . Bold is the lowest mean of each block, Rank is the mean rank over dataset–backbone cells, and the shaded row has the best rank of its block. Lower is better.
Figure 3: Mechanism consistency of the three sweeps. MC at δ=±0.1 , averaged over both directions, the declared mechanisms and the runs of Table 3 . White is chance, green agrees with the declared sign, red opposes it. Boxes mark the lowest-MAE choice of each row. None is omitted because its forecast does not move. Numbers are in Appendix E .
Figure 4: Directional supervision against the Gate control. Left: mechanism consistency ( MC at δ=+0.1 ) on the penalized mechanisms, mean over the 35 backbone–seed runs of each dataset, with the paired gain. Right: change in validation MAE relative to Gate per run, with the median (diamond), a 90% interval (bar) and the ±5% band (shaded).
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Adjusted forecast zt+k′
Details
None
zt+k
Plan discarded.
Concat
zt+k , backbone input [zs;as]
The history plan enters with the latent history; the future plan at+k enters through the native covariate interface of TimeXer and TiDE, and through MLP([zt+k;at+k]) otherwise.
Res
zt+k+g(ϕk)
Standard initialization.
Res-0
zt+k+g(ϕk)
Output layer of g zero-initialized.
FiLM-0
(1+γk)⊙zt+k+βk
[γk;βk]=g(ϕk) , output layer zero-initialized.
Gate
zt+k+αk⊙σ(bk)
[αk;bk]=g([ϕk;zt+k]) , σ the logistic function.
Appendix
Table 4: Plan-fusion methods. zt+k is the plan-blind forecast of the backbone at horizon step t+k and zt+k′ the adjusted forecast that is decoded; a is the encoded plan. Concat is the only method whose plan enters the backbone itself.
Encoder
Encoded plan as
Details
Instant
as
Identity.
Decay
as+η⊙(ms−as)
ms=(1−α)⊙as+α⊙ms−1 with α=σ(θ) learned per channel; the gate η is zero-initialized.
Conv
vs(6) , with v(0)=a
Six residual blocks vs(l)=vs(l−1)+GELU(cs(l)) with cs(l)=∑j=07wj(l)⊙vs−2l−1j(l−1) : depthwise kernels of eight taps with dilation 2l−1 , zero-initialized; receptive field 442 steps.
Attn
as+WoMHA(os,o≤s,o≤s)
os=Wias+ps′ with position embeddings ps′ ; causal mask; four heads, width 64 ; Wo zero-initialized.
Appendix
Table 5: Plan encoders. Each maps the plan at−L+1:t+H to at−L+1:t+H of the same width, causally and per channel unless stated otherwise; all parameters are per-channel vectors in RC except for Attn .
Forecasting protocol (Experiments 1–3, the diagnostic of Appendix G and Section 5 , whose added term is in Appendix F )
halve the learning rate after 5 epochs without held-out improvement
Stopping
at most 100 epochs; stop after 20 epochs without held-out improvement; the checkpoint with the lowest held-out loss is kept
Loss
MSE between the decoded forecast and the observed future states; gradients clipped at norm 1
Seeds
training seeds 0 – 4 (initialization, minibatch order, dropout); the split is fixed at seed 0 with a held-out fraction of 0.2
Appendix
Table 6: Hyperparameters. Shared by every dataset, backbone, design choice and seed unless a row says otherwise. L , H and dc are as in Section 3.1 ; d is the latent width; MAD is the median absolute deviation.
Sweep
Method
Mean rank
Δ MAE vs. ref.
Params.
Prediction space
Observation
2.52
–
–
AE
1.93
−9.87%
–
VAE
1.95
−7.54%
–
JEPA
3.61
+20.21%
–
Plan fusion
None
4.89
+5.73%
0
Concat
5.70
–
–
Appendix
Table 7: Additional summaries of the three sweeps. References are observation space, Concat , and Instant , respectively. Relative MAE and ranks use seed-averaged paired cells. Parameter counts are dataset-averaged fusion-adapter or temporal-encoder sizes, excluding the backbone, codec, and shared embeddings.
Exp.
Contrast
Δ MAE
95% CI (bootstrap over datasets)
Datasets
1
AE vs. observation
−9.87%
[−17.71,−3.93]%
8/8
1
VAE vs. observation
−7.54%
[−16.35,+1.96]%
7/8
1
JEPA vs. observation
+20.21%
[+9.28,+37.40]%
0/8
1
VAE vs. AE
+2.59%
[−1.93,+9.75]%
3/8
2
Gate vs. Concat
−12.73%
[−19.03,−7.12]%
8/8
2
Gate vs. None
−17.47%
[−24.22,−10.90]%
8/8
Appendix
Table 8: Paired effects of the three sweeps. Δ MAE is the geometric-mean change in validation MAE of the first method against the second, with its 95% bootstrap interval over datasets; Datasets counts the datasets on which the first method is better. Negative values favor the first method.
Option
Greenhouse
PleiaData
PREDIST
VitalDB
MIMIC-Cardio
Mean
Space
Observation
.61 ±.13
.49 ±.10
.97 ±.03
.51 ±.10
.35 ±.14
.59
AE
.55 ±.17
.51 ±.11
.92 ±.15
.55 ±.10
.33 ±.17
.57
VAE
.47 ±.26
.53 ±.15
.98 ±.03
.47 ±.05
.33 ±.08
.55
JEPA
.55 ±.15
.44 ±.16
.88 ±.13
.51 ±.11
.33 ±.10
.54
Fusion
Concat
.55 ±.17
.51 ±.11
.92 ±.15
.55 ±.10
.33 ±.17
.57
Res
.49 ±.11
.46 ±.10
.97 ±.02
.35 ±.09
.48 ±.09
.55
Appendix
Table 10: Mechanism consistency of the three sweeps, the numbers behind Figure 3 . The MC metric of Section 3.1 at δ=±0.1 , averaged over both directions and over the mechanisms a dataset declares, as mean ± std over the 35 backbone–seed runs of one dataset and one option, or 70 for encoding, which pools Gate and FiLM-0 . Chance is 0.5 and 1 is perfect agreement with the declared mechanism. None discards the plan, so its forecast does not move and it is omitted. Bold marks the highest unrounded mean of each block in each column; Mean averages the five datasets. Wastewater is set aside, see Appendix E.1 and Table 11 . Higher is better.
Dataset
Mechanism
Sign
Concat
Gate
FiLM-0
seed range
Greenhouse
Tpipe → Tair
+
.83
.97
.99
.03
CO2dosing → CO2air
+
.48
.08
.17
.29
VentLee → Tair
−
.61
.64
.36
.77
VentWind → Tair
−
.28
.19
.22
.66
PleiaData
setpoint_norm@mode=1 → indoor_temp
+
.80
.60
.50
.75
setpoint_norm@mode=2 → indoor_temp
−
.21
.37
.42
.57
Appendix
Table 11: Mechanism consistency per declared relation under the three fusion methods that Section 4.3 discusses. Mean over 35 backbone–seed cells; the four engineered datasets first, then the two clinical ones. Sign is the declared direction. Seed range is the median over backbones of the spread of the five seeds of one cell configuration on that relation.
δ=+0.1
δ=−0.1
lowest cell
Dataset
Action → State
Sign
Pen.
Gate
+ DS
Gate
+ DS
+ DS, δ=+0.1
Greenhouse
heating-pipe temp. → air temp.
+
✓
0.973
1.000
0.974
1.000
1.000
Greenhouse
leeward vent → air temp.
−
✓
0.653
1.000
0.634
1.000
0.991
Greenhouse
windward vent → air temp.
−
✓
0.196
1.000
0.181
1.000
1.000
Greenhouse
CO 2 dosing → CO 2 level
+
0.097
0.154
0.068
0.126
0.000
PleiaData
setpoint (heating) → indoor temp.
+
✓
0.606
1.000
0.601
1.000
1.000
Appendix
Table 12: Mechanism consistency of every declared relation with and without directional supervision. Mean over the 35 backbone–seed runs of Gate and of Gate + directional supervision ( + DS) under the upward and the downward shift; Pen. marks the twelve penalized mechanisms, the training shift is upward only. The last column is the lowest of the 35 runs of + DS under the upward shift. The last block holds the two relations that Section 4.1 sets aside; the on/off relation is categorical and has no downward shift.
Dataset
Pen.
Gate
+ DS
Δ MAE (median)
90% interval
Greenhouse
3
0.0632
0.0639
+0.03%
[−0.24,+2.11]%
PleiaData
1
0.0205
0.0198
− 1.54%
[−9.60,+3.34]%
PREDIST
2
0.1641
0.1641
+0.35%
[−0.34,+0.38]%
VitalDB
2
0.0313
0.0312
− 0.07%
[−0.70,+0.33]%
MIMIC-Cardio
4
0.0543
0.0542
− 0.15%
[−0.26,−0.03]%
Wastewater
0
0.1500
0.1498
+0.00%
[−0.59,+0.34]%
Appendix
Table 13: Validation MAE with and without directional supervision. Mean over the 35 backbone–seed runs of each dataset; Pen. is the number of penalized mechanisms. Δ MAE is the median of the paired relative change and the interval is a t -based 90% interval on the paired log ratios, as in Figure 4 . The pooled row pairs all 210 runs. Wastewater has no penalized mechanism and is the null control.
Dataset
λ=0
λ=0.1
λ=0.3
MAE (fused, H=16 )
Greenhouse
0.0632 ±.0005
0.0651 ±.0004
0.0658 ±.0009
VitalDB
0.0313 ±.0002
0.0312 ±.0001
0.0312 ±.0001
CGMacros
0.0086 ±.0001
0.0095 ±.0001
0.0098 ±.0002
Shanghai
0.1165 ±.0010
0.1174 ±.0010
0.1177 ±.0007
MIMIC-Cardio
0.0543 ±.0001
0.0545 ±.0001
0.0545 ±.0001
Appendix
Table 14: Auxiliary-loss results by dataset. The first two blocks report 16 -step full-model and pre-fusion MAE over seven backbones and five seeds. Rollout error is for steps 49 – 64 , not the cumulative horizon. It uses 273 matched triples: 35 per dataset except VitalDB ( 28 ; all TimeXer runs and two TimeKAN seeds missing). Subscripts are SDs over five seed means of available backbones; VitalDB coverage varies by seed. Bold marks the lowest unrounded mean of each row. Relative changes use seed-averaged cells in the first two blocks and matched backbone–seed triples for rollout.
Backbone dropped
Δ MAE
Δ MAE bypass
λ=0.1
λ=0.3
λ=0.1
λ=0.3
none (full set)
+1.75%
+2.78%
−44.73%
−47.83%
TimeXer
+2.35%
+3.73%
−45.24%
−48.53%
TiDE
+2.16%
+3.37%
−45.90%
−49.20%
DUET
+2.05%
+2.90%
−45.85%
−48.87%
PatchTST
+1.95%
+3.26%
−45.70%
−48.63%
Appendix
Table 15: Sensitivity to backbone inclusion. Each row removes one backbone and recomputes the geometric-mean change relative to λ=0 . Negative values indicate lower MAE. Removing CrossLinear changes the average full-model effect to approximately zero while retaining a large bypass gain, showing where the observed cost is concentrated.
Dataset
Backbone
Observation
AE
VAE
JEPA
Greenhouse
TimeXer
0.0627 ±.0008
0.0604 ±.0006
0.1051 ±.0009
0.0882 ±.0206
TiDE
0.0635 ±.0026
0.0658 ±.0027
0.1094 ±.0007
0.1158 ±.0152
DUET
0.0920 ±.0084
0.0709 ±.0016
0.0962 ±.0134
0.1453 ±.0261
PatchTST
0.0986 ±.0334
0.0924 ±.0334
0.0921 ±.0025
0.1019 ±.0152
TimeKAN
0.0620 ±.0007
0.0642 ±.0011
0.0770 ±.0022
0.0894 ±.0282
CrossLinear
0.1147 ±.0327
0.1066 ±.0024
0.0996 ±.0142
0.0827 ±.0065
Appendix
Table 16: Experiment 1, prediction space: the complete grid. Validation MAE, mean and SD over five training seeds.
Dataset
Backbone
None
Concat
Res
Res-0
FiLM-0
Gate
X-attn
Greenhouse
TimeXer
0.0587 ±.0006
0.0604 ±.0006
0.0625 ±.0019
0.0621 ±.0013
0.0617 ±.0008
0.0608 ±.0012
0.0651 ±.0021
TiDE
0.0609 ±.0005
0.0658 ±.0027
0.0636 ±.0017
0.0635 ±.0020
0.0629 ±.0012
0.0615 ±.0006
0.0641 ±.0011
DUET
0.0657 ±.0006
0.0709 ±.0016
0.0670 ±.0013
0.0678 ±.0013
0.0642 ±.0006
0.0649 ±.0017
0.0694 ±.0014
PatchTST
0.0563 ±.0008
0.0924 ±.0334
0.0606 ±.0013
0.0612 ±.0013
0.0605 ±.0026
0.0590 ±.0014
0.0694 ±.0039
TimeKAN
0.0587 ±.0004
0.0642 ±.0011
0.0656 ±.0020
0.0650 ±.0022
0.0645 ±.0014
0.0636 ±.0017
0.0659 ±.0026
CrossLinear
0.1626 ±.0056
0.1066 ±.0024
0.0731 ±.0009
0.0732 ±.0016
0.0721 ±.0039
0.0703 ±.0019
0.1310 ±.0060
Appendix
Table 17: Experiment 2, plan fusion: the complete grid. The Concat entries reuse the AE results from Table 16 .
Dataset
Backbone
Instant
Decay
Conv
Attn
Greenhouse
TimeXer
0.0608 ±.0012
0.0610 ±.0017
0.0611 ±.0018
0.0608 ±.0010
TiDE
0.0615 ±.0006
0.0614 ±.0007
0.0616 ±.0009
0.0614 ±.0011
DUET
0.0649 ±.0017
0.0657 ±.0011
0.0656 ±.0017
0.0646 ±.0009
PatchTST
0.0590 ±.0014
0.0586 ±.0016
0.0596 ±.0009
0.0589 ±.0018
TimeKAN
0.0636 ±.0017
0.0637 ±.0016
0.0635 ±.0018
0.0634 ±.0015
CrossLinear
0.0703 ±.0019
0.0706 ±.0028
0.0705 ±.0013
0.0713 ±.0007
Appendix
Table 18: Experiment 3, plan encoding with Gate : the complete grid. The Instant entries reuse the Gate results from Table 17 .
Dataset
Backbone
Instant
Decay
Conv
Attn
Greenhouse
TimeXer
0.0617 ±.0008
0.0613 ±.0007
0.0623 ±.0013
0.0618 ±.0011
TiDE
0.0629 ±.0012
0.0629 ±.0012
0.0631 ±.0013
0.0627 ±.0015
DUET
0.0642 ±.0006
0.0642 ±.0014
0.0649 ±.0005
0.0655 ±.0014
PatchTST
0.0605 ±.0026
0.0596 ±.0009
0.0609 ±.0024
0.0598 ±.0010
TimeKAN
0.0645 ±.0014
0.0654 ±.0026
0.0648 ±.0015
0.0637 ±.0013
CrossLinear
0.0721 ±.0039
0.0724 ±.0038
0.0722 ±.0033
0.0720 ±.0029
Appendix
Table 19: Experiment 3, plan encoding with FiLM-0 : the complete grid. The Instant entries reuse the FiLM-0 results from Table 17 ; Table 18 gives the same grid with Gate .
World Action Models (WAMs) enable decision-making through imagined rollouts by predicting future observations and actions. However, the reliability of these imagined futures remains under-examined: is a generated future merely visually plausible, or is it dynamically compatible with the action sequence it claims to model? In this work, we identify action-state consistency, the alignment between predicted actions and induced state transitions, as a missing reliability axis for WAMs. Through a systematic study across representative joint-prediction and inverse-dynamics models, we find that action-state consistency systematically separates successful and failed rollouts across many tasks and follows similar success-failure trends as learned value estimates. These results suggest that consistency captures decision-relevant structure beyond visual realism. We further identify background collapse as an important boundary condition, where low-dynamics failed trajectories can become deceptively consistent because static futures are easier to predict. Building on these findings, we introduce a value-free consensus strategy for test-time selection, which ranks candidate rollouts by agreement among predicted futures. This strategy improves success rates on RoboCasa and RoboTwin 2.0 without additional training or reward modeling. Taken together, our findings establish action-state consistency as both a diagnostic tool for evaluating WAM reliability and a practical signal for value-free planning.
A world model matters to an agent only through the state it constructs. That state must preserve some information, discard other information, and support some future function: prediction, control, planning, memory, grounding, or counterfactual reasoning. This paper treats world-model research as latent state design under sufficiency constraints. We propose a functional taxonomy that groups methods by what their latent state is for, rather than by architecture or application domain: predictive embedding, recurrent belief state, object/causal structure, latent action interface, grounded planning interface, and memory substrate. These roles expose distinctions that architecture-based groupings hide, including the gap between predictive sufficiency and control sufficiency, and the gap between passive video prediction and counterfactual action modeling. The taxonomy supports an evaluation framework that judges a model by the sufficiency constraint its latent state was built to satisfy. We compare methods along seven axes: representation, prediction, planning, controllability, causal/counterfactual support, memory, and uncertainty. We use the resulting matrix as a diagnostic for what a latent state preserves, discards, and enables. The conclusion that follows is that an actionable world model is the one whose state construction matches the task, not the one that preserves the most information.
Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.