Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.
Figures & tables
Figure 1 : Pythia’s default two-stage recipe. Stage 1 learns numerical predictions n from memory h and a residual correction using context memory m , aligning world states w with stop-gradient ( sg ) target states t encoded from full numerical sequences. Matching uses absolute-distance ( L1 ) and log root-mean-square (RMS) terms. A detached numerical reference guides only the context pathway. Stage 2 freezes the predictive representation and adapts three decoder blocks over ordered history, the special REG token, and world states using pinball loss and a parameter anchor; the final decoder block, normalization, and quantile head remain frozen. Targets provide supervision only.
Model
MASE
WSQL
Model
MASE
WSQL
Chronos2
0.7338
0.4493
TimeMixer
1.0162
–
TimesFM-2.5-200m
0.7381
0.4601
CALF
1.0484
–
Timer-S1
0.7419
0.4661
Moirai-1.0-R-large
1.0559
0.6931
Moirai-2.0-R-small
0.7477
0.4715
Timer-base-84m
1.0905
–
Sundial-base-128m
0.7788
0.5184
TimeCMA
1.1345
–
Toto-2.0-2.5B
0.7913
0.4941
TimeLLM
1.1521
–
Table 1: Main MUSE results, retaining all 31 baselines from MUSE Table 3 ( Chen et al., 2026b ) and their published zero-shot or dataset-specific training protocols. Scores are Seasonal-Naive-normalized geometric means; lower is better ( ↓ ). Model-size suffixes m/M and B denote millions and billions of parameters. Baselines are ordered by MASE down the left block, then the right. A dash denotes an unreported probabilistic score. Pythia configurations have different training budgets; their controlled capacity comparison is in Figure 5 .
Figure 2 : MUSE forecasting with Pythia-Tiny and Medium. The MASE panel includes leading numerical foundation models and representative multimodal, LLM-based, dataset-specific, and statistical methods. The WSQL panel includes all published baselines with a reported probabilistic score; the two panels therefore have different rosters. Bars start at zero, dashed lines mark Seasonal Naive, and hatching identifies Pythia. All 31 published baselines remain in Table 1 .
Figure 3 : Predictive pretraining and multimodal context, using Small for matched comparisons. Update counts use K for thousands. (a) Error increases over Pythia under equal 60K total updates: direct forecasting, joint readout training, numerical decoder adaptation (Num. tail), and numerical-only two-stage learning (Num. JEPA). (b) MUSE MASE during readout training after the same Stage 1, with the representation frozen or jointly updated. (c) Error reductions of the full model over entity, event, covariate (Covar.), and reference (Ref.) removal controls. (d) Group-level MASE reductions over the numerical-only and three source/reference controls; diamonds show aggregate reductions. Negative group effects remain visible. Exact scores and both-metric group results appear in Tables 9 and 16 and Figure 16 .
Figure 4 : Transferring Chronos-2 into a multimodal world model. (a,b) All nine Small encoder/decoder policies: frozen (F), anchored adaptation (A), and unanchored adaptation (U). Rows control the numerical encoder in Stage 1; columns control the decoder in Stage 2. Darker cells indicate lower errors within each metric. F decoder columns use zero readout updates; A/U columns use validation-selected checkpoints. (c) Changes to the default Small decoder recipe; LR denotes learning rate. (d) All six adapted-decoder trajectories: color denotes the Stage 1 encoder policy, solid circles/dashed squares the Stage 2 anchor policy. Paired A/U curves nearly coincide. All encoders are frozen in Stage 2. F/U selects 25K, while (d) retains its 30K endpoint; full scores and policy definitions appear in Appendix F.3 .
Figure 5 : Interacting training choices at final checkpoints. (a,b) The complete five-size, three-budget grid at batch 128 per device; circles mark each budget’s minimum. T/S/M/L/XL denote Tiny/Small/Medium/Large/XLarge. (c) Batch effects on Tiny and Medium (Med.); solid/dashed lines use 10K/30K per stage for Tiny and 30K/50K for Medium. (d) Medium at batch 256 with different total updates and Stage 1+Stage 2 allocations; diamonds share 60K total updates. Each declared budget has its own training schedule.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Size
Context layers
Predictor layers
Total
Stage 1 trained
Stage 2 trained
Tiny
1
1
149.64
30.16
28.32
Small
1
2
159.09
39.61
28.32
Medium
1
3
168.54
49.07
28.32
Large
2
4
185.08
65.60
28.32
XLarge
3
5
201.62
82.14
28.32
Appendix
Table 2 : Pythia configurations. Counts are in millions and include retained frozen numerical parameters. The external semantic encoder, training-only reference copy, optimizer states, and anchor buffers are excluded; shared parameters are counted once.
Source family
Stage 1 windows
Stage 2 windows
ACL18
6,099
6,271
CGTSF
100,686
100,266
EPA-Air
2,414
2,414
FNSPID
2,493,885
2,505,451
Fidel-TS
10,718,993
10,743,093
ILINet
278
279
Appendix
Table 3: Training source families and unique eligible window counts, after exclusions and training/validation separation. Context-Guided Time Series Forecasting is abbreviated CGTSF. Public family names link to their sources; internal data are disclosed only as Industrial. Inclusion of a family does not imply inclusion of its MUSE evaluation examples. The two stage-specific pools overlap and must not be added as distinct examples.
Figure 6 : Composition of the eligible pretraining window pools by domain and observation frequency, shown separately for the two stages. S1/S2 denote Stage 1/Stage 2; min, h, d, w, and Mon denote minute, hour, day, week, and month. Shares count unique eligible history origins after exclusions and validation removal, not repeated training draws. Internal sources are combined as Industrial. Small public-domain categories are grouped for readability.
Effective inputs
Pattern
Windows
Share
Entity only
100
436
3.43%
Entity + events
110
1,788
14.05%
Entity + covariates
101
2,878
22.61%
Entity + events + covariates
111
7,627
59.92%
Other four combinations
000–011
0
0.00%
Total
12,729
100.00%
Appendix
Table 4: Observed context availability in MUSE under the normal-input Small preprocessing. Bits denote effective entity/event/covariate inputs in that order. Counts are benchmark window instances; percentages use 12,729 as the denominator. All eight combinations were counted, including the four zero-count patterns.
Figure 7 : MUSE performance across forecast horizons and observation frequencies. Each point is the equal-weight geometric mean of its original normalized scoring groups; parentheses give group counts. Vertical dotted lines mark Seasonal Naive at 1. The same model checkpoint is used throughout. These slices have different task mixtures, so their scores do not isolate horizon or frequency effects on identical data.
Dataset
Freq.
Term
Tiny
Small
Medium
Large
ACL18 stock
1d
M
0.989/1.007
0.999/0.998
0.971/0.934
0.982/1.014
ACL18 stock
1d
S
1.014/0.959
1.133/1.168
1.014/0.949
1.018/0.933
CGTSF_LEU
30min
L
0.711/0.398
0.741/0.421
0.741/0.427
0.711/0.397
CGTSF_LEU
30min
M
0.717/0.437
0.747/0.462
0.745/0.465
0.712/0.433
CGTSF_LEU
30min
S
0.718/0.603
0.737/0.635
0.771/0.666
0.720/0.604
CGTSF_PTF
1h
L
0.511/0.290
0.544/0.311
0.544/0.315
0.540/0.307
Appendix
Table 5: All 29 MUSE scoring groups for the four reported Pythia configurations. Each pair gives normalized MASE / WSQL ( ↓ ); Seasonal Naive is 1 in every group. Freq. denotes observation frequency and Term denotes forecast horizon; S/M/L denote short/medium/long horizons. CGTSF subsets LEU and PTF denote London Electricity Usage and Paris Traffic Flow, respectively. The selected checkpoint is shared across all tasks.
Figure 8 : Two public MUSE forecast windows with Tiny and Medium predictions: hourly transport data from CGTSF-PTF and weekly healthcare data from TimeCAP. Curves show predictive medians and bands span the 0.1–0.9 quantiles, in the original observation scale. The overview strips show the final 96 exported history observations; the larger panels show one forecast horizon of history (48 hours or 12 weeks). Context cards summarize benchmark-supplied inputs retained for the prediction origin. The examples use fixed public configurations and windows selected independently of forecast quality.
Figure 9 : Quantile diagnostics under normal inputs. (a) Empirical decile coverage against nominal probability. (b) Central-80% coverage, with its nominal target. (c) The proportion of targets with any adjacent crossing. (d) Mean signed endpoint span normalized by the MASE scale. Overall and horizon-specific summaries average the original scoring groups equally: 29 overall, 14 short, 10 medium, and five long.
Model
Horizon
80% coverage (%)
Any crossing (%)
Scaled signed span
Pythia-Tiny
All
74.62
1.81
3.260
Short
74.47
1.26
3.072
Medium
74.75
1.68
4.082
Long
74.77
3.62
2.142
Pythia-Small
All
71.77
1.77
3.137
Short
71.87
1.26
2.989
Appendix
Table 6: Group-macro quantile diagnostics. The short, medium, and long subsets contain 253,250, 464,260, and 472,307 valid target points, respectively. Signed span retains any endpoint inversions instead of repairing the predictions.
Size
Stage 1 + Stage 2 (K)
Batch
MASE
WSQL
Tiny
10 + 10
128
0.6879
0.4269
Tiny
30 + 30
128
0.7079
0.4448
Tiny
50 + 50
128
0.7189
0.4460
Small
10 + 10
128
0.6999
0.4344
Small
30 + 30
128
0.6981
0.4373
Small
50 + 50
128
0.7152
0.4486
Appendix
Table 7: Capacity and budget results at final checkpoints. Updates are in thousands; batch size is per device on eight devices. Each run uses its own declared learning-rate horizon. More updates or larger batches change sampled exposure, including repetitions, rather than the number of unique training windows.
Figure 10 : All ten measured batch slices, using final checkpoints. Let sb denote the MASE or WSQL score at per-device batch size b . Cells show percentage changes from batch 128 within the same size and per-stage update budget: 100(sb/s128−1) . Teal denotes improvement and coral degradation, using a common symmetric scale of ±4% . Dashes mark unmeasured settings. Both metrics exhibit interactions with size and duration; larger batches are not uniformly better. Absolute scores are in Table 8 .
Size
Updates/stage (K)
b=96
b=128
b=256
Tiny
10
0.7003/0.4353
0.6879/0.4269
0.6950/0.4339
Tiny
30
0.7258/0.4476
0.7079/0.4448
0.7039/0.4372
Small
10
0.7171/0.4406
0.6999/0.4344
0.7036/0.4376
Small
30
0.7178/0.4423
0.6981/0.4373
0.7038/0.4374
Small
50
–
0.7152/0.4486
0.7207/0.4536
Medium
30
0.7076/0.4385
0.7108/0.4459
0.7111/0.4398
Appendix
Table 8: All measured batch-size slices at final checkpoints. Entries are MASE / WSQL; b is batch size per device on eight devices. Stage 1 and Stage 2 each use the indicated update budget. Dashes indicate unmeasured combinations.
Figure 11 : Allocation and sampled exposure at final checkpoints. (a,b) The full Large Stage 1 × Stage 2 budget grid, batch 128; each panel’s shading ranges from its lowest error (teal) to highest (coral). (c) Medium at batch 256, with MASE above and WSQL below each point; diamonds on the dashed diagonal share 60K total updates. (d) Equal-exposure exchanges between larger batches and more updates: Tiny/Large compare 10K+10K at batch 256 with 20K+20K at batch 128; Medium compares 25K+25K at 256 with 50K+50K at 128. Positive values favor the larger-batch recipe.
Figure 12 : MUSE test trajectories within Tiny and Medium runs at batch 128 per device. Colors identify independent pretraining budgets of 10K, 30K, or 50K updates in each stage; points within a curve are saved Stage 2 checkpoints from that run. Each curve therefore begins from its own pretrained representation and follows its own readout schedule. These curves describe transfer during training; checkpoint selection uses the separate pretraining validation subset.
Figure 13 : Latent representation and readout controls in Small. ID denotes identifier and LR denotes learning rate. (a–c) Percentage error increases relative to Pythia-Small for changes to latent objectives, readout inputs, and decoder adaptation. Axes differ to preserve the scale of the history-removal effect. (d) Final 30K results for a newly trained query decoder using history alone or history plus world states; the dashed horizontal line marks Seasonal Naive. The first three panels use validation-selected checkpoints. Details and absolute scores appear in Table 10 .
Variant
MASE ↓
WSQL ↓
Δ MASE
Learning recipe
Pythia-Small
0.6981
0.4373
+0.00%
Direct forecasting
0.7174
0.4442
+2.77%
Joint readout adaptation
0.7836
0.4942
+12.25%
Numerical decoder adaptation
0.7148
0.4419
+2.39%
Numerical-only two-stage
0.7148
0.4464
+2.40%
Appendix
Table 9: Small-model comparisons on MUSE. Pythia, numerical-only two-stage, and context controls use 30K updates per stage; direct forecasting and numerical decoder adaptation use 60K updates in one stage. The numerical-only model omits context and reference guidance throughout both stages. Joint readout adaptation reuses the same Stage 1 and updates the representation during Stage 2. All rows use validation-selected checkpoints. Δ is relative MASE change from Pythia-Small.
Variant
MASE ↓
WSQL ↓
Pythia-Small
0.6981
0.4373
Pretraining tasks and latent scale
Forecasting task only
0.7066
0.4392
Masked reconstruction only
0.6997
0.4333
Without task embedding
0.7294
0.4609
Without log-RMS alignment
0.7845
0.5002
Appendix
Table 10: Additional Small comparisons with normal inputs, global batch 1,024, and 30K updates per stage. Results use the validation-selected Stage 2 checkpoint: 25K for the history-only query readout and 30K otherwise. The query readouts share the pretrained representation and use a separately trained two-layer query decoder.
Context
Reference guidance
MASE ↓
WSQL ↓
None (numerical only)
No
0.7148
0.4464
Entity + events
No
0.7134
0.4439
Covariates
No
0.7168
0.4449
All context
No
0.7086
0.4404
All context
Yes
0.6981
0.4373
Appendix
Table 11: Context combinations and numerical-reference guidance under the Small 30K+30K protocol. All models are pretrained with their listed context inputs and share the task mixture, numerical initialization, readout policy, batch size, and validation selection rule. These are training controls, not test-time missing-modality interventions.
Auxiliary objective
MASE ↓
WSQL ↓
Numerical-reference guidance (default)
0.6981
0.4373
No reference objective
0.7086
0.4404
Fixed-temperature numerical reference
0.7127
0.4421
Numerical/null/random reference
0.7047
0.4365
Appendix
Table 12: Alternative reference objectives under the same Small training recipe and validation-selected 30K readout checkpoint. All rows retain numerical initialization, contextual inputs, latent matching, and log-RMS alignment.
Front
Decoder
Selected step (K)
Selected / zero-update
Final
F
F
0
1.3271/0.9020
1.3271/0.9020
F
A
30
0.7031/0.4337
0.7031/0.4337
F
U
25
0.7079/0.4369
0.7032/0.4339
A
F
0
1.9822/1.4013
1.9822/1.4013
A
A
30
0.8057/0.5046
0.8057/0.5046
A
U
30
0.8061/0.5047
0.8061/0.5047
Appendix
Table 13: Complete numerical-transfer grid in Small. F: frozen; A: trainable with a parameter anchor; U: trainable without an anchor. The first letter controls numerical front blocks 0–7 in Stage 1; the second controls decoder blocks 8–11 in Stage 2. All adapted blocks use learning-rate multiplier 0.1. The front end is frozen throughout Stage 2, and decoder normalization and the quantile head remain frozen in all nine settings. Stage 1 uses 30K updates; F decoder columns skip Stage 2 and have no validation-selected readout. Other columns train for 30K, with the selected readout step shown. Each score pair is MASE / WSQL.
Figure 14 : Latent correction at fixed Small checkpoints. (a) Group-macro numerical and corrected-state errors. (b) Each point is a scoring group’s window-averaged pair (b,a) ; points below the identity line improve after contextual correction. (c) Distribution of the 29 group-mean reference weights g=sigmoid((a−b)/b) . For the no-reference variant, g is computed as a diagnostic using the same formula.
Model
Numerical bˉ
Corrected aˉ
Reduction
MASE
WSQL
Pythia-Small
0.03614
0.02291
36.59%
0.6981
0.4373
Without reference objective
0.03891
0.02236
42.54%
0.7086
0.4404
Appendix
Table 14: Latent errors and forecasting results for the matched Small models. Let aˉ and bˉ denote arithmetic group-macro corrected and numerical errors, respectively. Latent reduction is 100(1−aˉ/bˉ) . MASE and WSQL retain the benchmark’s geometric aggregation.
Model
Normal
Null
Random
Pythia-Tiny
0.6879
0.7027
0.7315
Pythia-Medium
0.6943
0.6966
0.6961
Pythia-Small
0.6981
0.7089
0.7160
Small without reference objective
0.7086
0.7063
0.7077
Small with masked reconstruction only
0.6997
0.6990
0.7004
Appendix
Table 15: Selected test-time context interventions (MASE). Each row uses the same checkpoint in all three columns. These diagnostics do not affect checkpoint selection.
Figure 15 : Fixed-model responses to joint semantic-embedding replacement, stratified by the context available before intervention; cov. abbreviates covariates. Open and filled markers denote null and random embeddings; colors distinguish Tiny and Small. Within each group–availability slice, s denotes pooled MASE or WSQL, with subscripts indicating intervened/normal inputs; GM denotes the geometric mean. Each value is 100(GM(sintervened/snormal)−1) over nonempty scoring-group slices with that availability pattern. The four patterns contain 5/3/11/15 group slices and 436/2,878/1,788/7,627 windows, respectively; groups can contribute to more than one pattern. Numerical histories, event-time attributes, and covariate summaries retain their original values.
Figure 16 : Relative error reductions of full Pythia-Small over numerical-only two-stage learning (Num. only), no-entity, no-event, and no-reference controls in every MUSE scoring group. Writing sfull and scontrol for the corresponding normalized MASE or WSQL scores in that group, the reduction is 100(1−sfull/scontrol) . Teal favors the full model and coral favors the control; both panels use the same symmetric color scale. S/M/L denote short/medium/long horizons. All controls use matched Small recipes and selected checkpoints. Negative cells remain visible, showing that aggregate improvements need not hold for every task.
Dataset
Freq.
Term
Numerical only
Full Pythia
Reduction (%)
ACL18 stock
1d
M
0.9903/1.0053
0.9994/0.9984
-0.92/+0.68
ACL18 stock
1d
S
1.0783/1.0861
1.1331/1.1678
-5.08/-7.52
CGTSF_LEU
30min
L
0.7460/0.4264
0.7410/0.4215
+0.68/+1.17
CGTSF_LEU
30min
M
0.7503/0.4664
0.7475/0.4618
+0.39/+0.98
CGTSF_LEU
30min
S
0.7329/0.6207
0.7370/0.6351
-0.56/-2.33
CGTSF_PTF
1h
L
0.5612/0.3204
0.5435/0.3110
+3.16/+2.93
Appendix
Table 16: Paired Small results for numerical-only two-stage learning and full Pythia. Both use the selected 30K readout checkpoint after 30K latent-pretraining updates, with matching coverage in every group. Score pairs are normalized MASE / WSQL; reduction pairs are percentages, positive when full Pythia improves. These are training-time controls, not test-time removal of context.
Real-world time series come with text: metadata, descriptions, news, reports. Yet time series foundation models process numerical sequences in isolation, and the multimodal text-and-time-series models that attempt to bridge the two all adapt a pretrained language model post hoc, inheriting representations shaped without ever seeing temporal data. These models are also evaluated almost exclusively against other multimodal baselines, not against the strongest unimodal foundation models in either domain, leaving open whether joint training is needed at all. We present Chronicle, a compact 324M-parameter decoder-only transformer trained from scratch on natural language and time series within a single unified architecture. Both modalities share the same transformer blocks, attention mechanism, and residual stream; the bulk of pretraining uses unimodal batches so cross-modal capability emerges purely from shared parameters, with a short alignment stage that interleaves the two. To our knowledge, Chronicle is the first model jointly pretrained on text and time series from scratch, and the first multimodal model evaluated against dedicated foundation models in both domains. It matches Gemma-3-270M-PT on 19 NLU tasks, sets a new bar for frozen-embedding time series classification on 24 UCR/UEA datasets, and produces multimodal forecasts on Time-MMD that beat every supervised fusion baseline, all from a single backbone.
Paul Quinlan, Jeremy Levasseur, Qingguo Li +1
1InertialAI · Department of Electrical and Computer Engineering, Queen’s University · Department of Mechanical and Materials Engineering, Queen’s University
Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two types of modalities separately. For endogenous modalities, to capture how they evolve along with the underlying temporal dynamics, we introduce an Endo-Multimodal Predictor and Endo-Multimodal Supervision to explicitly learn their evolution from history to the future. For exogenous modalities, to generalize across domains and across various modality types and numbers under the scarcity of exo-multimodal pretraining data, we propose an Exo-Multimodal Retrieval Enhancer that enables rapid downstream adaptation without updating the TSFM parameters. We further introduce Endo-Modality Proxy Training to train this retrieval module without exogenous multimodal pretraining data. Extensive experiments across unimodal and multimodal benchmarks demonstrate strong forecasting performance in scenarios both with and without exogenous modalities.
Hanyin Cheng, Linfeng Wang, Zhengbo Qu +7
East China Normal University · Huawei Technologies Co., Ltd.
Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-numerical context that often shape real-world trajectories. In this work, we bridge this gap and argue for a multimodal time-series forecasting approach that post-trains LLMs to act as context-guided revisors over strong numerical TSFM priors. We introduce PostTime, a post-training recipe combining Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR), along with a methodology to generate automated reasoning traces for forecast revisions. PostTime teaches an LLM to generate context-conditioned forecast interventions -- decisions to revise, preserve, or ignore the TSFM prior based on the multimodal context. We evaluate this approach on the TimesX multimodal forecasting benchmark using a Gemma-3-4B LLM and TimesFM-2.5 TSFM, and show that it significantly outperforms standalone TSFMs, LLM-only baselines, and existing multimodal forecasting approaches.