Pythia: Toward Foundation World Models for Multimodal Time Series
Organizations: Ant International
Abstract
Time-series foundation models offer a unified approach to forecasting across heterogeneous domains. Textual context and auxiliary observations provide complementary information about temporal dynamics, yet reusable multimodal predictive representations remain underexplored. We introduce Pythia, a foundation world model that learns context-conditioned latent dynamics across datasets through a joint-embedding predictive architecture. A stop-gradient numerical reference guides contextual corrections to predicted future states. A separate probabilistic decoder then adapts to the frozen predictive representation and observed history, decoupling world-model pretraining from observation-space forecasting. On MUSE, Pythia-Tiny's normalized mean absolute scaled error (MASE) and weighted sum quantile loss (WSQL) are 0.6879 and 0.4269, reducing errors by 6.26% and 5.00% relative to the strongest model evaluated in the published MUSE leaderboard. Through a series of controlled experiments, we investigate how to design a time-series world model through shared pretraining and how joint-embedding predictive learning can incorporate multimodal information. The results support separating predictive representation learning from probabilistic readout and show complementary contributions from entity descriptions, events, and covariates.
Figures & tables
| Model | MASE | WSQL | Model | MASE | WSQL |
| Chronos2 | 0.7338 | 0.4493 | TimeMixer | 1.0162 | – |
| TimesFM-2.5-200m | 0.7381 | 0.4601 | CALF | 1.0484 | – |
| Timer-S1 | 0.7419 | 0.4661 | Moirai-1.0-R-large | 1.0559 | 0.6931 |
| Moirai-2.0-R-small | 0.7477 | 0.4715 | Timer-base-84m | 1.0905 | – |
| Sundial-base-128m | 0.7788 | 0.5184 | TimeCMA | 1.1345 | – |
| Toto-2.0-2.5B | 0.7913 | 0.4941 | TimeLLM | 1.1521 | – |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Size | Context layers | Predictor layers | Total | Stage 1 trained | Stage 2 trained |
| Tiny | 1 | 1 | 149.64 | 30.16 | 28.32 |
| Small | 1 | 2 | 159.09 | 39.61 | 28.32 |
| Medium | 1 | 3 | 168.54 | 49.07 | 28.32 |
| Large | 2 | 4 | 185.08 | 65.60 | 28.32 |
| XLarge | 3 | 5 | 201.62 | 82.14 | 28.32 |
| Source family | Stage 1 windows | Stage 2 windows |
| ACL18 | 6,099 | 6,271 |
| CGTSF | 100,686 | 100,266 |
| EPA-Air | 2,414 | 2,414 |
| FNSPID | 2,493,885 | 2,505,451 |
| Fidel-TS | 10,718,993 | 10,743,093 |
| ILINet | 278 | 279 |
| Effective inputs | Pattern | Windows | Share |
| Entity only | 100 | 436 | 3.43% |
| Entity + events | 110 | 1,788 | 14.05% |
| Entity + covariates | 101 | 2,878 | 22.61% |
| Entity + events + covariates | 111 | 7,627 | 59.92% |
| Other four combinations | 000–011 | 0 | 0.00% |
| Total | 12,729 | 100.00% |
| Dataset | Freq. | Term | Tiny | Small | Medium | Large |
| ACL18 stock | 1d | M | 0.989/1.007 | 0.999/0.998 | 0.971/0.934 | 0.982/1.014 |
| ACL18 stock | 1d | S | 1.014/0.959 | 1.133/1.168 | 1.014/0.949 | 1.018/0.933 |
| CGTSF_LEU | 30min | L | 0.711/0.398 | 0.741/0.421 | 0.741/0.427 | 0.711/0.397 |
| CGTSF_LEU | 30min | M | 0.717/0.437 | 0.747/0.462 | 0.745/0.465 | 0.712/0.433 |
| CGTSF_LEU | 30min | S | 0.718/0.603 | 0.737/0.635 | 0.771/0.666 | 0.720/0.604 |
| CGTSF_PTF | 1h | L | 0.511/0.290 | 0.544/0.311 | 0.544/0.315 | 0.540/0.307 |
| Model | Horizon | 80% coverage (%) | Any crossing (%) | Scaled signed span |
| Pythia-Tiny | All | 74.62 | 1.81 | 3.260 |
| Short | 74.47 | 1.26 | 3.072 | |
| Medium | 74.75 | 1.68 | 4.082 | |
| Long | 74.77 | 3.62 | 2.142 | |
| Pythia-Small | All | 71.77 | 1.77 | 3.137 |
| Short | 71.87 | 1.26 | 2.989 |
| Size | Stage 1 + Stage 2 (K) | Batch | MASE | WSQL |
| Tiny | 10 + 10 | 128 | 0.6879 | 0.4269 |
| Tiny | 30 + 30 | 128 | 0.7079 | 0.4448 |
| Tiny | 50 + 50 | 128 | 0.7189 | 0.4460 |
| Small | 10 + 10 | 128 | 0.6999 | 0.4344 |
| Small | 30 + 30 | 128 | 0.6981 | 0.4373 |
| Small | 50 + 50 | 128 | 0.7152 | 0.4486 |
| Size | Updates/stage (K) | |||
| Tiny | 10 | 0.7003/0.4353 | 0.6879/0.4269 | 0.6950/0.4339 |
| Tiny | 30 | 0.7258/0.4476 | 0.7079/0.4448 | 0.7039/0.4372 |
| Small | 10 | 0.7171/0.4406 | 0.6999/0.4344 | 0.7036/0.4376 |
| Small | 30 | 0.7178/0.4423 | 0.6981/0.4373 | 0.7038/0.4374 |
| Small | 50 | – | 0.7152/0.4486 | 0.7207/0.4536 |
| Medium | 30 | 0.7076/0.4385 | 0.7108/0.4459 | 0.7111/0.4398 |
| Variant | MASE | WSQL | MASE |
| Learning recipe | |||
| Pythia-Small | 0.6981 | 0.4373 | +0.00% |
| Direct forecasting | 0.7174 | 0.4442 | +2.77% |
| Joint readout adaptation | 0.7836 | 0.4942 | +12.25% |
| Numerical decoder adaptation | 0.7148 | 0.4419 | +2.39% |
| Numerical-only two-stage | 0.7148 | 0.4464 | +2.40% |
| Variant | MASE | WSQL |
| Pythia-Small | 0.6981 | 0.4373 |
| Pretraining tasks and latent scale | ||
| Forecasting task only | 0.7066 | 0.4392 |
| Masked reconstruction only | 0.6997 | 0.4333 |
| Without task embedding | 0.7294 | 0.4609 |
| Without log-RMS alignment | 0.7845 | 0.5002 |
| Context | Reference guidance | MASE | WSQL |
| None (numerical only) | No | 0.7148 | 0.4464 |
| Entity + events | No | 0.7134 | 0.4439 |
| Covariates | No | 0.7168 | 0.4449 |
| All context | No | 0.7086 | 0.4404 |
| All context | Yes | 0.6981 | 0.4373 |
| Auxiliary objective | MASE | WSQL |
| Numerical-reference guidance (default) | 0.6981 | 0.4373 |
| No reference objective | 0.7086 | 0.4404 |
| Fixed-temperature numerical reference | 0.7127 | 0.4421 |
| Numerical/null/random reference | 0.7047 | 0.4365 |
| Front | Decoder | Selected step (K) | Selected / zero-update | Final |
| F | F | 0 | 1.3271/0.9020 | 1.3271/0.9020 |
| F | A | 30 | 0.7031/0.4337 | 0.7031/0.4337 |
| F | U | 25 | 0.7079/0.4369 | 0.7032/0.4339 |
| A | F | 0 | 1.9822/1.4013 | 1.9822/1.4013 |
| A | A | 30 | 0.8057/0.5046 | 0.8057/0.5046 |
| A | U | 30 | 0.8061/0.5047 | 0.8061/0.5047 |
| Model | Numerical | Corrected | Reduction | MASE | WSQL |
| Pythia-Small | 0.03614 | 0.02291 | 36.59% | 0.6981 | 0.4373 |
| Without reference objective | 0.03891 | 0.02236 | 42.54% | 0.7086 | 0.4404 |
| Model | Normal | Null | Random |
| Pythia-Tiny | 0.6879 | 0.7027 | 0.7315 |
| Pythia-Medium | 0.6943 | 0.6966 | 0.6961 |
| Pythia-Small | 0.6981 | 0.7089 | 0.7160 |
| Small without reference objective | 0.7086 | 0.7063 | 0.7077 |
| Small with masked reconstruction only | 0.6997 | 0.6990 | 0.7004 |
| Dataset | Freq. | Term | Numerical only | Full Pythia | Reduction (%) |
| ACL18 stock | 1d | M | 0.9903/1.0053 | 0.9994/0.9984 | -0.92/+0.68 |
| ACL18 stock | 1d | S | 1.0783/1.0861 | 1.1331/1.1678 | -5.08/-7.52 |
| CGTSF_LEU | 30min | L | 0.7460/0.4264 | 0.7410/0.4215 | +0.68/+1.17 |
| CGTSF_LEU | 30min | M | 0.7503/0.4664 | 0.7475/0.4618 | +0.39/+0.98 |
| CGTSF_LEU | 30min | S | 0.7329/0.6207 | 0.7370/0.6351 | -0.56/-2.33 |
| CGTSF_PTF | 1h | L | 0.5612/0.3204 | 0.5435/0.3110 | +3.16/+2.93 |