We introduce Timer-M1, a pretrained multivariate time series foundation model that learns with primitives for zero-shot forecasting. Across domains, time series share elementary temporal and relational patterns, termed primitives, yet differ in how these primitives manifest and evolve across different contexts. Despite progress in zero-shot and task-general forecasting, existing foundation models may still struggle to generalize to complex real-world scenarios. To this end, we develop a primitive-based data synthesis and pretraining pipeline. The synthesis pipeline generates series with temporal primitives shared across domains and then assembles real and generated series into multivariate samples using relational primitives. Afterwards, samples are organized into episodes by assigning distinct channel roles as target variates, past-only covariates, and known-future covariates, ensuring that the model is optimized on predictable variates using available exogenous information. Technically, Timer-M1 further adapts gated two-dimensional Transformer blocks that dynamically allocate cross-variate attention across layers. Across three large-scale forecasting benchmarks, Timer-M1 ranks first on both FEV and TIME and second on GIFT-Eval among most recent time series foundation models. These results support effective primitive-based pretraining as a route to robust general forecasting technique across domains and task settings.
Figures & tables
Figure 1: Learning with primitives. Temporal and relational patterns observed in real series are composed into diverse samples; episodes vary what is observed and predicted while introducing noise and interference to improve robustness, linking to generalization to new forecasting scenarios.
Figure 2: Timer-M1 overview. (a) Temporal and relational primitives combine real and synthetic sources into multivariate samples through synthesis and augmentation. (b) Forecasting episodes specify variate roles and observation availability, with masks controlling forecasting and within-episode attention. (c) Gated two-dimensional Transformer blocks integrate temporal and cross-variate information, with learned layer-wise gates regulating cross-variate updates.
Figure 3: Persisted records and scalar positions by generation route, before sampling. Real–synthetic joint generation includes corpus-derived and mixed-source generation routes.
Figure 4: Evaluation overview. Public benchmarks assess forecasting across settings, while primitive evaluations examine temporal extrapolation, relational forecasting, and input reliability. Rankings are among compared methods.
Figure 5: Forecasting results on FEV across 100 matched configurations. Top: aggregate MASE and SQL for 12 methods, normalized by Seasonal Naive. Bottom: SQL across five task types with different variate roles and observation availability. Timer-M1 achieves the lowest aggregate errors. Colors denote forecasting interfaces above and individual models below.
Figure 6: Comparison on TIME across 98 matched configurations. Timer-M1 achieves state-of-the-art aggregate performance, attaining the lowest normalized MASE and CRPS among the compared methods across diverse forecasting task settings. Colors distinguish forecasting interfaces.
Figure 7: Point and probabilistic forecasting performance on GIFT-Eval across 97 matched configurations. Timer-M1 ranks second among 12 compared pretrained models on both MASE and CRPS, demonstrating competitive generalization across diverse domains and forecast horizons.
Figure 8: Primitive evaluation. Top: predefined temporal groups, seven relational mechanisms, and six reliability categories, measured by MSE (lower is better). Unsupported interfaces are omitted; detailed results for the displayed groups appear in Appendix C.3 . Bottom: six strength cases spanning distinct mechanisms, not an estimate of their frequency. Shading denotes M1’s marginal 10–90% quantiles. † marks unavailable squared-error archives for some baseline.
Figure 9
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Primitive
Representative construction
Temporal
Periodicity
Oscillatory components and their superposition
Trend
Gradual growth, decline, and shared drift
Stochastic dependence
Noise and autoregressive dynamics
Regime changes
Piecewise changes in level or evolution
Relational
Signed responses
Positive and negative source mixing
Delayed responses
Temporal offsets between source and response
Appendix
Table 2: Operator families in primitive-based data construction. The taxonomy follows the temporal and relational primitives introduced in the main text. Examples describe mechanisms rather than a fixed generation recipe.
Component
Configuration
Parameters
120,998,412 (approximately 121M)
Transformer layers
12
Representation / feed-forward width
768 / 3072
Attention heads / head dimension
12 / 64
Patch length
16 observations
Prediction quantiles
0.1,0.2,…,0.9
Appendix
Table 3: Model configuration. The gate contributes one learned scalar per layer; the total parameter count includes the embedding and prediction head.
Figure 10: Layer-wise allocation of cross-variate interaction. The learned coefficients generally rise with depth, while local deviations allow different layers to assign different weights to variate-attention updates. The dashed line marks the earlier stable-model mean used by the fixed- 0.202 control, not the mean of the displayed curve. Together with the gate ablation, the profile motivates adaptive allocation rather than a uniform residual scale.
Capability
Group
Controlled pattern or observation condition
Temporal
Periodic
Period, waveform, modulation, and gradual frequency change
Trend
Direction, growth shape, and observed slope change
Shift
Observed level change and unannounced boundary changes
Relational
Sign; scale
Shared dynamics with opposite signs or different scales
Lag; phase
Delayed responses and phase-shifted periodic targets
Cointegration
Multiple responses sharing a nonstationary source
Appendix
Table 4: Displayed primitive evaluation groups and their controlled conditions. The counts of scoring records and model-specific results appear in Tables 8 – 10 . All groups report MSE; NMAE additionally supports the comparisons discussed in the text.
Model
FEV MASE
FEV SQL
TIME MASE
TIME CRPS
GIFT MASE
GIFT CRPS
Timer-M1
0.3771
0.4887
0.6372
0.5363
0.6806
0.4688
TimesFM-3
0.3742
0.4866
0.6398
0.5363
0.6668
0.4557
Chronos-2
0.3550
0.4728
0.6620
0.5563
0.6978
0.4854
TiRex-2
0.3374
0.4550
–
–
0.6973
0.4781
Toto 2.0 (2.5B)
0.3254
0.4442
0.6419
0.5394
0.6956
0.4759
TimesFM-2.5
0.3562
0.4668
0.6686
0.5674
0.7050
0.4903
Appendix
Table 5: Aggregate benchmark results. FEV reports skill ( ↑ ); TIME and GIFT-Eval report normalized errors ( ↓ ). Red bold and blue underline mark the best and second-best displayed scores per column. Dashes indicate unavailable matched results. Training exposure differs across methods (Appendix B.4 ).
Task type
n
Timer-M1
TimesFM-3
Chronos-2
TiRex-2
Toto 2.0
Univariate
32
0.2968
0.2998
0.2799
0.2629
0.2653
Multivariate
26
0.4266
0.4252
0.4017
0.3889
0.4366
Past-only
12
0.4256
0.3933
0.3655
0.3969
0.3426
Known-future
18
0.4481
0.4549
0.4323
0.3979
0.3332
Known + past
12
0.2988
0.2923
0.3029
0.2465
0.1710
Appendix
Table 6: FEV MASE skill by task type. Groups are mutually exclusive and include all 100 configurations; higher is better. Best and second-best values are marked within each row.
Task type
n
Timer-M1
TimesFM-3
Chronos-2
TiRex-2
Toto 2.0
Univariate
32
0.3828
0.3847
0.3703
0.3538
0.3563
Multivariate
26
0.5954
0.5953
0.5795
0.5668
0.6033
Past-only
12
0.4668
0.4382
0.4239
0.4387
0.3950
Known-future
18
0.5363
0.5458
0.5257
0.4971
0.4344
Known + past
12
0.4295
0.4171
0.4254
0.3765
0.3022
Appendix
Table 7: FEV SQL skill by task type. Groups are mutually exclusive and include all 100 configurations; higher is better. Best and second-best values are marked within each row.
Group
n
Timer-M1
TimesFM-3
Chronos-2
TiRex-2
Toto 2.0
Shift
6
6.4278
7.4830
6.8490
6.9092
6.4418
Periodic
13
0.0271
0.6803
0.0919
0.0735
0.2300
Trend
10
0.1312
0.0299
0.4522
0.4512
0.0083
Appendix
Table 8: Temporal extrapolation groups corresponding to Figure 8 . Mean squared error is averaged over the complete records of each displayed group; lower is better.
Group
n
Timer-M1
TimesFM-3
Chronos-2
TiRex-2
Toto 2.0
Sign
2
0.0021
0.0036
0.0077
0.0100
0.0073
Cointegration
2
0.2477
0.3953
2.2627
1.3930
4.7848
Promotion
2
0.3008
0.3042
1.0424
0.5567
–
Lag
2
0.0012
0.0005
0.0185
0.0074
0.0049
Phase
2
0.000053
0.000043
0.0002
0.0087
0.0010
Scale
2
0.0688
0.0955
0.8690
2.2971
4.7189
Appendix
Table 9: Relational forecasting groups corresponding to Figure 8 . Mean squared error is averaged over the complete records of each displayed group; lower is better.
Group
n
Timer-M1
TimesFM-3
Chronos-2
TiRex-2
Toto 2.0
Irrelevant context
7
0.3681
0.2588
0.2554
0.2585
0.2451
Irrelevant targets
3
0.0035
–
0.0167
0.0633
0.0055
Missing event inputs
7
5.7917
5.7594
3.9130
2.2120
–
Target corruption
21
12.0938
2.6081
23.9802
0.2260
3.8370
Target missingness
21
0.1957
0.1477
0.7231
0.2184
0.3595
Variate order
7
0.0034
–
0.0221
0.0623
0.0055
Appendix
Table 10: Input reliability groups corresponding to Figure 8 . Mean squared error is averaged over the complete records of each displayed group; lower is better.
Gate
Overall
Univariate
Multivariate
Past-only
Known-future
Known + past
MASE Skill ↑
Adaptive
0.3771
0.2968
0.4266
0.4256
0.4481
0.2988
Fixed 0.202
0.3714
0.2969
0.4299
0.4266
0.4084
0.3081
Fixed 1.0
0.3680
0.2963
0.4260
0.4238
0.3979
0.3104
SQL Skill ↑
Adaptive
0.4887
0.3828
0.5954
0.4668
0.5363
0.4295
Appendix
Table 11: Gate ablation on FEV overall and five task types. Both metrics report skill scores. Red bold and underline indicate the best and second-best results within each column and metric. Adaptive gates achieve the highest overall skill, with the clearest advantage in known-future forecasting.
Recipe
FEV MASE ↑
FEV SQL ↑
TIME MASE ↓
TIME CRPS ↓
GIFT MASE ↓
GIFT CRPS ↓
Real Only
0.3196
0.4431
0.6579
0.5526
0.6998
0.4850
Synthetic Only
0.3568
0.4710
0.6438
0.5456
0.6884
0.4741
No-Joint
0.3578
0.4725
0.6379
0.5359
0.6819
0.4722
Full Recipe
0.3771
0.4887
0.6372
0.5363
0.6806
0.4688
Appendix
Table 12: Data-construction ablations across three benchmarks. Recipe names refer to the source-generation routes under comparison; shared benchmark-associated training data are unchanged across variants.
Figure 11: Temporal extrapolation across periodicity, modulation, frequency drift, and observed regime changes. Lower strips show Timer-M1’s absolute forecast error. Each panel uses its own value scale.
Figure 12: Relational forecasting with shared, signed, delayed, phase-shifted, cointegrated, and event-driven dynamics. Lower strips show available related inputs or target histories; they are not additional forecast targets unless specified by the task.
Figure 13: Input reliability under missing histories, imperfect event information, irrelevant inputs, and historical outliers. Lower strips show Timer-M1’s absolute forecast error.