We develop a Unified Scaling Law and a Unified Theory of Time Series Learning to understand how model capacity and historical information support forecasting. Across different lookback lengths and forecast horizons, we analyze 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks spanning six domains. Our empirical methodology integrates local resource relations into a parsimonious, fitted five-parameter law: capacity gains increase with history, context gains diminish toward saturation, and horizon effects enter as a common shift. Fitted without Toto 2.0, the law predicts its horizon-averaged capacity-scaling curves with mean absolute percentage errors of 1.09% and 1.50% at input lengths 2048 and 4096. To understand how history supports prediction, our learning theory uses Gaussian regression to analyze rule identification and predictive capability. We hypothesize that full-shot models learn by accumulating information in weights, while frozen time series foundation models (TSFMs) use history by extracting information through activations. Matched-history comparisons establish the predictive value of additional history. Controlled parameter exchanges and activation interventions provide evidence that history-derived rule information can be retained, reused across queries, and used to recover a contribution to long-context prediction. Together, these findings inform capacity scaling, context allocation, and the development of models that retain and apply historical rules. Code and main results are available at https://github.com/Fifthky/UniScale.
Figures & tables
Source
TSFM Checkpoints
Dataset-Frequency Task
Experimental Cells
Points
Controlled GIFT-Horizon
21
23
3,128
390
Controlled Grid-Horizon
21
23
15,640
680
Combined Controlled
21
23
18,768
1,070
GIFT Public
37
23
2,257
199
Table 1: Grid-first selection retains dataset-level results (cells); points aggregate responses for fitting and prediction. Controlled GIFT-Horizon contains only cells absent from Controlled Grid-Horizon.
Figure 1: The Unified Scaling Law and observed resource profiles. (a) Predicted MASE over N,L at H=48,192,720 (blue, green, orange). (b)–(d) Observed means (markers) and mean predictions (lines) over identical capacity, input-length, and horizon groups, using all 1,070 Controlled points. Resource axes are logarithmic; R is linear in (a) and logarithmic in (b)–(d). 1k denotes 1,024 steps.
Figure 2: Observed resource relations and reduced-law restrictions: (a) capacity response, (b) changing context returns, (c) history-dependent capacity gains, and (d) context elasticity for the representative 22M group. Unified curves use fitted coefficients; orange dashed curves impose restrictions without refitting. Faint points show individual Controlled contrasts; emphasized markers show medians, with interquartile-range (IQR) bars where shown. Appendix C.4 specifies each comparison.
Figure 3: Controlled/Public prediction and horizon-averaged Toto 2.0 capacity curves.
Figure 4: Forecasting with matched history budgets on 23 tasks. Full-shot curves average three seeds; TSFM curves show the mean and median over a fixed set of available checkpoints covering the grid. Seed error bars are too small to distinguish at this scale: mean pointwise relative standard deviations are ±0.208% for DLinear and ±0.502% for PatchTST. The line marks Seasonal Naive.
Process
Query only
Transfer
Natural
Erased
Recovered
AR(1)
0.26289
0.24181
0.17697
0.23986
0.21229
Lag-8
0.26972
0.25106
0.19702
0.25074
0.23755
Threshold
0.30168
0.27873
0.24983
0.28404
0.26195
Table 2: Shared-site transfer and recovery in TimesFM 2.5. Conditional-mean MASE, averaged over three rule magnitudes and two donor repetitions. Natural, erased, and recovered predictions share the 512-step recipient history.
Figure 5: Context gains and fitted saturation scales. Left: mean and median MASE improvements between adjacent inputs for matched checkpoint–horizon pairs. Right: fitted logR at H=96 for three capacities; L⋆ markers locate approximate saturation regions within the evaluated range.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Sampling intervals
Domain
Bitbrains fast storage
5 min
Cloud/web operations
Bitbrains random
5 min
Cloud/web operations
BizITObs L2C
5 min, 1 h
Cloud/web operations
Electricity
15 min, 1 h
Energy
ETT1
15 min, 1 h
Energy
ETT2
15 min, 1 h
Energy
Appendix
Table 3: The fixed evaluation panel and its application domains. Sampling intervals span five, ten, and fifteen minutes, one hour, one day, and one week. Nineteen pairs have short, medium, and long GIFT configurations; the four daily/weekly pairs contribute four short configurations, for 61 in total.
Checkpoint
Active N
Total N
Checkpoint
Active N
Total N
Toto 2.0 4M
4.14
4.14
Moirai 1.1 large
310.97
310.97
Toto 2.0 22M
21.92
21.92
Moirai 2.0 small
11.39
11.39
Toto 2.0 313M
312.68
312.68
TimesFM 1.0
200.00
200.00
Toto 2.0 1B
1041.03
1041.03
TimesFM 2.0
500.00
500.00
Toto 2.0 2.5B
2454.28
2454.28
TimesFM 2.5
200.00
231.29
Chronos 2 small
27.93
27.93
TiRex 1.1 GIFT
35.00
35.00
Appendix
Table 4: Controlled checkpoints and active and total parameter counts, in millions.
Figure 6: Resource coverage across active capacity N , input length L , and scored horizon H (colors and markers). Public coordinates are grid-binned; Controlled coordinates use data-supported inputs. The Controlled GIFT-Horizon Panel shows 662 protocol points and contributes 390 additional points after Grid-first selection (Table 1 ). Small horizontal offsets separate overlapping horizons.
Form
Controlled
Public
Toto
Constant
14.83
16.92
17.10
Capacity-independent
9.96
13.01
11.84
Unified
10.13
12.82
11.53
Appendix
Table 5: Reference prediction MAPEs (%) on Controlled, Public, and held-out Toto response points.
Figure 7: Complete Toto 2.0 evaluation. (a,b) Observed versus predicted logR at all 150 original NHL coordinates for the all-checkpoint and leave-Toto-out fits; color encodes active parameters and the diagonal denotes exact prediction. (c,d) Capacity response after pairing adjacent input lengths and averaging the five horizons. Solid circles are observed Toto means and dashed crosses are predictions from the corresponding fit. Each context curve contains all five Toto checkpoints.
Model
Input C
History T=4096
History T=8192
DLinear
96
1.04264
1.03958
DLinear
512
0.88869
0.88328
PatchTST
96
0.91400
0.91361
PatchTST
512
0.86638
0.84896
PatchTST-compact
96
0.99099
0.98986
PatchTST-compact
512
0.90357
0.89159
Appendix
Table 6: Validation-selected MASE in the crossed experiment, averaged over 23 tasks and three seeds. T is training history and C is immediate input.
Process
Frozen model
L=96
L=2048
Rule projection
AR(1)
TimesFM 2.5
0.32424
0.12650
0.812
AR(1)
Chronos 2
0.33295
0.12958
0.666
Lag-8
TimesFM 2.5
0.33272
0.16004
0.739
Lag-8
Chronos 2
0.34423
0.26440
0.133
Threshold
TimesFM 2.5
0.37300
0.22949
0.454
Threshold
Chronos 2
0.37281
0.25215
0.369
Appendix
Table 7: Conditional-mean MASE with fixed population statistics at a=0.3 . Rule projection measures the predicted contrast along the known conditional-mean contrast; one corresponds to the oracle amplitude.
Process
Model
L=1024
L=4096
L=8192
Transfer projection
AR(1)
DLinear
0.27871
0.16235
0.12634
0.986
AR(1)
PatchTST
0.32585
0.14790
0.09567
1.028
Lag-8
DLinear
0.32463
0.17748
0.12204
0.985
Lag-8
PatchTST
0.34635
0.14133
0.10034
0.946
Threshold
DLinear
0.35156
0.25287
0.22841
0.584
Threshold
PatchTST
0.37115
0.20446
0.18716
0.755
Appendix
Table 8: Conditional-mean MASE of trained predictors, averaged over three seeds. Transfer projection measures the donor-rule response after parameter exchange and shared-history normalization calibration at L=8192 .
Model
Objective
Baseline
Matched
Opposite
Random
Shuffled
TimesFM 2.5
Transfer
0.27810
0.25720
0.30316
0.31468
0.26804
TimesFM 2.5
Recovery
0.25821
0.23726
0.29228
0.29514
0.25137
Chronos 2
Transfer
0.28774
0.28776
0.29987
0.28710
0.29144
Chronos 2
Recovery
0.21064
0.19798
0.20436
0.20984
0.20454
Appendix
Table 9: Conditional-mean MASE for the primary TimesFM setting and independently confirmed Chronos setting. Baseline is query-only prediction for transfer and erased prediction for recovery. Matched and opposite refer to the donor rule.
Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more history into the model improves forecast quality. But in stochastic domains, distant history is often just high-frequency noise, not signal. Hence, the proposed work tests whether this premise actually holds by running continuous context architectures (PatchTST included) through the ETTh1 benchmark. The obtained results contradict the premise: an inverse scaling law shows up clearly, with forecasting error rising as context gets longer. A 3,000-step window causes performance to drop by over 68%, evidence that attention mechanisms are poor at ignoring irrelevant historical volatility. Retrieval-Augmented Forecasting (RAFT) is evaluated as an alternative. RAFT achieves a mean squared error (MSE) of 0.379 with a fixed 720-step window and selective retrieval, outperforming both long-context configurations and zero-shot foundation models (Chronos, Moirai) despite requiring far less computation. In addition, the retrieval step injects only the most relevant historical segments as dynamic exogenous variables, which gives the model a context-informed inductive bias it cannot build on its own from raw sequences. Therefore, foundation models going forward need to shift architecturally toward selective retrieval.
Rishi Ahuja, Kumar Prateek, Simranjit Singh +1
Department of Information Technology, Dr. B.R. Ambedkar National Institute of Technology Jalandhar, Punjab, 144008, India.
Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers to general-purpose foundation models, on the assumption that capacity is what unlocks accuracy. We take the opposite position: most of the gap can be closed at far lower cost by tuning preprocessing rather than scaling models. We use Ridge regression as the testbed, since it has a closed-form solution and interpretable weights, which let the optimal hyperparameters be read off the search directly. We search over context length, local normalization, regularization, and augmentation on eight standard benchmarks and find three patterns. (1) Optimal lookback is strongly series-specific and often non-monotonic in forecast horizon, with fitted power-law exponents ranging from +0.46 on ETTm2 to −0.19 on Exchange and Traffic, challenging the convention that longer horizons need longer history. (2) Normalizing over a learned trailing fraction of the context, rather than its entirety, is almost universally preferred. (3) Series within the same dataset often disagree on hyperparameters; the optimal degree of cross-series sharing varies from fully shared to fully per-series. The resulting models beat prior linear forecasters on most dataset-horizon entries and exceed Transformer, MLP, and CNN baselines on six of eight benchmarks. The optimized hyperparameters also serve as a diagnostic on the data itself, revealing structures that larger models absorb silently into their learned parameters. We provide an accompanying interactive online demonstration and the code at https://sakanaai.github.io/SearchCast/.
Lang Huang, Jinglue Xu, Luke Darlow
1Sakana AI, Tokyo, Japan · 2National Institute of Informatics, Japan
This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.
Kang Yang, Gaofeng Dong, Liying Han +1
Department of Electrical and Computer Engineering University of California, Los Angeles