Longer histories can improve time-series foundation models (TSFMs), but require substantially higher inference cost. We therefore ask whether contextual information can be provided more efficiently through a compact set of learned token embeddings. We introduce PaCTS, which generates a small set of instance-adaptive latent prompts in the form of continuous embedding tokens conditioned on the visible context. These prompts serve as compact context surrogates for frozen TSFMs. PaCTS constructs them from instance-specific global statistics and further refines them with segment-level temporal information, capturing both global characteristics and local temporal variations. The prompt module is jointly trained and deployed across heterogeneous time series with the frozen backbone. Extensive experiments demonstrate the effectiveness of prompts as context, consistently improving forecasting across context lengths and model architectures. With a shorter input context, PaCTS can outperform the same frozen backbone using double context while requiring substantially less inference computation. Compared with weight-space adaptation methods, PaCTS achieves stronger improvements and better out-of-distribution generalization.
Figures & tables
Figure 1: PaCTS improves forecasting at lower additional cost than extending context.
Figure 2: Overview of PaCTS. (a) A lightweight prompt generation module maps the context to instance-adaptive latent prompts: continuous vectors prepended to the time-series embeddings. (b) Global context statistics determine a weighted combination of learnable prompt basis patterns, yielding the adaptive component Pa(x) . Segment-level statistics are encoded with relative-position embeddings and retrieved through cross-attention to produce a refinement R(x) . The adaptive component and its refinement are combined, gated, and added to the shared component Ps to form the final prompts.
Context length 4,096
Context length 8,192
Method
MASE ↓
Δ (%)
CRPS ↓
Δ (%)
MASE ↓
Δ (%)
CRPS ↓
Δ (%)
Chronos-2 zero-shot
0.708
—
0.496
—
0.704
—
0.491
—
Linear probing
0.710
+0.28
0.491
-1.01
0.705
+0.14
0.485
-1.22
BitFit
0.708
0.00
0.489
-1.41
0.704
0.00
0.485
-1.22
LayerNorm-only
0.705
-0.42
0.488
-1.61
0.701
-0.43
0.483
-1.63
Full fine-tuning
0.703
-0.71
0.487
-1.81
0.698
-0.85
0.482
-1.83
Table 1: Forecasting performance on GIFT-Eval across 97 configurations under the univariate protocol based on Chronos-2. Adaptation results are averaged over three seeds. PaCTS achieves best performance across both context lengths and metrics, even lower MASE and CRPS than the zero-shot and five adaptation baselines with twice the context length (PaCTS 4,096 vs. others 8,192).
Method (context)
MASE ↓
CRPS ↓
GFLOPs ↓
Memory (MiB) ↓
Latency (ms) ↓
Params (M)
Zero-shot (4,096)
0.708
0.496
1,018.4
651.3
40.4
119.478
Zero-shot (8,192)
0.704
0.491
2,087.4
809.4
75.0
119.478
PaCTS (4,096)
0.697
0.480
1,057.9
660.7
47.4
119.631
Δ vs. zero-shot (8,192)
−0.99%
−2.24%
−49.32%
−18.37%
−36.80%
+0.128%
Table 2: Forecasting performance and inference costs of Chronos-2 on GIFT-Eval. Compared with the pretrained model at doubled context length, our method achieves better forecasting performance at lower inference cost, with only 0.128% additional parameters, demonstrating the effectiveness of latent prompts as context surrogates.
Variant
MASE ↓
CRPS ↓
Zero-shot
0.708
0.496
Static prompts
0.707
0.490
+ Adaptive generation
0.704
0.486
+ Segment refinement
0.697
0.480
Table 3: Ablations on prompt generation. Prompt module is tuned based on Chronos-2 at 4,096 steps. Both adaptive generation and segment refinement improves forecasting.
Figure 3: Effects of history length and forecast horizon. The bars show percentage changes in MASE and CRPS relative to zero-shot Chronos-2 at 4,096 steps. Negative values indicate improvements. The left two panels group configurations by history lengths, and the right two by forecasting horizons. PaCTS performs consistently better than other methods and double-context zero-shot forecasts. Detailed numbers are in Appendix D .
Figure 4: In-distribution and out-of-distribution generalization performance comparisons. Relative changes in MASE and CRPS are measured against zero-shot Chronos-2 at 4,096 steps. Among the compared methods, PaCTS achieves the best in-distribution performance on GIFT-Eval and is the only method that improves both metrics over zero-shot inference on the TIME benchmark.
Table 4: Extension to multivariate forecasting with univariate-trained prompts. PaCTS is trained only on the GIFT-Eval training split under the univariate setting and evaluated directly using the official multivariate inference protocol of Chronos-2, without additional training or tuning. GIFT-Eval and TIME serve as in-distribution and out-of-distribution benchmarks, respectively. The univariate-trained prompts improve both metrics on both benchmarks.
Figure 5: Forecasting performance (top) and computational costs (bottom) across backbones on GIFT-Eval. Lower values are better for all metrics. As with Chronos-2, shorter-context PaCTS achieves competitive or better performance than longer-context zero-shot baselines, with substantially lower computational costs.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
No.
Feature
Implementation
1
Mean
Mean of observed values
2
Standard deviation
Clipped above at 10
3
Minimum
Clipped to [−10,10]
4
Maximum
Clipped to [−10,10]
5
Half-window change
Second-half mean minus first-half mean; [−5,5]
6
Mean first difference
Consecutive observed pairs; [−5,5]
Appendix
Table 5: Ten input statistics used by the prompt generator.
Method
Updated parameters
Learning rate
Full fine-tuning
All backbone parameters
10−6
LoRA
Rank 8, α=16 ; attention and output layers
10−5
LayerNorm-only
Layer-normalization parameters
10−3
BitFit
Selected bias parameters
10−3
Linear probe
Output patch embedding
10−3
Appendix
Table 6: Chronos-2 adaptation baseline settings. Every row uses the official fit pipeline, 5,000 optimizer steps, and batch size 256.
GIFT-Eval
TIME
Method
Context steps
MASE
Δ (%)
CRPS
Δ (%)
MASE
Δ (%)
CRPS
Δ (%)
Zero-shot
4,096
0.708
-
0.496
-
0.664
-
0.558
-
Zero-shot
8,192
0.704
-0.56
0.491
-1.01
0.666
+0.30
0.559
+0.18
Full fine-tuning
4,096
0.703
-0.71
0.487
-1.81
0.665
+0.15
0.560
+0.36
LoRA
4,096
0.703
-0.71
0.493
-0.60
0.666
+0.30
0.560
+0.36
LayerNorm-only
4,096
0.705
-0.42
0.488
-1.61
0.666
+0.30
0.559
+0.18
Appendix
Table 7: Detailed results of in-distribution and out-of-distribution generalization comparisons. PaCTS achieves the best in-distribution performance on GIFT-Eval and is the only method that improves both metrics over zero-shot inference on the out-of-distribution TIME benchmark.
History <4,096 ( n=38 )
History ≥4,096 ( n=59 )
Method
Context steps
MASE
CRPS
MASE
CRPS
Zero-shot
4,096
0.700
0.533
0.712
0.474
Zero-shot
8,192
0.700
0.533
0.706
0.466
LoRA
4,096
0.690
0.524
0.711
0.474
Full fine-tuning
4,096
0.690
0.519
0.711
0.468
PaCTS
4,096
0.685
0.511
0.705
0.461
Appendix
Table 8: Detailed comparisons of different history lengths on GIFT-Eval.
Short ( n=55 )
Medium ( n=21 )
Long ( n=21 )
Method
Context steps
MASE
CRPS
MASE
CRPS
MASE
CRPS
Zero-shot
4,096
0.676
0.506
0.740
0.485
0.763
0.483
Zero-shot
8,192
0.675
0.506
0.729
0.468
0.757
0.477
Full fine-tuning
4,096
0.670
0.497
0.736
0.476
0.760
0.474
LoRA
4,096
0.670
0.501
0.734
0.481
0.761
0.484
PaCTS
4,096
0.667
0.491
0.728
0.466
0.749
0.463
Appendix
Table 9: Detailed comparisons of different forecast horizons on GIFT-Eval.
MASE ↓
CRPS ↓
Context
Zero-shot
PaCTS
Δ (%)
Zero-shot
PaCTS
Δ (%)
4,096
0.716
0.701
−2.09
0.490
0.484
−1.22
8,064
0.704
0.692
−1.70
0.481
0.476
−1.04
Appendix
Table 10: GIFT-Eval performance of PatchTST-FM at different context lengths. The 8,192-step model window reserves a 128-step masked forecast span for the 96-step horizon, leaving 8,064 context steps.
MASE ↓
CRPS ↓
Context
Zero-shot
PaCTS
Δ (%)
Zero-shot
PaCTS
Δ (%)
4,096
0.717
0.706
−1.53
0.514
0.507
−1.36
8,192
0.709
0.700
−1.27
0.506
0.500
−1.19
16,384
0.706
0.698
−1.13
0.503
0.498
−0.99
Appendix
Table 11: GIFT-Eval performance of TimesFM-2.5 at matched context lengths.
Backbone
Width d
Context lengths
Trainable parameters
Trainable fraction
Chronos-2
768
4,096 / 8,192
153,286
0.128%
PatchTST-FM
1,024
4,096 / 8,064
199,110
0.077%
TimesFM-2.5
1,280
4,096 / 8,192 / 16,384
244,934
0.106%
Appendix
Table 12: Prompt-module configurations in the cross-backbone experiments. The trainable fraction is relative to the parameter counts of each frozen backbone.
Method (context)
MASE ↓
CRPS ↓
GFLOPs ↓
Memory (MiB) ↓
Latency (ms) ↓
Params (M)
Zero-shot (4,096)
0.716
0.490
2,245.5
1,199.4
50.1
257.896
Zero-shot (8,064)
0.704
0.481
4,536.8
1,370.8
101.9
257.896
PaCTS (4,096)
0.701
0.484
2,333.1
1,211.2
69.8
258.095
Δ vs. zero-shot (8,064)
−0.43%
+0.62%
−48.57%
−11.64%
−31.50%
+0.08%
Appendix
Table 13: GIFT-Eval performance and inference costs of PatchTST-FM r1.
Method (context)
MASE ↓
CRPS ↓
GFLOPs ↓
Memory (MiB) ↓
Latency (ms) ↓
Params (M)
Zero-shot (4,096)
0.717
0.514
1,719.3
1,006.7
123.6
231.289
Zero-shot (8,192)
0.709
0.506
3,546.0
1,099.5
242.7
231.289
PaCTS (4,096)
0.706
0.507
1,854.0
1,015.6
135.9
231.534
Δ vs. zero-shot (8,192)
−0.42%
+0.20%
−47.72%
−7.63%
−43.98%
+0.11%
Appendix
Table 14: GIFT-Eval performance and inference costs of TimesFM-2.5 at the 4,096-step comparison point.
Method (context)
MASE ↓
CRPS ↓
GFLOPs ↓
Memory (MiB) ↓
Latency (ms) ↓
Params (M)
Zero-shot (8,192)
0.709
0.506
3,546.0
1,099.5
242.7
231.289
Zero-shot (16,384)
0.706
0.503
7,521.6
1,286.7
480.3
231.289
PaCTS (8,192)
0.700
0.500
3,689.1
1,108.8
246.0
231.534
Δ vs. zero-shot (16,384)
−0.85%
−0.60%
−50.95%
−13.83%
−48.77%
+0.11%
Appendix
Table 15: GIFT-Eval performance and inference costs of TimesFM-2.5 at longer contexts.
In-context learning (ICL) enables task adaptation at inference time by conditioning on demonstrations rather than updating model parameters. Although recent time-series foundation models incorporate contextual conditioning, retrieval, or example-based prompting, they typically rely on implicit positional structure or task-specific objectives rather than explicit instruction-conditioned input-output demonstrations. We introduce iAmTime, a time-series foundation model trained with instruction-conditioned amortized meta-learning to infer tasks directly from example demonstrations. iAmTime represents each episode as a structured prompt over historical context and future-known variables using specialized semantic tokens that attend to designated time-series regions, exchange information across demonstrations, and inject task information into the query representation. The model combines a Hierarchical Multi-Scope Transformer Encoder, which captures temporal and covariate dynamics while inferring latent task structure from demonstrated input-output mappings, with a Task-Conditioned Patch Decoder, which adapts decoding through expert-based routing. We train iAmTime on large-scale real and synthetic corpora using supervised and self-supervised instruction-conditioned tasks, including forecasting, imputation, reconstruction, classification, anomaly detection, and source de-mixing. Across diverse domains, frequencies, and horizons, iAmTime improves zero-shot adaptation over strong time-series foundation baselines on probabilistic and point forecasting benchmarks, while achieving competitive or superior performance on four non-forecasting tasks.
While Next-Token Prediction (NTP) has unified LLM pretraining, its adaptation to unbounded, continuous time series (TS) remains open. To bridge the gap, we introduce UniTok, a universal tokenizer that transforms TS into discrete tokens, and UniTok-FM, a foundation model pretrained via NTP on these tokens. UniTok-FM is a general-purpose foundation model that supports zero-shot and prompt-boosted forecasting, as well as few-shot generation and classification via training-free in-context inference--a capability not achieved by prior works. Technically, UniTok is a vector-quantized autoencoder incorporating prefix normalization for scale stabilization, a progressive-resolution causal architecture for encoding and decoding, and a structure-preserving reconstruction loss for training. UniTok-FM adopts an off-the-shelf LLM architecture without TS-specific modifications. Instead of pretraining on isolated TS, it performs NTP on context windows formed by multiple series with similar patterns, aiming to capture their shared dynamics. Experiments on forecasting, generation, and classification show that a single unified UniTok-FM consistently outperforms statistical and supervised baselines, achieves competitive performance with task-specific foundation models, and uniquely enables training-free in-context inference across tasks.
Yunhao Zhang, Ruiying Qi, Jiale Zheng +3
Shanghai Jiao Tong University · Huawei Noah’s Ark Lab
Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more history into the model improves forecast quality. But in stochastic domains, distant history is often just high-frequency noise, not signal. Hence, the proposed work tests whether this premise actually holds by running continuous context architectures (PatchTST included) through the ETTh1 benchmark. The obtained results contradict the premise: an inverse scaling law shows up clearly, with forecasting error rising as context gets longer. A 3,000-step window causes performance to drop by over 68%, evidence that attention mechanisms are poor at ignoring irrelevant historical volatility. Retrieval-Augmented Forecasting (RAFT) is evaluated as an alternative. RAFT achieves a mean squared error (MSE) of 0.379 with a fixed 720-step window and selective retrieval, outperforming both long-context configurations and zero-shot foundation models (Chronos, Moirai) despite requiring far less computation. In addition, the retrieval step injects only the most relevant historical segments as dynamic exogenous variables, which gives the model a context-informed inductive bias it cannot build on its own from raw sequences. Therefore, foundation models going forward need to shift architecturally toward selective retrieval.
Rishi Ahuja, Kumar Prateek, Simranjit Singh +1
Department of Information Technology, Dr. B.R. Ambedkar National Institute of Technology Jalandhar, Punjab, 144008, India.