CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning
Organizations: Graduate School of Information, Yonsei University, Seoul, Republic of Korea
Abstract
Time series forecasting underpins critical decision-making across diverse domains. While large language models (LLMs) offer promising reasoning capabilities, existing LLM-based time series forecasting approaches either reduce them to numerical predictors that bypass their strengths, or allow direct forecast generation that destabilizes predictions in non-stationary settings. We introduce CTRL, a framework that decouples semantic reasoning from quantitative prediction. A frozen backbone generates base forecasts, while specialized LLM agents function as controllers that analyze backbone prediction errors through decomposed trend, seasonal, and irregular components, grounding reasoning in interpretable temporal structure. Each agent outputs compact control signals that a lightweight residual decoder translates into forecast corrections. CTRL incorporates label-free test-time adaptation that detects distribution shift from input statistics alone and readapts control signals with only 3-24 LLM calls via caching. CTRL is explicitly designed to improve robustness under non-stationary temporal dynamics and distribution shift, while remaining competitive on highly stationary time series where adaptive correction provides limited additional benefit.
Figures & tables
| Data | H | Ours(DL) | Ours(PT) | DLinear | PatchTST | TimeLLM † | CALF | TEMPO | GPT4TS | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | ||
| ETTh1 | 48 | .378 | .340 | .381 | .342 | .382 | .346 | .386 | .346 | .398 | .369 | .400 | .372 | .419 | .395 | .425 | .412 |
| 96 | .398 | .371 | .399 | .372 | .399 | .372 | .405 | .380 | .429 | .414 | .410 | .393 | .469 | .471 | .464 | .473 | |
| 192 | .429 | .410 | .439 | .429 | .430 | .412 | .439 | .431 | .432 | .425 | .436 | .430 | .494 | .514 | .527 | .598 | |
| 336 | .457 | .442 | .430 | .415 | .459 | .444 | .431 | .416 | .452 | .450 | .476 | .490 | .477 | .485 | .599 | .738 | |
| 720 | .510 | .498 | .458 | .443 | .511 | .497 | .459 | .444 | .477 | .462 | .522 | .540 | .502 | .502 | .600 | .783 | |
| Exch. | ETT | ECL | Weather | ||||
|---|---|---|---|---|---|---|---|
| h2 | m2 | h1 | m1 | ||||
| ADF | -1.9 | -4.1 | -5.7 | -5.9 | -15.0 | -8.5 | -26.7 |
| Shift | 2.33 | 1.48 | 1.49 | 0.50 | 0.50 | 0.62 | 0.71 |
| ETTm2 | Weather | |||
|---|---|---|---|---|
| MSE | MAE | MSE | MAE | |
| LLMTime | .245 | .151 | .237 | .150 |
| CTRL | .197 | .070 | .132 | .039 |
| Improv. | 19.6% | 53.2% | 44.3% | 74.2% |
| Configuration | ETTh2 | ETTm2 |
|---|---|---|
| Baseline | ||
| Backbone only | +9.45 | +7.20 |
| Learned control signal | +3.13 | +4.42 |
| Random control signal | +5.84 | +4.54 |
| Zero control signal | +5.70 | +5.18 |
| Test-time Adaptation | ||
| Ours | LLMTim | GPT4TS | CALF | TEMPO | |
|---|---|---|---|---|---|
| Params | 400K | 0 | 4.4M | 18.2M | 12.4M |
| Train | 1.8m | 0 | 2.1m | 62m | 106m |
| LLM | 3–24 | 0 | 0 | 0 |
| Dataset | Test Samples | Checks | Max Calls |
|---|---|---|---|
| ETTh1/h2 | 2,402 | 2 | 9 |
| ETTm1/m2 | 11,042 | 7 | 24 |
| Weather | 10,060 | 7 | 24 |
| ECL | 4,781 | 3 | 12 |
| Exchange | 1,038 | 1 | 6 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Total | Train | Val | Test | Freq. |
|---|---|---|---|---|---|
| ETTh1/h2 | 17,420 | 12,194 | 1,742 | 3,484 | Hourly |
| ETTm1/m2 | 69,680 | 48,776 | 6,968 | 11,042 | 15-min |
| Weather | 52,696 | 36,887 | 5,269 | 10,540 | 10-min |
| ECL | 26,304 | 18,412 | 2,631 | 5,261 | Hourly |
| Exchange | 7,588 | 5,311 | 759 | 1,518 | Daily |
| Dataset | LR | Epochs | STL (Seas., Trend) |
|---|---|---|---|
| ETTh1 | 1e-4 | 40 | 24, 48 |
| ETTh2 | 1e-4 | 40 | 24, 48 |
| ETTm1 | 1e-4 | 40 | 24, 48 |
| ETTm2 | 1e-4 | 40 | 24, 48 |
| ECL | 5e-6 | 40 | 12, 24 |
| Exchange | 1e-4 | 40 | 7, 24 |
| Configuration | MSE (%) |
|---|---|
| LLM Sampling (Llama 3.3 70B) | |
| temp=0.2, top-p=0.8 | +0.22 |
| temp=0.3, top-p=0.8 | – |
| temp=0.5, top-p=0.8 | +0.26 |
| Adaptation Threshold ( ) | |
| =OFF (no adaptation) | +1.03 |
| Component | Params |
|---|---|
| DLinear backbone (frozen) | 74K |
| GPT-2 encoder (frozen) | 124M |
| MLP projector (trained) | 100K |
| Residual decoder wo proj. (trained) | 215K |
| Total trainable | 400K |
| Zero | w/o TTA | Full | |
|---|---|---|---|
| Mean Imp. | +4.3% | +12.6% | +13.3% |
| Win Rate | 49% | 76% | 82% |
| Comparison | Improv. |
|---|---|
| w/o Adaptation vs Zero | +8.3% |
| CTRL (Full) vs w/o Adaptation | +0.7% |
| CTRL (Full) vs Zero | +9.0% |
| Scale | Bias | Gate | Conf. | |
| Before Adaptation (Stage 1) | ||||
| Trend | 0.90 | 0.10 | 0.40 | 0.60 |
| Seasonal | 0.95 | 0.10 | 0.40 | 0.60 |
| Irregular | 0.95 | 0.10 | 0.50 | 0.60 |
| After Adaptation (Stage 3) | ||||
| Trend | 0.50 | 0.50 | 0.69 | 0.21 |
| Scale | Bias | Gate | Conf. | |
| Before Adaptation (Stage 1) | ||||
| Trend | 0.90 | 0.10 | 0.40 | 0.60 |
| Seasonal | 0.95 | 0.10 | 0.40 | 0.60 |
| Irregular | 0.95 | 0.10 | 0.50 | 0.60 |
| After Adaptation (Stage 3) | ||||
| Trend | 0.58 | 0.29 | 0.59 | 0.28 |
| Component | (ETTh2) |
|---|---|
| Irregular text emb. | 0.032 |
| Trend bias | 0.009 |
| Trend confidence | 0.006 |
| Trend scale | 0.005 |
| Trend gate | 0.005 |
| Seasonal confidence | 0.001 |
| Text Variant | |
|---|---|
| No text (zeroed embedding) | 0.0323 |
| Random unrelated words | 0.0074 |
| Technical jargon | 0.0045 |
| Crisis/volatility description | 0.0010 |
| Random characters | 0.0005 |
| Calm/stable description | 0.0004 |
| Trend Specialist Agent | |
|---|---|
| Role | First specialist handling TREND correction |
| Focus | Low-frequency patterns: direction, level shifts, drift, slope accuracy |
| Key Questions | (1) Slope sign match? (2) Slope value match? (3) Systematic over/under-prediction? |
| Output: 4D control signal | |
| scale | Trend amplitude adjustment ( 1 if underestimate, 1 if overestimate) |
| bias | Level correction ( if too low, if too high) |
| Seasonal Specialist Agent | |
|---|---|
| Role | Second specialist handling SEASONAL correction |
| Focus | High-frequency periodic patterns: daily/weekly cycles, amplitude, phase |
| Key Questions | (1) Correct periodicity? (2) Amplitude correct? (3) Phase aligned? |
| Output: 4D control signal | |
| scale | Amplitude adjustment ( 1 if dampened, 1 if exaggerated) |
| bias | Baseline shift (usually ) |
| Irregular Pattern Agent | |
|---|---|
| Role | Third specialist handling noise/anomalies not explained by trend or seasonality |
| Task | Generate concise text description converted to numerical embedding |
| Analysis priorities (ordered) | |
| 1. Noise | Variance level: high/low/changing |
| 2. Anomalies | Outliers, spikes present? |
| 3. Autocorr. | Correlated or white noise? |
| Trend/Seasonal Adaptation | |
|---|---|
| Input | Component shift analysis: validation slope/dominance vs test, shift ratio (%), direction change |
| Output (JSON) | |
| shift_diagnosis | Description of detected shift |
| impact_assessment | Expected effect on predictions |
| scale_multiplier | Multiplicative adjustment to scale |
| bias_shift | Additive adjustment to bias |
| Irregular Adaptation | |
|---|---|
| Input | Original analysis text, statistics comparison (std, mean, spike ratio, autocorr) for validation vs test |
| Output (JSON) | |
| needs_adjustment | Boolean flag |
| updated_text | Revised irregular analysis |
| adapted_params | {scale, bias, gate, confidence} adjustments |
| reasoning | Explanation |
| Unified Distribution Shift Adaptation | |
|---|---|
| Input: STL decomposition comparison | |
| Trend shift | Validation vs test slope, shift ratio, significance |
| Seasonal shift | Validation vs test dominance, shift ratio, significance |
| Raw statistics | Mean, std comparison, overall shift score |
| Output (JSON) | |
| Base params | shift_diagnosis, impact_assessment, scale_multiplier, bias_shift, gate_multiplier, confidence_penalty |