We present TimeBraid, a series of unified time-series and language models that align pretrained language models and pretrained time-series foundation models through interleaved global residual attention layers. Each model inherits knowledge, instruction following, and reasoning from one side, continuous-signal perception and zero-shot forecasting from the other, and fuses the two in a shared representation space where both modalities are understood and generated. We study the design choices that make such unified modeling work: where to align the two representation spaces, how to ground language in temporal structure, how to balance understanding with generation, and how to keep joint optimization stable. The resulting recipe combines a unified prompting scheme for diverse time-series and text tasks, stabilized joint training, and supervision from 2.2M curated series--text pairs and 4.9M instruction-tuning samples. Across benchmarks spanning time-series perception, understanding, reasoning, and both context-aided and unimodal forecasting, TimeBraid remains competitive with far larger general-purpose models and task-specific counterparts.
Figures & tables
Figure 2: Two-stage training recipe. Stage 2 counts show the final one-pass sample inventory, while percentages show training-time sampling shares.
Figure 3: TimeBraid architecture: the model components and the flow of information between the time-series and language modalities.
Model
TSAQA
TSExam
TB-MCQ
A.D. ↑
CLS ↑
Char. ↑
Comp. ↑
D.T. ↑
T.R. ↑
PZ ↑
Overall ↑
Overall ↑
Avg. ↑
Closed-source Reference Models
GPT-5.4
53.32
49.57
81.98
74.47
54.94
82.96
51.02
63.10
67.83
38.50
GPT-4.1
55.85
50.38
89.36
76.99
51.13
79.09
45.77
62.82
67.89
36.91
GPT-4o
54.32
47.20
84.15
69.07
53.24
75.58
45.61
60.73
55.96
32.30
Gemini-2.5-Flash
52.08
49.07
81.08
72.21
60.17
84.49
60.84
65.08
53.89
34.61
Table 1: Time-series understanding accuracy (%) on TSAQA (per-category and overall), TimeSeriesExam (TSExam), and the 16 TemporalBench multiple-choice tasks (TB-MCQ, unweighted macro-average). The Closed-source Reference Models group provides references excluded from ranking. A.D.: anomaly detection; CLS: classification; Char.: characterization; Comp.: comparison; D.T.: data transformation; T.R.: temporal relation; PZ: puzzling/ordering format. Char./Comp./D.T./T.R. report the multiple-choice format. SFT denotes TSAQA-specific LoRA fine-tuning ( Jing et al., 2026 ) . Per-format and per-category results are in Tables 13 , 12 , and 16 .
Model
Lexical and Semantic Alignment
Numeric ↑
DeBERTa-F1 ↑
SimCSE ↑
BLEU ↑
ROUGE-L ↑
METEOR ↑
Closed-source Reference Models
GPT-5.4
0.660
0.843
0.064
0.239
0.264
0.778
Gemini 2.0 Flash
0.694
0.884
0.113
0.283
0.304
0.757
GPT-4o
0.685
0.886
0.090
0.259
0.314
0.739
Open-source Large Language Models
Table 2: Contextual description on the CaTS-Bench human-rewritten split: five metrics of alignment with reference captions and the released Numeric Fidelity score, all in [0,1] . The Closed-source Reference Models group provides references excluded from ranking. Rows marked finetuned are trained on the CaTS-Bench training split; the full comparison is in Table 14 .
Model
TimeMMD (MSE ↓ )
CGTSF (MSE ↓ )
Agri.
Clim.
Econ.
Ener.
Env.
Health
Sec.
Soc.
Traf.
MSPG
LEU
PTF
Time-Series Foundation Models
TimesFM 2.5
0.127
0.859
0.017
0.239
0.569
0.665
114.8
0.845
0.152
0.585
0.777
0.243
Chronos-2
0.085
0.988
0.015
0.262
0.566
1.217
110.3
1.110
0.201
0.649
0.715
0.230
Moirai
0.102
0.949
0.020
0.277
0.581
1.280
109.8
0.859
0.163
1.000
0.681
0.284
Sundial
0.102
0.865
0.022
0.266
0.561
1.240
111.7
0.856
0.160
0.763
0.683
0.242
Table 3: Real-world contextual forecasting. Every text-conditioned model receives the paired text of each window, and a starred name (PatchTST ∗ ) marks a time-series backbone extended with a text branch. TimeMMD reports per-domain MSE averaged over four horizons; TimeBraid, the foundation models, and MIGAS-1.5 use a fixed 128-step history, while ChatTime, Aurora, and the text-conditioned supervised baselines use the official 8/36/96 lookbacks. CGTSF reports training-split-standardized MSE on MSPG, LEU, and PTF, averaged over four history-length settings. The remaining baselines, detailed results, and full model names are in Tables 18 and 19 .
Model
Ctrl-F
CAF
CiK
MSE ↓
Top-1 (%) ↑
CRPS ↓
RCRPS ↓
Closed-source Reference Models
GPT-5.4
5.188
51.33
0.233
0.145
GPT-4o
5.447
60.00
0.234
0.257
Gemini-2.5-Flash
6.624
56.00
0.247
0.110
Open-source Large Language Models
Table 4: Controlled forecasting with verifiable targets. The Closed-source Reference Models group provides references excluded from ranking. Ctrl-F: history- z -normalized MSE and Top-1 correct-sibling retrieval accuracy. CAF: normalized CRPS on all 904 correct-context test cases. CiK: RCRPS under official task weights. Details in Tables 21 , 21 , and 15 .
Metric
Statistical Methods
Task-Specific Models (Supervised)
Time Series Foundation Models
TimeBraid (Ours)
Naive
Seasonal Naive
Auto ARIMA
DeepAR
TiDE
N-BEATS
PatchTST
TimesFM 2.5
TabPFN-TS
Chronos 2
Moirai2
Sundial Base
TiRex
Toto-2.0 FnF
Chronicle
MASE
0.763
1.270
1.000
1.074
1.343
1.091
0.938
0.849
0.705
0.771
0.698
0.728
0.750
0.716
0.676
1.053
CRPS
0.546
1.591
1.000
0.912
0.853
0.772
0.816
0.587
0.490
0.544
0.485
0.516
0.559
0.488
0.463
0.754
Table 5: Aggregate GIFT-Eval forecasting performance. Lower is better.
Dataset
Zero-shot
Full-shot
TimeBraid-2.5B
Sundial L
TimeMoE U
Moirai L
TimesFM 2.5
Chronos-2
iTransformer
TimeMixer
PatchTST
DLinear
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
0.373
0.375
0.331
0.369
0.356
0.391
0.422
0.391
0.375
0.370
0.373
0.356
0.407
0.409
0.381
0.395
0.387
0.400
0.403
0.406
ETTm2
0.256
0.311
0.254
0.315
0.288
0.344
0.329
0.343
0.271
0.305
0.254
0.291
0.288
0.332
0.275
0.323
0.280
0.326
0.350
0.400
ETTh1
0.393
0.407
0.395
0.420
0.412
0.426
0.480
0.439
0.396
0.405
0.420
0.405
0.454
0.447
0.448
0.442
0.468
0.454
0.455
0.451
ETTh2
0.338
0.377
0.334
0.387
0.371
0.399
0.367
0.377
0.343
0.369
0.342
0.366
0.383
0.406
0.364
0.395
0.386
0.406
0.558
0.515
Table 6: Unimodal forecasting on ETT and Weather. Each dataset entry averages MSE or MAE over horizons {96,192,336,720} . Avg. rank aggregates the 20 dataset–horizon settings; lower is better. Per-horizon results are in Table 22 .
Figure 4: Training-loss trajectories for the design ablations; each panel compares 10k-step runs that differ only in the ablated factor: (A) cross-modal interface, (B) shared learning rate, (C) plain vs. robust RevIN, (D) understanding-to-forecasting data ratio, (E) dual- vs. tri-tower. Insets zoom into the last 2k steps; lower is better.
Figure 5: Effect of the alignment stage over the course of SFT. We compare aligned initialization with direct SFT from the pretrained towers on (a) TSAQA accuracy (10% test subset, higher is better), (b) TimeMMD MSE (lower is better), and (c) CAF normalized qCRPS (lower is better).
Figure 6: Effect of text guidance across domains and history lengths. Each panel sweeps the text-mixing weight λ on one TimeMMD domain, reporting the relative MSE change against the text-free forecast ( λ=0 ) at input histories 8–128; lower is better. At least one nonzero weight helps at every history length in Public Health, Economy, Agriculture, Energy, Environment, and Traffic; Security prefers λ=0 ; Climate and Social Good change sign with history. We report this sweep as a sensitivity analysis because the released text is not point-in-time certified; it is separate from model selection, for which a single λ=0.3 is selected on validation data and fixed across all domains.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Alignment data curation for univariate understanding, multivariate understanding, and univariate forecasting.
Family
Curation route
n
Share
Understanding 668,850 rows, 30.0%
Morphology captions
morphology-grounded VLM annotation
50,000
2.2%
Context-rich captions
context-grounded VLM annotation
50,000
2.2%
ChatTS ( Xie et al., 2024 ) univariate
attribute-programmed synthesis
260,000
11.7%
ChatTS ( Xie et al., 2024 ) multivariate
attribute-programmed synthesis
108,850
4.9%
Multivariate SCM
SCM-grounded relation synthesis
100,000
4.5%
Appendix
Table 7: Alignment corpus: 12 families, 2,229,500 examples. H is the forecast horizon. Rewrite variants reuse the base rows and change only the user-visible text.
Source
Final samples
Training Share
Forecasting 2,640,566 samples, 54.49%
CGTSF ( Wang et al., 2025a )
53,926
2.04%
FinMultiTime S&P 500 ( Xu et al., 2025b )
144,387
5.46%
MoTime (News/Wiki) ( Zhou et al., 2025b )
32,589
1.23%
CAF-7M ( Zheng et al., 2026 )
2,314,216
42.15%
Time-MMD ( Liu et al., 2024a )
20,768
0.79%
Appendix
Table 8: Final one-pass Stage-2 sample inventory and training-time sampling shares. References identify the source datasets; counts refer to our selected and reformatted examples.
Benchmark
Status
Understanding
TemporalBench (TB-MCQ)
Out-of-domain.
TSExam
Independently generated and rewritten samples from the same data-generating process are included in the alignment mixture.
TSAQA
Training split is included in the SFT mixture.
CaTS-Bench
Training split is included in the SFT mixture.
Forecasting
Appendix
Table 9: Training status of the evaluation suites. Status is determined by split membership in the mixture.
Variant
Language tower
Fusion
Time-series tower
Total
TimeBraid-1.2B
0.596
0.378
0.231
1.205
TimeBraid-2.5B
1.721
0.545
0.231
2.497
TimeBraid-6.7B
4.022
2.265
0.389
6.676
Appendix
Table 10: Parameter inventory across TimeBraid scales, in billions of parameters.
Stage 1 (alignment)
Stage 2 (supervised fine-tuning)
Optimizer
AdamW, 8-bit states
Learning rate
2×10−5 , all parameters
Schedule
constant with warmup
Warmup steps
500
Weight decay
0
Gradient clipping
1.0
Appendix
Table 11: Training hyperparameters for both stages. Rows spanning both columns are identical across stages.
Table 13: TSAQA test-set accuracy (%). A.D.: anomaly detection; CLS: classification; TF/MC/PZ: true-or-false, multiple-choice, and puzzling/ordering formats. SFT denotes TSAQA-specific LoRA fine-tuning ( Jing et al., 2026 ) . Closed-source Reference Models use zero-shot evaluation. Best, second, and third results are marked within the primary comparison, including the TSAQA-specific SFT rows and excluding Closed-source Reference Models; ties share a rank.
Model
Lexical and Semantic Metrics
Numeric
DeBERTa F1
SimCSE
BLEU
ROUGE-L
METEOR
Closed-source Reference Models
GPT-5.4
0.660
0.843
0.064
0.239
0.264
0.778
Gemini 2.0 Flash
0.694
0.884
0.113
0.283
0.304
0.757
GPT-4o
0.685
0.886
0.090
0.259
0.314
0.739
Open-source Vision-Language Models
Appendix
Table 14: Full results on the CaTS-Bench human-rewritten split.
Model
Average RCRPS
Intemporal Information
Historical Information
Future Information
Covariate Information
Causal Information
Closed-source Reference Models
Gemini-2.5-Flash
0.110 ± 0.002
0.134 ± 0.003
0.142 ± 0.001
0.036 ± 0.001
0.116 ± 0.003
0.252 ± 0.012
GPT-5.4
0.145 ± 0.000
0.182 ± 0.001
0.116 ± 0.001
0.049 ± 0.000
0.160 ± 0.001
0.414 ± 0.001
GPT-4o
0.257 ± 0.001
0.305 ± 0.002
0.135 ± 0.001
0.159 ± 0.000
0.238 ± 0.001
0.613 ± 0.006
GPT-5.4-mini
0.277 ± 0.000
0.302 ± 0.001
0.173 ± 0.001
0.212 ± 0.000
0.261 ± 0.001
0.553 ± 0.000
Open-source Large Language Models
Appendix
Table 15: Results on CiK: RCRPS weighted by the official task weights, reported with standard errors and stratified by context type. Asterisks mark models that do not use natural-language context.
Model
FreshRetailNet
PSML
Causal Chambers
MIMIC
Average
T1
T2
T3
T4
T1
T2
T3
T4
T1
T2
T3
T4
T1
T2
T3
T4
Closed-source Reference Models
GPT-5.4
45.45%
28.03%
41.48%
42.42%
50.50%
31.33%
33.20%
57.33%
18.67%
53.33%
37.20%
45.33%
35.11%
29.08%
35.56%
31.91%
38.50%
GPT-4.1
38.64%
31.82%
45.45%
38.64%
48.50%
27.33%
40.00%
53.33%
12.00%
46.00%
48.80%
41.33%
18.62%
29.08%
42.68%
28.37%
36.91%
GPT-4o
63.07%
16.67%
28.98%
39.39%
69.00%
23.33%
35.20%
36.67%
10.00%
22.67%
34.00%
42.00%
46.81%
19.86%
0.00%
29.08%
32.30%
Gemini-2.5-Flash
59.66%
22.73%
29.55%
38.64%
72.50%
23.33%
22.00%
46.67%
10.67%
45.33%
37.20%
44.00%
42.55%
29.08%
0.00%
29.79%
34.61%
Appendix
Table 16: TemporalBench results under the official evaluation setup: (a) accuracy on the 16 multiple-choice tasks, with Average the unweighted macro-average, and (b) forecasting error, not averaged across datasets because the metrics differ.
Table 18: TimeMMD forecasting with paired text, reporting per-domain MSE and MAE macro-averaged over four horizons. Every multimodal forecasting model receives each window’s paired text and runs at the official 8/36/96 lookbacks, as do ChatTime and Aurora; TimeBraid, the time-series foundation models, and MIGAS-1.5 use a fixed 128-step history. Avg Rank averages each model’s rank across the nine domains; lower is better.
Model
Per-dataset (MSE / MAE ↓ )
Macro
MSPG
LEU
PTF
MSE ↓
MAE ↓
Time-Series Foundation Models
Chronos-2
0.649 / 0.388
0.715 / 0.414
0.230 / 0.285
0.531
0.363 †
TimesFM-2.5
0.585 / 0.384
0.777 / 0.442
0.243 / 0.306
0.535
0.377
Sundial
0.763 / 0.507
0.683 / 0.453
0.242 / 0.306
0.563
0.422
Moirai-2
1.000 / 0.590
0.681 / 0.422
0.284 / 0.332
0.655
0.448
Appendix
Table 19: CGTSF context-guided forecasting ( Wang et al., 2025a ) on MSPG (Melbourne solar power generation), LEU (London electricity usage), and PTF (Paris traffic flow). Errors are standardized using training-split statistics; per-dataset entries average four history-length settings, and Macro MSE and Macro MAE weight the twelve settings equally. We replicate the official evaluation setting and evaluate every baseline ourselves.
Table 26
Figure 8: Aggregate GIFT-Eval forecasting performance across models and context lengths. Lower is better.
TimeBraid
Time-Series Foundation Models
Traditional Time-Series Models
Models
TimeBraid-2.5B
Sundial Large
Time-MoE Ultra
Moirai Large
TimesFM 2.5
Chronos-2
iTransformer
TimeMixer
PatchTST
DLinear
(Ours)
(Zero-shot)
(Zero-shot)
(Zero-shot)
(Zero-shot)
(Zero-shot)
(Full-shot)
(Full-shot)
(Full-shot)
(Full-shot)
Metric
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.310
0.336
0.273
0.329 †
0.281
0.341
0.380
0.361
0.307
0.326
0.301 †
0.312
0.334
0.368
0.320
0.357
0.329
0.367
0.345
0.372
192
0.356
0.364
0.312
0.357
0.305
0.358 †
0.412
0.383
0.358
0.358 †
0.352 †
0.343
0.377
0.391
0.361
0.381
0.367
0.385
0.380
0.389
336
0.388 †
0.385
0.343
0.378
0.369
0.395
0.436
0.400
0.389
0.381 †
0.388 †
0.367
0.426
0.420
0.390
0.404
0.399
0.410
0.413
0.413
Appendix
Table 22: Per-horizon long-term forecasting on ETT and Weather, reporting MSE and MAE at horizons {96,192,336,720} and their average. TimeBraid-2.5B uses a 2,880-step lookback; the baselines retain their official input lengths. Avg Rank averages each model’s rank across the twenty dataset–horizon settings.
Stage
MMLU (%)
Base (Qwen3-1.7B language tower)
60.30
Align
40.88
Align, w/ unimodal replay
53.38
SFT
50.68
SFT, w/ unimodal replay during alignment
50.09
Appendix
Table 23: Five-shot MMLU accuracy (%) of TimeBraid-2.5B across training stages ( Hendrycks et al., 2020 ) . Base is the pretrained Qwen3-1.7B language tower evaluated standalone on text-only prompts; higher is better.
Figure 9: Context-conditioned forecasting with and without scenario text.
Figure 10: Context-conditioned and text-controlled forecasting on test examples. In the lower panels, solid blue, green, and purple denote conditions 1–3 over the same held-out history; dashed blue in the upper panels denotes forecasting without text. Only the differing condition text is displayed.
Figure 11: Text-controlled forecasting on held-out test histories. Each group uses three textual conditions with the same numerical input. Colored curves are recorded model outputs; the text below each panel quotes the differing condition. These are constructed scenarios, not observed alternative futures. Full histories and complete forecast horizons are shown on separate horizontal axes with the same vertical scale.
Figure 12: Additional text-controlled forecasts on held-out histories. All three predictions use the displayed textual conditions, quoted from the full inputs. The constructed scenarios vary trends, levels, and local changes; they illustrate the model’s responses to distinct instructions. Both axes use the same vertical scale and display the complete input and output sequences.
Figure 13: Selected Time-MMD test forecasts with matched contextual-text ablations. Orange uses the full dataset-provided textual context; dashed blue removes that context while retaining the same forecast instruction, numerical history, and horizon. Dashed gray is the observed future. Text below each panel is a verbatim excerpt of the full context used for prediction. MAE is measured over the complete future horizon in source units; reductions describe these selected examples.
Figure 14: Additional matched Time-MMD test forecasts with market context. Orange forecasts use textual context; dashed blue forecasts remove only that context. Displayed excerpts are taken from the full inputs. The same checkpoint and inference settings are used for both arms. Complete histories and future horizons are shown, and the two axes in each example share their vertical scale.
Figure 15: Selected CGTSF test forecasts using weather, calendar, and location text together with the numerical history. The contextual passages supplied to the model are displayed below the curves. Orange is the recorded text-conditioned prediction and dashed gray is the observed future. These examples show conditional forecast quality; a matched context-removal comparison is not included for this page.
Figure 16: Series-property QA, sleep-stage classification, and multivariate relation analysis.
Figure 17: Scenario-based reasoning and time-series captioning.
Figure 18: Sensor anomalies, human activity, bone age, and anomaly localization.
Figure 19: Question answering, paraphrase consistency, and time-series classification.
Real-world time series come with text: metadata, descriptions, news, reports. Yet time series foundation models process numerical sequences in isolation, and the multimodal text-and-time-series models that attempt to bridge the two all adapt a pretrained language model post hoc, inheriting representations shaped without ever seeing temporal data. These models are also evaluated almost exclusively against other multimodal baselines, not against the strongest unimodal foundation models in either domain, leaving open whether joint training is needed at all. We present Chronicle, a compact 324M-parameter decoder-only transformer trained from scratch on natural language and time series within a single unified architecture. Both modalities share the same transformer blocks, attention mechanism, and residual stream; the bulk of pretraining uses unimodal batches so cross-modal capability emerges purely from shared parameters, with a short alignment stage that interleaves the two. To our knowledge, Chronicle is the first model jointly pretrained on text and time series from scratch, and the first multimodal model evaluated against dedicated foundation models in both domains. It matches Gemma-3-270M-PT on 19 NLU tasks, sets a new bar for frozen-embedding time series classification on 24 UCR/UEA datasets, and produces multimodal forecasts on Time-MMD that beat every supervised fusion baseline, all from a single backbone.
Paul Quinlan, Jeremy Levasseur, Qingguo Li +1
1InertialAI · Department of Electrical and Computer Engineering, Queen’s University · Department of Mechanical and Materials Engineering, Queen’s University
Real-world time-series applications increasingly require models that can handle time series forecasting, context-conditioned prediction, and language-based temporal reasoning. Yet current time-series foundation models remain fragmented across these capabilities: numerical specialists often provide the strongest forecasts, while language-based models offer broader contextual understanding and analysis. A central challenge is to unify these heterogeneous capabilities without reducing their individual performance. We introduce OpenTSLM TeeMoE, a generalist time-series language model that can forecast directly from observed time series, reason over textual context and temporal patterns, and synthesize and refine predictions from external numerical forecasting specialists. We independently train three low-rank experts for forecast aggregation, native forecasting, and temporal analysis over a shared backbone. A learned LoRA mixture-of-experts controller then weights their frozen parameter updates for each request. Our proposed model achieves strong performance on widely used benchmarks for time series forecasting, context-conditioned prediction, and language-based temporal reasoning, ranking among the top three on GIFT-Eval by mean MASE rank, Context is Key by RCRPS, and TimeSeriesExam by accuracy.
Tony Chen, Timo Stoffregen, Maxwell Xu +8
Columbia University, USA · Stanford University, USA · Aionic Labs, Switzerland +4
Most time series (TS) models are specialized for a single task, either understanding (i.e., returning text answers about a TS) or generation (i.e., returning a numeric forecast). Only recently have unified models begun to handle the two within a single architecture. Even these models, however, produce the two outputs as task-separated paths and cannot predict a series and explain why that prediction arises within a single coherent response. In this paper, we argue for a task-fused model that jointly produces 1) prediction (generation) and 2) selfexplanation (understanding), thereby integrating 1) numerical TS forecasting and 2) interpretable text reasoning within a single response. To enable the systematic study of this capability, we present both a benchmark and a recipe that jointly address the two tasks. The benchmark, ReasonTS-Bench, identifies five fundamental patterns underlying TS and enables the joint evaluation of both tasks. ReasonCast, our recipe for finetuning any LLM to perform both tasks jointly, yields a model that generates a reasoning chain and a forecast together in a single autoregressive pass. Extensive experiments show that ReasonCast outperforms both LLMs and TS models on prediction accuracy while producing verifiable, causal reasoning. Code is available at: https://github.com/seunghan96/reasoncast.