OpenTSLM TeeMoE: A Unified Time-Series Language Model for Forecasting, Contextual Prediction, and Reasoning
Authors: Tony Chen, Timo Stoffregen, Maxwell Xu, Thomas Kaar, Martin Maritsch, Geremia Pompei, Nicolas Zumarraga, Robert Jakob, +3 more
Organizations: Columbia University, USA · Stanford University, USA · Aionic Labs, Switzerland · Google · Agentic Systems Lab, ETH Zürich, Switzerland · University of Pisa, Italy · National University of Singapore
Real-world time-series applications increasingly require models that can handle time series forecasting, context-conditioned prediction, and language-based temporal reasoning. Yet current time-series foundation models remain fragmented across these capabilities: numerical specialists often provide the strongest forecasts, while language-based models offer broader contextual understanding and analysis. A central challenge is to unify these heterogeneous capabilities without reducing their individual performance. We introduce OpenTSLM TeeMoE, a generalist time-series language model that can forecast directly from observed time series, reason over textual context and temporal patterns, and synthesize and refine predictions from external numerical forecasting specialists. We independently train three low-rank experts for forecast aggregation, native forecasting, and temporal analysis over a shared backbone. A learned LoRA mixture-of-experts controller then weights their frozen parameter updates for each request. Our proposed model achieves strong performance on widely used benchmarks for time series forecasting, context-conditioned prediction, and language-based temporal reasoning, ranking among the top three on GIFT-Eval by mean MASE rank, Context is Key by RCRPS, and TimeSeriesExam by accuracy.
Figures & tables
Figure 1: OpenTSLM TeeMoE shares one language backbone across three capability adapters. An adapter-disabled pass forms the request representation for the controller, which selects adapter weights and an output path. A second pass applies the mixed adapters to the full expert inputs. When the aggregation expert has a mixture weight above 0.5, external forecasts enter through the numerical connector. Dashed arrows denote mixture weights.
Table 1: The same composed OpenTSLM TeeMoE model across three benchmarks. DP: direct prompting; CorDP: direct prompting for forecast correction; SW: SampleWise. TSFMs receive no textual context on CiK; n/a denotes unsupported output. * marks our evaluations; OOT denotes projected cost above our compute budget (Appendix E.1 ). TimeOmni-VL’s TSE score likely reflects unsuccessful transfer despite using its official reasoning interface and the TSE question and scoring protocol (Appendix E.3 ). Comparison sources appear in Appendix C .
Table 2: Individual experts and TeeMoE. Deltas measure improvements over the pre-adaptation baseline: the pre-edit ensemble on GIFT and prompted unadapted Qwen on CiK and TSE. Parentheses in the baseline row give absolute scores. The aggregation expert uses its numerical interface on GIFT and its adapter with the language-model head on CiK and TSE. OOT follows Appendix E.1 .
Table 3: Composition controls and joint training: improvements over the pre-adaptation baseline (Table 2 ). Composition controls use the same frozen specialists; joint training uses one shared adapter and a separate output selector.
Table 4: Training and expert-plus-controller revision costs in H100 GPU-hours, excluding shared numerical preparation.
Table 5: Numerical ensembling and refinement, ordered by GIFT mean rank (best first). Numerical baselines use history only; TeeMoE also receives context on CiK. Protocols are in Appendix B .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Rule and joint value
LoRA rank
rJ=32
LoRA scaling / dropout
αJ=2rJ=64 ; dropout 0 , inherited
Global batch size
BJ=median(32,8,8)=8
Peak learning rate
ηJ=min(3×10−5,10−5,10−5)=10−5
Weight decay
λwd,J=round1sf((0+0.01+0)/3)=0.003
Schedule
mode(constant,cosine,cosine)=cosine
Appendix
Table 6: Joint settings derived from the three specialist recipes. Tuples follow aggregation, native forecasting, and analysis; round1sf rounds to one significant digit.
Category
MASE
MASE rank
CRPS
CRPS rank
m4_yearly/A/S
2.952
5.000
0.105
7.000
m4_quarterly/Q/S
1.098
4.000
0.071
8.000
m4_monthly/M/S
0.881
5.000
0.089
7.000
m4_weekly/W/S
1.820
3.000
0.034
3.000
m4_daily/D/S
3.165
18.000
0.021
12.000
m4_hourly/H/S
0.644
10.000
0.018
8.000
Appendix
Table 7: TeeMoE on all 97 GIFT-Eval dataset–frequency–horizon categories (371,330 forecast windows). Ranks use the same 134-entry population as Table 1 . Category errors are raw. Overall follows the official leaderboard: geometric means of seasonal-naive-normalized errors and mean ranks, equally weighting categories ( Salesforce, 2026 ) . S/M/L denote short/medium/long horizons. Lower is better throughout.
One-shot evaluation: IC-DP supplies a labelled instance from the same CiK task type, including its ground-truth future. We include DP/CorDP without demonstrations in Table 1 and report IC-DP separately ( Ashok et al., 2026 ) .
TimeClaw
Average RCRPS 0.115
This score covers the held-out portion of a benchmark train/test split, rather than the full 355-instance panel ranked in Table 1 . Scores over these different evaluation populations are not directly comparable ( Li et al., 2026 ) .
These values are reported in the same TimeClaw table and inherit its held-out evaluation subset ( Li et al., 2026 ) .
Multimodal forecasting study
MAE, MSE, WQL, and ordinary CRPS
The arXiv v2 CiK experiments explicitly omit RCRPS, so none of these metrics is the canonical ranking statistic ( Zhang et al., 2026b ) .
TsLLM
sMAPE 64.500% ; MASE 64.700%
No RCRPS is reported and the evaluated instances are not specified; the point metrics cannot be converted into weighted RCRPS ( Parker et al., 2026 ) .
CodeCast
MSE 22,038.000 ; MAE 30.520 , plus ELO/rank/win rate
These are point-forecast or preference metrics rather than probabilistic RCRPS ( Zhang et al., 2026a ) .
Appendix
Table 8: Additional CiK reports and the protocol differences that exclude them from the full 355-instance comparison.
Work or setting
Reported
Comparison scope
Original TSE v1.0
Category accuracies, including strong one-shot image/text baselines
The original 763-question population differs from v1.1; its rounded category results are not inserted into the v1.1 table ( Cai et al., 2024 ) .
TSEA
GPT-4o 73.000%; Gemini-2.5-Pro 71.000%
Category counts describe 763 questions, not 746 ( Gwiazda et al., 2026 ) .
The category denominators are consistent with the older 763-question population, without establishing question-level identity ( Yu et al., 2026 ) .
TsLLM
Easy 87.300%; hard 85.500%
The inspected arXiv v2 does not specify the release, partition sizes, or overall accuracy, preventing placement on the full v1.1 ranking ( Parker et al., 2026 ) .
ITFormer
High category accuracies
TSE fine-tuning is reported without an established held-out split, release, or overall accuracy ( Wang et al., 2025b ) .
TimeOmni-1
47.800% average
Uses 746 questions, but reports an equal-category mean rather than all-question accuracy ( Guan et al., 2026a ) .
Appendix
Table 9: Other TSE reports and their relationship to the v1.1 comparison. Reported values are not converted into full-population accuracy when counts or averaging rules are unavailable.
Figure 2: Task-text and full-request views of an illustrative contextual forecasting request. Blue fields appear in both views; the full view also contains the purple history and forecast-timestamp fields. The controller combines both representations, and the expert receives the full request. The shared system instruction is unchanged.
Model
Mean-based projection
Faster-request sensitivity
OpenTSLM SP Llama 3.2 1B
35,000
28,000
ChatTS-14B
450,000
330,000
TS-Reasoner-7B
500,000
350,000
TimeOmni-VL
100,000
92,000
Native forecasting expert
130,000
110,000
Analysis expert
140,000
110,000
Appendix
Table 10: Projected GIFT-Eval cost in allocated H100 GPU-hours for 25 trajectories per window, rounded to two significant figures. All six profiled models exceed the 1,000-hour budget under both estimates.
Configuration
Capped
Uncapped
TeeMoE and adapter controls
OpenTSLM TeeMoE
0.115
0.115
Qwen3.6-27B (no adapters)
0.151
0.169
Aggregation expert (text transfer)
0.154
0.168
Analysis expert
0.138
0.138
Native forecasting expert
0.123
0.123
Appendix
Table 11: CiK RCRPS with and without the official per-instance cap of 5 (lower is better). Local results use all 355 instances and the same saved 25 trajectories for both columns. Published Table 1 contenders are listed separately: their reported scores are retained, but paired capped/uncapped values are unavailable. N/A denotes an unsupported full-benchmark evaluation.
Component
H100 GPU-hours
FM forecast caches, training
58.530
FM forecast caches, evaluation
16.770
XGBoost, full-data fit and ten cross-fitting folds
1.390
Aggregation editor, 4,096 training examples
0.480
Native forecasting, 20,000 examples
21.839
Analysis, 12,000 examples
6.930
Appendix
Table 12: Cache preparation and training costs in H100 GPU-hours.
Most time series (TS) models are specialized for a single task, either understanding (i.e., returning text answers about a TS) or generation (i.e., returning a numeric forecast). Only recently have unified models begun to handle the two within a single architecture. Even these models, however, produce the two outputs as task-separated paths and cannot predict a series and explain why that prediction arises within a single coherent response. In this paper, we argue for a task-fused model that jointly produces 1) prediction (generation) and 2) selfexplanation (understanding), thereby integrating 1) numerical TS forecasting and 2) interpretable text reasoning within a single response. To enable the systematic study of this capability, we present both a benchmark and a recipe that jointly address the two tasks. The benchmark, ReasonTS-Bench, identifies five fundamental patterns underlying TS and enables the joint evaluation of both tasks. ReasonCast, our recipe for finetuning any LLM to perform both tasks jointly, yields a model that generates a reasoning chain and a forecast together in a single autoregressive pass. Extensive experiments show that ReasonCast outperforms both LLMs and TS models on prediction accuracy while producing verifiable, causal reasoning. Code is available at: https://github.com/seunghan96/reasoncast.
Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context-aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, and validate forecasts against temporal and domain constraints. In this work, we propose CastFSR, an agentic framework that formulates context-aware forecasting as a Fast--Slow--Reflect workflow. In fast thinking, CastFSR profiles observations and selects lightweight forecasters to construct a data-driven forecast prior. In slow deliberation, it retrieves contextual evidence, adaptively determines informative look-back windows, and reasons about how contexts reshape future dynamics. In reflection, it iteratively refines forecasts to ensure temporal, contextual, and domain consistency. CastFSR supports both training-free inference with off-the-shelf LLMs and efficient deployment through a two-stage SFT and reinforcement learning strategy that transfers its orchestration capability to compact LLMs. Extensive experiments on public datasets demonstrate that CastFSR consistently outperforms representative baselines. Our code is available at https://github.com/Xiaoyu-Tao/CastFSR.
Xiaoyu Tao, Mingyue Cheng, Bokai Pan +6
State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China, Hefei, Anhui Province, China
Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for events and constraints that the past alone cannot reveal. This requires both reliable numerical forecasting and the ability to interpret contextual information. Time-series foundation models (TSFMs) provide strong numerical forecasts, while large language models (LLMs) can reason over text, but combining their strengths remains challenging because asking an LLM to generate or revise forecast values directly can distort the temporal structure captured by the TSFM. We instead formulate forecasting as a planning problem over TSFM-generated trajectories. The frozen TSFM acts as a simulator that proposes numerical continuations, while the LLM acts as a policy and value function that guides candidate selection and evaluates completed trajectories against the context. We instantiate this as \rc{} (\textbf{L}LM \textbf{A}s \textbf{F}orecasting \textbf{P}lanner), a training-free framework that bridges the modality gap without retraining either model, using Monte Carlo tree search (MCTS) over the forecast horizon with a \emph{Ranker} LLM as policy and a \emph{Judge} LLM as value function. Experiments on Context-is-Key and Time-MMD across two TSFM backbones (Chronos and TimesFM) and four LLMs show that \rc{} delivers consistent improvements across model choices, supporting sequential search as an effective training-free approach to text-conditioned forecasting.
Huu Hiep Nguyen, Dung Nguyen, Minh Hoang Nguyen +2
Applied Artificial Intelligence Initiative Deakin University Geelong, Australia