Time Series Language Models (TSLMs) offer a promising path toward time series understanding by reasoning over temporal signals and producing natural language answers and explanations. A common approach is Chain-of-Thought (CoT), which generates step-by-step rationales linking relevant signal patterns to final answers. Although these models learn from reference CoT traces during post-training, generating faithful descriptions of input time series at inference remains challenging. Expressing high-dimensional, continuous temporal representations in discrete language tokens may cause the model to neglect task-relevant patterns or describe them inaccurately. Because later reasoning steps build on these descriptions, early errors propagate, leading to incorrect answers with plausible explanations that are inconsistent with the input signal. We propose Lapras (Latent Post-trained Reasoning Across Series), a post-training framework that equips TSLMs with latent reasoning. A model trained with Lapras reasons through a sequence of continuous thoughts in the joint time series-language space, producing text only for the final answer. It learns this through teacher-student self-distillation, where a teacher trained on CoT reference traces reasons explicitly through text. The student aligns its hidden states with the teacher's at the answer stage, transferring the teacher's reasoning ability into its latent computation. We evaluate Lapras across four TSLM backbones on five time series question answering benchmarks. Lapras improves average F1 by up to 10.79% over explicit CoT while generating 23.9x fewer tokens. Lapras's continuous thoughts can also be decoded into readable reasoning traces via standard language decoding, preserving textual explanations. Together, these results highlight Lapras as a promising post-training paradigm for efficient, effective, and interpretable TSLM reasoning.
Figures & tables
Figure 1: Three reasoning paradigms for TSLMs. (i) Direct answering (No-CoT) maps the input directly to a textual answer without generating intermediate reasoning steps. (ii) Explicit reasoning (CoT) expresses intermediate reasoning as discrete text tokens and conditions on these steps to produce the answer. (iii) Latent reasoning (Lapras) performs intermediate reasoning through continuous thoughts in latent space and decodes text only for the final answer.
Figure 2: Early answering with predicted and reference CoT. A substantial accuracy gap between predicted and reference traces is evident by the end of signal description and persists through subsequent reasoning. Mean and standard deviation are computed over the full test set of each benchmark.
Figure 3: Training framework of Lapras. The student (left) produces K continuous thoughts between ⟨bot⟩ and ⟨eot⟩ , feeding each hidden state back as the next input embedding without emitting any token, and predicts the answer under Lansstudent . The teacher (right) writes a full CoT (e.g., [xi,xi+1,…,xi+j] ) and predicts the answer under LCEteacher . A distillation loss LKD aligns their hidden states at the anchor position (the first answer token). The teacher is discarded at inference.
ECG
Sleep
HAR
TSR
Engine
Avg.
Model
Finetune Method
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
Δ
F1
Δ
ChatTS-7B
No-CoT
81.65
81.63
87.65
78.05
69.42
63.41
66.54
66.50
27.98
27.92
66.65
63.50
CoT
59.25
58.56
79.63
65.48
72.79
67.27
61.90
61.87
35.23
35.49
61.76
(-4.89)
57.73
(-5.77)
Latent Reasoning:
iCoT
76.83
76.76
87.00
78.54
73.74
69.55
67.20
67.16
33.68
32.81
67.69
(+1.04)
64.96
(+1.46)
COCONUT
84.45
84.41
86.67
77.31
73.52
68.36
65.58
65.56
36.27
35.24
69.30
(+2.65)
66.18
(+2.67)
Table 1: Main Result. Accuracy and Macro-F1 (%) of five finetuning methods on four TSLMs across five reasoning tasks. Our method is shaded in green . Bold and underline denote the best and second-best per column within each model. Δ reports the change in average relative to No-CoT.
Figure 4: Case study of decoded continuous thoughts of Lapras. Lapras’s continuous thoughts yield readable, task-relevant text whose content can be compared with the reference rationale. The decoded steps progress from a signal description of the series, through inferences built on it, to the final answer . Highlight colors match those in the figure.
Figure 5: Answer probability across reasoning steps on the SLIP-1B backbone for Lapras (top) and CoT (bottom), shown for one representative example per task. Curves show the cumulative probability of Answer in top-1 and Answer in top-2 . Lapras keeps a second candidate alive well into the chain and commits near the end, whereas CoT tends to commit from the first step.
Figure 6: Inference efficiency of the three reasoning methods. Generated tokens and wall-clock time per sample from SLIP-1B with one NVIDIA A6000.
ECG
Sleep
HAR
TSR
Engine
Avg.
Δ
K=1
81.65
78.22
75.18
54.86
34.20
64.82
K=3
81.65
77.36
75.05
64.44
33.16
66.33
(+1.51)
K=6
82.74
78.12
74.84
67.32
38.86
68.38
(+3.56)
K=8
82.43
77.36
74.85
67.39
38.34
68.07
(+3.25)
Table 2: Number of continuous thoughts. Answer accuracy (%) with SLIP-1B, varying only the number of continuous thoughts K . Default K=6 in green . Δ reports the change in average relative to K=1 .
Method
ECG
Sleep
HAR
TSR
Engine
Avg.
Δ
Lapras (SLIP-1B) (Ours)
82.74
78.12
74.84
67.32
38.86
68.38
w/o projection π
80.87
77.68
73.60
50.00
29.02
62.23
(-6.15)
w/o reasoning distillation
78.54
77.25
74.45
56.23
35.23
64.34
(-4.04)
w/o mask
79.41
78.44
72.18
69.17
33.68
66.58
(-1.80)
Table 3: Lapras ablation experiments with the SLIP-1B backbone on the five reasoning benchmarks. We report answer accuracy (%). Default settings are marked in green .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Test
Channels
Len
Format
ECG ( Oh et al., 2026 )
5,800
643
12
1000
Binary (Yes/No)
HAR ( Langer et al., 2026 )
68,542
8,222
3
128
Multiple choice
Sleep ( Langer et al., 2026 )
7,413
923
1
1500
Multiple choice
TSR ( Merrill et al., 2024 )
22,566
4,094
1
128–1020
Multiple choice
Engine ( Wang et al., 2025b )
2,305
193
33
600
Multiple choice
Appendix
Table 4: Overview of datasets.
Backbone
LM
Sensor encoder
Connector
∣Hs∣
Fusion (depth)
ChatTS-7B
7B
per-patch MLP
none (native space)
C⌈L/P⌉
concat (all layers)
ITFormer-0.5B
0.5B
PatchTST
question-cond. resampler
25
concat (all layers)
SLIP-1B
1B
language-pretrained
attention pooler
64
cross-attn (last 4 )
OpenTSLM-1B
1B
PatchTST
Perceiver resampler
64
cross-attn (every layer)
Appendix
Table 5: The four backbones along three design axes. ∣Hs∣ is the number of sensor tokens handed to the language model; C is the channel count, L the series length, and P the patch size of Table 7 . Concat places Hs in the token sequence and reads it with self-attention; cross-attn keeps it outside and injects it through dedicated layers.
ECG
Sleep
HAR
TSR
Engine
Avg.
Model
Finetune Method
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
Δ
F1
Δ
OpenTSLM-SoftPrompt
No-CoT
61.59
60.98
84.29
73.93
73.46
69.85
49.15
48.65
31.09
30.64
59.92
56.81
CoT
52.10
51.94
83.42
72.87
71.14
67.00
53.03
53.02
32.12
32.11
58.36
(-1.55)
55.39
(-1.42)
iCoT
63.76
63.62
84.07
72.10
72.55
69.34
54.86
54.85
26.42
18.03
60.33
(+0.41)
55.59
(-1.22)
COCONUT
65.16
64.48
83.42
71.25
74.18
71.18
56.74
56.50
34.72
33.83
62.84
(+2.93)
59.45
(+2.64)
Lapras (Ours)
65.16
64.80
85.81
75.62
74.25
70.94
58.04
58.00
37.31
36.08
64.11
(+4.20)
61.09
(+4.28)
Appendix
Table 6: Effect of the fusion pathway. Accuracy and Macro-F1 (%) on the two OpenTSLM variants, which share a 1 B language model and differ only in fusion. Ours in green ; bold and underline are best and second-best per column within a variant; Δ is the change in average over No-CoT.
Backbone
Dataset
Patch Size
Epochs
lr
wd
λ
ChatTS-7B
ECG
32
20
1e-5
1e-2
10
Sleep
8
20
1e-5
1e-2
10
HAR
8
4
1e-5
1e-2
10
TSR
8
4
1e-5
1e-2
1
Engine
50
5
1e-5
1e-2
20
SLIP-1B
ECG
32
20
5e-5
1e-2
10
Appendix
Table 7: Per-task hyperparameters for the four backbones and the OpenTSLM-SoftPrompt variant. One row covers all five finetuning methods, which share every setting except the distillation weight λ of Lapras.
Figure 7: Per-step logit-lens readout on one TSR example. At each step, we plot the next-token distribution read through the logit lens, with tokens ranked by probability and colored on a log scale, together with its entropy (blue). Top: under CoT the distribution is sharply peaked at nearly all of the 140 steps. Bottom: under latent reasoning the 6 continuous thoughts spread probability over many tokens ( Neff≈22 ), entropy decreases after the first thought, and the distribution reaches the commitment point ( top-1≥0.95 ) only at the first answer token.
Figure 8: Attention maps of ChatTS-7B on TSR at layers 0 , 21 , and 35 for (a) No-CoT, (b) explicit chain-of-thought, and (c) latent reasoning. Color gives attention relative to the uniform level, red above and blue below, and the bars on the two axes mark the time series , question , reasoning , and answer spans.
Layer
No-CoT
CoT
Lapras
0 (first)
2.95
3.39
17.35
21 (middle)
11.18
4.65
6.54
35 (last)
0.72
0.27
1.21
Appendix
Table 8: Attention on the time series with ChatTS-7B. Allocation λts (%) is the share of a reasoning token’s attention that lands on the time series tokens, at the first, middle, and last layers.
Figure 9: Verbalization bottleneck in CoT reasoning. An operational decision over a temperature series with four choices. CoT (bottom right) runs a CoT trace that first verbalizes the signal as text, then reasons from that textual summary rather than the original representation, and reaches an incorrect final answer. Spans are colored grounded , correct but insufficient , and incorrect .
Figure 10: Early-answering accuracy under predicted and reference CoT. Accuracy of a frozen CoT-SFT model answering from only the first k CoT sentences, using either its own predicted CoT or the reference annotation. The verbalization bottleneck Δ is the accuracy difference between the two at the end of the describe stage. Shaded regions mark the describe and reason stages.
Figure 11: Case study of Lapras’s interpretability , obtained by decoding each of the six continuous thoughts z1,…,z6 into text and placing it beside the explicit chain. The decoded steps progress from a global description of the series, through local descriptions of individual segments, to inferences that combine them, and finally to the answer . Highlight colors match the box fills in the figure.