Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.
Figures & tables
Figure 1 : Flow-based Latent Reasoning (FLaRe). (a) A VAE encodes a symbolic CoT rsym into a small code z , corrupted in training for a smooth latent space, and its decoder reads it back as rsym or, given the question q , as a natural language CoT rlang . (b) In stage 1, the shared flow model θ learns to denoise the code given the question and to answer from its own thought z^ . (c) In stage 2, on new questions with answers only, θ proposes K thoughts per question, and those whose decoded answer is correct are re-encoded as flow targets. The model is updated with the flow loss on these targets and an answer loss Lroll backpropagated through the full denoising path.
Figure 2Table 3Figure 4Table 5
Part of the loop
Variant
Direct
Decoded
Stage 1
55.3
61.3
Stage 2
59.1
62.6
Training
No online flow loss
59.0
61.6
No stage 1 replay
59.0
62.5
Targets
No verification
57.1
60.7
One trace
58.8
61.8
Table 6 : Stage 2 design (%). Each ablation changes one part of the loop.
Useful
Diverse
Explainable
Refinable
Efficient
Method
follows
pass@16
vote
values in order
0.1→0.5× CoT compute
speedup at % CoT acc.
Explicit CoT
97.2
–
–
–
–
1.0× at 100%
Coconut
7.3
+4.6
−1.5
19.1
+2.8
2.6× at 61%
CODI
31.8
+3.5
−0.7
38.5
+6.1
2.0× at 94%
PCCoT
27.7
+0.8
−1.9
7.8
−10.5
3.9× at 91%
FLaRe, stage 1
39.7
+16.6
+3.3
46.0
+0.4 / +18.2
3.9× at 93%
Table 7 : The five requirements on GSM8K. Useful: twin pairs (%) whose answer follows the injected thought. Diverse: gain over greedy with 16 samples. Explainable: questions (%) whose decoded thought has all reference intermediate results in order. Refinable: gain from 0.1 to 0.5× explicit CoT compute. Efficient: speedup over explicit CoT at the share of its accuracy. FLaRe: decoded reading, except refinable (direct / decoded) and efficient (direct, S=2 ).
Qwen2.5-0.5B-Instruct
Llama-3.2-1B-Instruct
Llama-3.2-3B-Instruct
IID
OOD
IID
OOD
IID
OOD
Method
GSM8K
Hard
SVAMP
MArith
GSM8K
Hard
SVAMP
MArith
GSM8K
Hard
SVAMP
MArith
Explicit CoT
59.7
17.0
61.6
98.3
62.5
14.7
66.5
97.8
72.4
20.8
73.4
98.3
No-CoT
27.3
6.4
38.5
57.8
35.0
7.4
38.1
67.8
39.2
9.3
56.7
96.7
Coconut
28.2
6.8
41.9
58.3
36.1
8.3
48.4
87.2
45.3
10.9
57.5
96.7
iCoT
32.9
7.4
41.0
66.7
36.5
8.4
40.8
72.8
42.1
9.6
50.7
91.1
Table 8 : Direct answer accuracy (%) of models trained on GSM8K-Aug, on GSM8K (IID) and on GSM8K-Hard, SVAMP and MultiArith (OOD). Bold: best latent method per column and scale. Baselines come from the original papers when reported, else from our runs (Appendix A.5 ).
Method
Family
Thought lives in
Produced by
Trained with
Decodable
Budget
Explicit CoT
Horizontal
Tokens
AR decoding
CoT SFT
Written out
Model-chosen
iCoT
Vertical
Hidden states
Single pass
CoT curriculum
No
Fixed
Pause tokens
Vertical
Hidden states
Single pass
Answer loss
No
Fixed
Coconut
Horizontal
Embeddings
AR feedback
CoT curriculum
LM head
Fixed
CODI
Horizontal
Embeddings
AR feedback
CoT distillation
LM head
Fixed
PCCoT
Parallel
Embeddings
Jacobi iteration
CoT distillation
LM head
Fixed
Table 9 : Latent reasoning methods along the design axes of the paper.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
VAE
The VAE is initialized from a model trained for one epoch on the natural language CoTs of OMI-2 ( Toshniwal et al., 2025 ) with 32 slots and a single natural language route. It is then reduced to M=8 slots at initialization and trained with the dual objective of equation 3 with both terms weighted equally. Token substitution draws the replacement uniformly from the vocabulary with clean decoder targets, and the latent noise and dropout act on the sampled code before decoding. The statistics (m,s) are the mean and standard deviation of the posterior means of 20K training rows.
Backbones ϕ , ψ
Llama-3.2-1B / 3B, fine-tuned
Token substitution psub
0.3 on the encoder input
Initialization
OMI-2 language VAE, 32×512
Latent noise
VP, δ=0.7 , on half of the codes
Latent code M×d
8×512
Latent dropout pdrop
0.4
Input, decoder routes
rsym , dual routes, equal weights
Optimizer
AdamW, (0.9,0.98) , wd 10−5 , clip 1.0
Training data
paired GSM8K-Aug, 367K rows
Learning rate
10−4 encoder, 2⋅10−5 decoder
Appendix
Table 10 : Implementation details of the final recipe: the VAE, the flow model of stages 1 and 2, the inference and the evaluation protocol.
Figure 5 : Noise draws per code. GSM8K accuracy (%) over training for one, two, four and eight draws per code, under the direct and decoded readings, for the EMA weights (solid, filled markers) and the LIVE weights (dashed, open markers).
Method
Qwen2.5-0.5B-Instruct
Llama-3.2-1B-Instruct
Llama-3.2-3B-Instruct
Explicit CoT
T
T
R (a)
No-CoT
T
T
T
iCoT
T
T
T
Coconut
T
R (b)
T
CODI
T
P
T
PCCoT
T
P on GSM8K, R (c) on the OOD sets
T
Appendix
Table 11 : Source of each baseline number of Tab. 8 . P: printed by the paper that introduced the method, CODI ( Shen et al., 2025 ) , PCCoT ( Wu et al., 2025 ) or KaVa ( Kuzina et al., 2026 ) . T: trained by us, with LaDiR reimplemented as in Appendix A.6 . R: released weights evaluated by us, the public explicit CoT checkpoint yingfanbot/gsm-cot-llama3b (a), and those of Dilgren and Wiegreffe (2026) (b) and Wu et al. (2025) (c).
Method
Tuning
Learning rate
Epochs
Other settings
Explicit CoT
LoRA, r=128 , α=32
8⋅10−4 / 8⋅10−4 / –
10
batch 128, cosine, 3% warmup, wd 0.1
No-CoT
LoRA, r=128 , α=32
2⋅10−4 / 8⋅10−4 / 2⋅10−4
10
as explicit CoT, answer only
iCoT
full, single precision
10−5
20
batch 32, 8 CoT tokens removed per epoch (Qwen: 11)
Coconut
full, bf16
5⋅10−5
3 + 6 + 1
c=1 , a CoT stage, stages 1 to 6 and a fully latent stage, batch 128
CODI
LoRA, r=128 , α=32
8⋅10−4 / – / 3⋅10−4
10 / – / 8
6 latents, distillation weight 20, batch 128
PCCoT
LoRA, r=128 , α=32
5⋅10−4 / – / 2⋅10−4
10
24 latents, T=3 iterations, batch 128
Appendix
Table 12 : Recipes of our baseline runs on GSM8K-Aug. Learning rates are for Qwen2.5-0.5B / Llama-3.2-1B / Llama-3.2-3B.
Corpus
Origin
Rows
Questions
GSM8K-Aug, flow
whynlp/gsm8k-aug train
385K
385K
GSM8K-Aug, paired
aligned with whynlp/gsm8k-aug-nl
367K
367K
Diverse CoT
Qwen2.5-32B-Instruct on GSM8K-Aug questions
216K
148K
Diverse CoT, paired
same, with a verified natural language twin
205K
142K
IID data
OMI-2 augmented_gsm8k , converted
166K
39K
OOD data
OMI-2 augmented_math , converted
463K
80K
Appendix
Table 13 : Corpora used in the paper. Rows are (question, CoT, answer) triples and questions are distinct normalized question texts.
VAE training data
Recon.
Direct
Decoded
GSM8K-Aug CoT
99.1
50.7
56.9
+ Diverse CoT
99.0
48.8
54.8
+ IID data
99.0
51.6
56.9
+ OOD data
99.3
50.9
56.0
+ IID and OOD data ( 5× budget)
99.6
51.1
58.3
Appendix
Table 14 : Training data of the VAE, with the flow data fixed. Reconstruction (%) and GSM8K accuracy (%) under the direct and decoded readings.
IID pair, from the augmented_gsm8k share
Question
Lucy has orange trees that produce 4 oranges each. If an orange can be sold for 2andLucywantstoearn80, how many orange trees does she need to harvest?
rlang
Each orange tree produces 4 oranges. Each orange can be sold for 2.SotheamountofmoneyLucyearnsfromeachtreeis4∗2 = 8.IfLucywantstoearn80, then she needs to harvest 80/8 = 10 trees. Thus, Lucy needs to harvest 10 orange trees.
rsym
<< 4*2=8 >> << 80/8=10 >> << 10 >>
a
10
OOD pair, from the augmented_math share
Question
The greatest common divisor of a , 12 , and 21 is 3 . What is the smallest possible value of a that is greater than 3 ?
Appendix
Table 15 : Two published pairs, one IID and one OOD, with the natural language CoT verbatim from OMI-2 and the symbolic CoT written by the converter.
Model
Follows
Keeps original
Neither
Clean follow
Explicit CoT (reference steps)
97.2
0.2
2.5
99.9
CODI
31.8
19.8
48.4
68.4
PCCoT
27.7
20.8
51.6
62.3
Coconut
7.3
29.5
63.2
19.2
FLaRe, stage 1
39.7
8.7
51.6
91.4
FLaRe, stage 2
37.3
15.0
47.7
76.1
Appendix
Table 16: Twin test on 824 pairs. Share of pairs (%) whose answer follows the injected thought (the twin’s answer), keeps the original answer, or is neither. Clean follow: the share that follows, counted only on the pairs where the model answers both the question and its twin correctly without any injection.
Figure 6 : pass@ k on the GSM8K test set, estimated from 32 samples per question with the unbiased estimator of Chen et al. (2021) . Our models generate flow samples from noise, the baselines add noise to their thought at the calibrated scale, and explicit CoT samples its tokens at temperature 0.7.
Model
σ
Greedy
pass@1
pass@ K
Majority
Distinct ans.
Committed err.
CODI
0.5
55.6
54.6
59.1
54.9
1.46
51%
PCCoT
1.0
54.1
51.9
54.9
52.2
1.31
58%
Coconut
2.0
36.0
33.8
40.6
34.5
1.89
34%
FLaRe, stage 1, decoded
–
61.2
60.1
77.8
64.4
3.41
6%
FLaRe, stage 1, direct
–
55.3
–
67.1
55.6
2.68
–
FLaRe, stage 2, decoded
–
62.6
62.0
77.3
64.4
2.89
9%
Appendix
Table 17: Sampling K=16 thoughts per question: accuracies (%), distinct answers per question, and committed errors as a share of the greedy errors.
Table 20
Figure 7 : Accuracy against the latent computation spent at inference: thoughts for Coconut and CODI, Jacobi iterations for PCCoT, loops for the depth-recurrent model, and Euler steps for our flow model, on GSM8K-Aug and GSM8K-Aug-NL. The dotted lines mark the training budget.