Chain-of-thought (CoT) reasoning often improves language-model performance by giving models additional computation before answering. However, explicit CoT expresses this computation as a sequence of autoregressively generated tokens. Latent reasoning replaces these tokens with compact continuous states, but most autoregressive latent-reasoning methods retain a left-to-right dependency among latent vectors. We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace. A shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Using the refined state, a probabilistic head predicts a distribution from which latent tokens are sampled in parallel and used to condition an autoregressive decoder for answer generation. Training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states. Across HumanEval and MBPP, LLoCoT achieves the highest mean among the evaluated methods, performing on par in accuracy with Reasoning SFT, our explicit-CoT baseline, while outperforming the base model, answer-only SFT and NF-CoT. Relative to Reasoning SFT, LLoCoT reduces time to the first answer token by approximately 36× and reasoning-phase latency by approximately 42×, while increasing end-to-end throughput by 9.2%. This design replaces serial thought generation with parallel latent-slot refinement while retaining probabilistic latent modeling and autoregressive answer decoding.
Figures & tables
Figure 1: Overview of the LLoCoT architecture. A shared transformer repeatedly refines K loop hidden states, with each non-final Gaussian mean providing feedback to the next loop. The final Gaussian head generates all final latent thoughts jointly for autoregressive answer decoding; VAE-derived targets and auxiliary intermediate supervision are used only during training.
Mode
HumanEval
HumanEval+
MBPP
MBPP+
Mean
Qwen3-1.7B-Base
59.9±1.2
52.6±1.3
63.3±0.8
53.4±0.8
57.3 (+0.0)
Base SFT
68.6±1.1
63.9±1.1
63.4±0.7
55.0±0.6
62.7 (+5.4)
Reasoning SFT
70.3±1.0
64.4±1.0
63.5±0.7
54.7±0.7
63.2 (+5.9)
NF-CoT
69.2±1.1
63.5±1.1
62.2±0.7
52.9±0.7
62.0 (+4.7)
No-feedback LLoCoT, R=2
69.3±1.1
63.8±1.1
64.3±0.7
55.0±0.6
63.1 (+5.8)
LLoCoT
69.3±1.0
63.9±1.0
65.1±0.6
56.2±0.6
63.6 (+6.3)
Table 1: Main code-generation results. Values are percentages. Benchmark entries are point estimate ± the 95% fixed-task sampling-interval half-width. The mean is the unweighted arithmetic mean of the four benchmark point estimates. Parenthetical values in the Mean column are changes relative to Qwen3-1.7B-Base. Best and second-best available entries are shown in bold and underlined , respectively.
Mode
HumanEval
HumanEval+
MBPP
MBPP+
Mean
No-feedback LLoCoT, R=1
69.1±1.0
63.3±1.1
64.4±0.7
54.9±0.6
62.9 (-0.2)
No-feedback LLoCoT, R=2
69.3±1.1
63.8±1.1
64.3±0.7
55.0±0.6
63.1 (+0.0)
No-feedback LLoCoT, R=3
69.4±1.0
64.0±1.1
64.3±0.6
55.0±0.6
63.2 (+0.1)
No-feedback LLoCoT, R=4
68.2±1.0
63.1±1.0
63.8±0.7
54.8±0.6
62.5 (-0.6)
No-feedback LLoCoT, R=6
68.6±1.1
63.6±1.1
60.4±0.7
52.0±0.6
61.2 (-1.9)
No-feedback LLoCoT, R=8
68.0±1.0
63.0±1.0
62.2±0.6
53.6±0.6
61.7 (-1.4)
Table 2: Loop-depth ablation for No-feedback LLoCoT. Values are percentages and report point estimate ± the 95% fixed-task sampling-interval half-width. The mean is the arithmetic average of the four benchmark point estimates. Parenthetical values in the mean column are changes relative to No-feedback LLoCoT, R=2 . Green, red, and gray indicate positive, negative, and zero changes, respectively.
Mode
HumanEval
HumanEval+
MBPP
MBPP+
Mean
No-feedback LLoCoT, R=2
69.3±1.1
63.8±1.1
64.3±0.7
55.0±0.6
63.1 (+0.0)
+ Answer-mean
61.7±1.2
55.3±1.2
60.2±0.7
51.6±0.7
57.2 (-5.9)
+ Auxiliary-NLL loss
68.1±1.1
62.6±1.0
65.0±0.7
56.0±0.6
62.9 (-0.2)
+ Feedback-mean + auxiliary NLL (LLoCoT)
69.3±1.0
63.9±1.0
65.1±0.6
56.2±0.6
63.6 (+0.5)
+ Feedback-sampled + auxiliary NLL
67.9±1.1
62.3±1.0
64.2±0.6
55.3±0.6
62.4 (-0.7)
Table 3: Feedback, conditioning, and auxiliary-supervision ablations. Values are percentages and report point estimate ± the 95% fixed-task sampling interval half-width. The mean is the arithmetic average of the four benchmark point estimates. Parenthetical values in the mean column are changes relative to No-feedback LLoCoT, R=2 . Green, red, and gray indicate positive, negative, and zero changes, respectively.
Metric
Base model
Reason. SFT
NF-CoT
LLoCoT
Time to first answer token (s) ↓
0.03
3.17
2.01 ( ×0.6 )
0.09 ( ×0.03 )
End-to-end latency (s) ↓
22.59
15.16
13.73 ( ×0.9 )
11.47 ( ×0.8 )
Serial reasoning/latent phase (s) ↓
0
2.47
1.98 ( ×0.8 )
0.06 ( ×0.02 )
Conditioning/selection phase (s) ↓
0
0.70
0.03 ( ×0.04 )
0.03 ( ×0.04 )
Peak allocated (GiB) ↓
3.4
3.3
6.3 ( ×1.9 )
6.3 ( ×1.9 )
Reasoning/latent forward calls ↓
1.0
81.5
64.0 ( ×0.8 )
2.0 ( ×0.02 )
Table 4: Sampled inference-efficiency measurements. Values are point estimates. Parenthetical multiplicative ratios for NF-CoT and LLoCoT are relative to Reasoning SFT. Ratios are rounded to the displayed precision; values below 0.1 retain one significant digit. Best and second-best values in each row are shown in bold and underlined, respectively. Colors indicate favorable, unfavorable, and neutral changes relative to Reasoning SFT.
Language models typically reason via explicit chain-of-thought (CoT), generating intermediate steps token-by-token. Latent CoT offers an alternative: it performs multi-step reasoning in the model's hidden states, replacing decoded tokens with continuous representations for greater efficiency. However, existing latent CoT methods underperform explicit CoT beyond 1B parameters, and the gap widens with scale. Looped, or recurrent-depth, Transformers, which reuse their weights to increase computation depth without adding parameters, are a natural fit for latent reasoning. We therefore ask whether looped Transformers can bridge this gap. We answer affirmatively with a simple recipe: a looped padded Transformer that processes K latent blocks in parallel for R iterations, with a cross-entropy loss on each latent position's gold CoT-step token, similar to explicit CoT supervision. We instantiate it as LOTUS (Looped Transformers with parallel supervision on latents). LOTUS is, to our knowledge, the first latent-CoT method to bridge the gap to explicit CoT at the 3B scale, while cutting thought-phase latency by 2.5x-6.9x from compact math expressions to natural language. Projecting LOTUS's post-loop latents through the base LM head recovers the gold reasoning steps and even surfaces alternative valid intermediate steps, evidence that its latent space is interpretable and CoT-aligned. Ablations confirm that both the looped backbone and the parallel supervision on gold CoT tokens are essential. Code is available at https://github.com/yingfan-bot/lotus.
Chain-of-thought (CoT) prompting improves reasoning in large language models (LLMs) by externalizing intermediate computation as discrete text tokens, but this textual interface also introduces redundancy and inference overhead. Latent reasoning offers a promising alternative by carrying part of the computation in continuous representations. However, existing methods typically predefine when latent computation is invoked and how it is allocated during decoding, leaving a key problem unresolved: when to invoke latent computation, what type of computation to perform, and how much budget to allocate. We propose \textbf{Ty}ped \textbf{L}at\textbf{e}nt \textbf{R}easoning (Tyler), a typed and budget-aware framework for latent reasoning during autoregressive decoding. Tyler learns a policy that, at each decoding step, chooses between emitting a text token and switching to a latent computation module specialized for a particular reasoning function. Once invoked, an operator maps the current reasoning state into latent tokens that support global planning, local state updates, or reusable procedural abstraction. Across extensive experiments on three backbone LLMs, Tyler improves accuracy by up to 14.49 points over CoT and by up to 4.30 points over the strongest competing baseline. It further generalizes across diverse reasoning domains and achieves the best final-stage performance with the lowest forgetting.
Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning step must be verbalized before the model can proceed, even when the underlying update is semantic, uncertain, or only partially formed. Latent reasoning offers a higher-bandwidth alternative by performing intermediate computation in compact continuous states before committing to text. Yet existing latent-reasoning methods often sacrifice key advantages that make CoT effective in autoregressive language models, including native left-to-right generation, probabilistic sampling, compatibility with KV-cache decoding, and tractable likelihood estimation. We propose NF-CoT, a latent reasoning framework that preserves these advantages by modeling continuous thoughts with normalizing flows. NF-CoT instantiates a TARFlow-style normalizing flow inside the LLM backbone, defining a tractable probability model over compact continuous thoughts distilled from explicit CoT. Continuous-thought positions are generated by an NF head, while text positions are generated by the standard LM head within the same causal stream. This design provides exact likelihoods for latent thoughts, enables probabilistic left-to-right decoding with the original KV cache, and supports direct policy-gradient optimization in the latent reasoning space. On code-generation benchmarks, NF-CoT improves pass rates over explicit-CoT and prior latent-reasoning baselines while substantially reducing intermediate-reasoning cost.