Chain-of-thought (CoT) reasoning often improves language-model performance by giving models additional computation before answering. However, explicit CoT expresses this computation as a sequence of autoregressively generated tokens. Latent reasoning replaces these tokens with compact continuous states, but most autoregressive latent-reasoning methods retain a left-to-right dependency among latent vectors. We introduce LLoCoT: a looped latent-reasoning framework that replaces left-to-right latent generation with iterative refinement of a compact latent workspace. A shared transformer is reapplied for a small number of refinement iterations, jointly updating the latent slots based on the prompt and the evolving workspace state. Using the refined state, a probabilistic head predicts a distribution from which latent tokens are sampled in parallel and used to condition an autoregressive decoder for answer generation. Training uses continuous representations derived from explicit CoT together with a final-answer prediction loss and likelihood-based supervision of the latent states. Across HumanEval and MBPP, LLoCoT achieves the highest mean among the evaluated methods, performing on par in accuracy with Reasoning SFT, our explicit-CoT baseline, while outperforming the base model, answer-only SFT and NF-CoT. Relative to Reasoning SFT, LLoCoT reduces time to the first answer token by approximately 36× and reasoning-phase latency by approximately 42×, while increasing end-to-end throughput by 9.2%. This design replaces serial thought generation with parallel latent-slot refinement while retaining probabilistic latent modeling and autoregressive answer decoding.
Figures & tables
Figure 1: Overview of the LLoCoT architecture. A shared transformer repeatedly refines K loop hidden states, with each non-final Gaussian mean providing feedback to the next loop. The final Gaussian head generates all final latent thoughts jointly for autoregressive answer decoding; VAE-derived targets and auxiliary intermediate supervision are used only during training.
Mode
HumanEval
HumanEval+
MBPP
MBPP+
Mean
Qwen3-1.7B-Base
59.9±1.2
52.6±1.3
63.3±0.8
53.4±0.8
57.3 (+0.0)
Base SFT
68.6±1.1
63.9±1.1
63.4±0.7
55.0±0.6
62.7 (+5.4)
Reasoning SFT
70.3±1.0
64.4±1.0
63.5±0.7
54.7±0.7
63.2 (+5.9)
NF-CoT
69.2±1.1
63.5±1.1
62.2±0.7
52.9±0.7
62.0 (+4.7)
No-feedback LLoCoT, R=2
69.3±1.1
63.8±1.1
64.3±0.7
55.0±0.6
63.1 (+5.8)
LLoCoT
69.3±1.0
63.9±1.0
65.1±0.6
56.2±0.6
63.6 (+6.3)
Table 1: Main code-generation results. Values are percentages. Benchmark entries are point estimate ± the 95% fixed-task sampling-interval half-width. The mean is the unweighted arithmetic mean of the four benchmark point estimates. Parenthetical values in the Mean column are changes relative to Qwen3-1.7B-Base. Best and second-best available entries are shown in bold and underlined , respectively.
Mode
HumanEval
HumanEval+
MBPP
MBPP+
Mean
No-feedback LLoCoT, R=1
69.1±1.0
63.3±1.1
64.4±0.7
54.9±0.6
62.9 (-0.2)
No-feedback LLoCoT, R=2
69.3±1.1
63.8±1.1
64.3±0.7
55.0±0.6
63.1 (+0.0)
No-feedback LLoCoT, R=3
69.4±1.0
64.0±1.1
64.3±0.6
55.0±0.6
63.2 (+0.1)
No-feedback LLoCoT, R=4
68.2±1.0
63.1±1.0
63.8±0.7
54.8±0.6
62.5 (-0.6)
No-feedback LLoCoT, R=6
68.6±1.1
63.6±1.1
60.4±0.7
52.0±0.6
61.2 (-1.9)
No-feedback LLoCoT, R=8
68.0±1.0
63.0±1.0
62.2±0.6
53.6±0.6
61.7 (-1.4)
Table 2: Loop-depth ablation for No-feedback LLoCoT. Values are percentages and report point estimate ± the 95% fixed-task sampling-interval half-width. The mean is the arithmetic average of the four benchmark point estimates. Parenthetical values in the mean column are changes relative to No-feedback LLoCoT, R=2 . Green, red, and gray indicate positive, negative, and zero changes, respectively.
Mode
HumanEval
HumanEval+
MBPP
MBPP+
Mean
No-feedback LLoCoT, R=2
69.3±1.1
63.8±1.1
64.3±0.7
55.0±0.6
63.1 (+0.0)
+ Answer-mean
61.7±1.2
55.3±1.2
60.2±0.7
51.6±0.7
57.2 (-5.9)
+ Auxiliary-NLL loss
68.1±1.1
62.6±1.0
65.0±0.7
56.0±0.6
62.9 (-0.2)
+ Feedback-mean + auxiliary NLL (LLoCoT)
69.3±1.0
63.9±1.0
65.1±0.6
56.2±0.6
63.6 (+0.5)
+ Feedback-sampled + auxiliary NLL
67.9±1.1
62.3±1.0
64.2±0.6
55.3±0.6
62.4 (-0.7)
Table 3: Feedback, conditioning, and auxiliary-supervision ablations. Values are percentages and report point estimate ± the 95% fixed-task sampling interval half-width. The mean is the arithmetic average of the four benchmark point estimates. Parenthetical values in the mean column are changes relative to No-feedback LLoCoT, R=2 . Green, red, and gray indicate positive, negative, and zero changes, respectively.
Metric
Base model
Reason. SFT
NF-CoT
LLoCoT
Time to first answer token (s) ↓
0.03
3.17
2.01 ( ×0.6 )
0.09 ( ×0.03 )
End-to-end latency (s) ↓
22.59
15.16
13.73 ( ×0.9 )
11.47 ( ×0.8 )
Serial reasoning/latent phase (s) ↓
0
2.47
1.98 ( ×0.8 )
0.06 ( ×0.02 )
Conditioning/selection phase (s) ↓
0
0.70
0.03 ( ×0.04 )
0.03 ( ×0.04 )
Peak allocated (GiB) ↓
3.4
3.3
6.3 ( ×1.9 )
6.3 ( ×1.9 )
Reasoning/latent forward calls ↓
1.0
81.5
64.0 ( ×0.8 )
2.0 ( ×0.02 )
Table 4: Sampled inference-efficiency measurements. Values are point estimates. Parenthetical multiplicative ratios for NF-CoT and LLoCoT are relative to Reasoning SFT. Ratios are rounded to the displayed precision; values below 0.1 retain one significant digit. Best and second-best values in each row are shown in bold and underlined, respectively. Colors indicate favorable, unfavorable, and neutral changes relative to Reasoning SFT.