Organizations: School of Computer Science and Engineering, Northeastern University, China · Hong Kong Polytechnic University, China · NiuTrans Research, Shenyang, China
Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start is therefore unknown during denoising, and the model has to decide the trace length, the place of every trace token, and the answer at the same time. We propose Plan Canvas to fix the boundary between the trace and the answer. A plan region of fixed capacity holds a compact trace, supervised padding fills its unused positions, and the answer starts at a fixed position. The fixed regions also allow separate denoising clocks for the plan and for the answer. With the trace text, backbone, and canvas length of the free-trace baseline held fixed, Plan Canvas improves accuracy on ProsQA and on Deep ProsQA, a graph benchmark with longer proofs. On Deep ProsQA, accuracy rises from 73.0% to 87.0%, the share of questions answered with a valid path rises from 30.8% to 59.1%, and the gain is largest on the longest proofs.
Figures & tables
Figure 1: Training and sampling use the same ELF backbone with fixed plan and answer regions. Each training example selects either flow matching with region corruption or token cross-entropy with separate decoder corruption. The branch probabilities are 0.8 for flow matching and 0.2 for token cross-entropy in the reported configuration. During sampling, the prompt stays fixed while region clocks control Euler updates of every target position. The model predicts clean states, which define velocity, and switches to token decoding once both regions reach time one.
ProsQA
Deep ProsQA
Layout
Full canvas S
Readout G
Full canvas S
Readout G
Direct answer
720
16
784
16
Full trace
800
96
960
192
Free trace / answer then trace
752
48
848
80
Plan Canvas, either clock interface
752
48
848
80
Table 1: Sequence budgets in token positions, shared by training and evaluation within each target layout. Prompt caps are 704 positions for ProsQA and 768 positions for Deep ProsQA across all layouts. The readout window is computed from these fixed caps and never from a reference answer’s content or length.
ProsQA
Deep ProsQA
Target layout
Accuracy
Valid path
Correct / invalid
Accuracy
Valid path
Correct / invalid
Direct answer
87.6
–
–
57.7
–
–
Full trace
91.0
79.8
12.2
73.5
58.5
20.3
Free trace
91.4
75.0
17.0
73.0
30.8
42.5
Answer then trace
88.2
79.4
8.8
80.2
38.5
41.7
Plan Canvas, single clock
96.8
90.6
6.2
82.1
55.0
27.1
Table 2: Answer accuracy and path checks on the test splits, in percent of all test questions. A valid path starts at a fact about the entity, follows directed edges of the question graph, and ends at the predicted concept. Correct / invalid counts correct answers with an invalid or missing path. Free trace, answer then trace, and both Plan Canvas variants share the trace text, canvas length, and readout window, and shaded rows use the fixed plan region.
fimpus nulpus molpus slilpus rupus chepus, then padding
chepus
Segment clocks
fimpus nulpus molpus slilpus rupus chepus, then padding
chepus
Table 3: Emitted plans for Deep ProsQA test question 450, which asks whether Eva is a chepus or a pripus among 40 stated facts and rules. The underlined pair repeats a concept and then makes a hop that is not an edge of the question graph.
Figure 2: Shape of the emitted paths on the Deep ProsQA test split, in percent of all test questions. Panel (a) counts paths that contain the same concept twice, grouped by shortest proof length. Panel (b) gives the distribution of the emitted path length minus the shortest proof length.
Figure 3: Deep ProsQA test accuracy under three changes. Panel (a) groups the test split by shortest proof length, with 111 or 112 questions per length. Panel (b) evaluates the same checkpoints with 32, 64, 96, and 128 synchronous Euler iterations, counting two guided backbone calls per iteration and one decoder call. Panel (c) trains the segment-clock canvas with different plan capacities, and the dashed line marks the free trace.
Configuration
Fixed
Marker
Prefix
Time
Deep ProsQA
Free trace
–
–
–
–
73.0
Fixed plan region
✓
–
–
–
82.1
+ segment markers
✓
✓
–
–
84.0
+ time-prefix groups
✓
✓
✓
–
85.9
+ position time
✓
✓
✓
✓
87.0
Table 4: Deep ProsQA accuracy in percent as the parts of the segment-clock interface are added under tied training times. Fixed denotes the plan region, and Marker, Prefix, and Time denote segment markers, time-prefix groups, and the position-dependent time term.
Train times
Schedule
(op,oa)
Joint
Calls
ProsQA
Deep
Tied
Synchronous
(0,0)
64
129
95.2
87.0
Tied
Plan-first
(0,64)
128
257
96.0
87.6
Tied
Answer-first
(64,0)
128
257
94.6
85.4
Mixed
Synchronous
(0,0)
64
129
95.2
87.9
Mixed
Plan-lead
(0,32)
96
193
95.4
89.0
Mixed
Plan-first
(0,64)
128
257
95.4
88.7
Table 5: Accuracy in percent under different offsets with the same weights and 64 active steps per region. Joint counts Euler iterations, Calls counts backbone evaluations including the decoder, and mixed training shares region times for half the examples.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Layout
P valid
P test [95% CI]
D valid
D test [95% CI]
Direct answer
88.0
87.6 [84.4, 90.2]
58.2
57.7 [54.6, 60.7]
Full trace
90.3
91.0 [88.2, 93.2]
78.6
73.5 [70.7, 76.1]
Free trace
89.3
91.4 [88.6, 93.5]
71.0
73.0 [70.2, 75.7]
Answer then trace
90.3
88.2 [85.1, 90.7]
78.2
80.2 [77.6, 82.5]
Plan Canvas (single clock)
94.0
96.8 [94.9, 98.0]
82.0
82.1 [79.6, 84.4]
Plan Canvas (segment clocks)
94.3
95.2 [93.0, 96.8]
86.8
87.0 [84.8, 88.9]
Appendix
Table 6: Validation and test answer accuracy in percent, with P denoting ProsQA and D denoting Deep ProsQA. Bracketed values give marginal 95% Wilson intervals for the complete test splits under each fixed checkpoint.
Set
First / second
Δ (pp)
Paired 95% CI
W/L
p
P
Canvas single / Free trace
+5.4
[+3.0, +8.0]
34/7
2.53×10−5
P
Canvas single / Full trace
+5.8
[+3.2, +8.4]
37/8
1.54×10−5
P
Canvas single / Answer then trace
+8.6
[+5.6, +11.6]
52/9
1.80×10−8
P
Segment tied, sync / Canvas single
-1.6
[-3.0, -0.2]
3/11
0.057
P
Segment mixed, sync / Segment tied, sync
+0.0
[-1.4, +1.4]
7/7
1.000
P
Tied, Plan-first / Segment tied, sync
+0.8
[+0.0, +1.8]
5/1
0.219
Appendix
Table 7: Paired test contrasts with first minus second accuracy differences, with P denoting ProsQA and D denoting Deep ProsQA. Differences and interval endpoints use percentage points, while W/L counts discordant questions won or lost by the first model. Canvas single denotes the single-clock canvas, segment variants specify tied or mixed training, and sync denotes synchronous generation. The percentile bootstrap interval and exact discrete test need not cross their respective thresholds at the same contrast. The last three rows are the component steps of Table 4 , whose bootstrap uses 10,000 replicates with seed 0.
Flow matching enables language generation in few steps, but whether additional integration steps improve reasoning remains unclear. We prove that a flow parameterized by a two-layer Transformer can solve graph reachability, with the required number of integration steps increasing with the target's distance from the root. Yet, standard flow language models can fail to benefit from additional steps on reasoning tasks. We attribute this limitation to objectives that supervise each time point independently, without explicitly training successive steps to build on one another. To address this, we instead train through the model's own latent rollout over a randomly sampled subinterval of [0, 1], decoding only at the endpoint. On ProsQA, this raises accuracy to 97% and enables performance to improve with additional integration steps. For the longer rollouts required by reasoning tasks such as Sudoku and Maze, retracting the latent state onto a sphere stabilizes the dynamics and yields substantial gains over baselines with more than three times as many parameters. Sampling multiple rollouts further improves performance when paired with a parameter-free selection score, although reliable selection remains challenging for longer answers. Together, these results establish a theoretical basis for reasoning with flows and show how rollout training, stable latent dynamics, and rollout selection help realize this capacity in practice.
Faissal Izermine, Hanru Bai, Oscar Davis +1
Max Planck Institute for Intelligent Systems · ELLIS Institute Tübingen · ETH Zurich +3
Chain-of-Thought (CoT) empowers Large Language Models (LLMs) to tackle complex problems, but remains constrained by the computational cost and early token commitments in discrete reasoning traces. Recent latent reasoning approaches attempt to optimize efficiency by performing reasoning within continuous hidden states. However, many such methods optimize latent states end to end without a trained interface for intermediate textual readout, and several representative configurations use a pre-defined number of latent steps during inference. In this work, we introduce \textbf{PLaT} (\textbf{P}lanning with \textbf{La}tent \textbf{T}houghts), a framework that decouples latent planning from verbalization. The Planner deterministically evolves latent planning states, while an independent Decoder provides textual readouts when needed. Answer-aware textual stopping allows the latent rollout to use a problem-dependent number of groups rather than a pre-specified chain length. PLaT achieves competitive coverage at larger k in several mathematical settings, with lower Pass@1: on Llama-1B GSM8K, it reaches 80.59% Pass@128 versus CODI's 72.37%. These results support PLaT as a candidate-generation interface supplying multiple textual readouts for downstream verification or reranking.
Flow Language Models (FLMs) have emerged as a continuous-state alternative to discrete diffusion language models, yet the role of their continuous representations in reasoning remains unclear. We investigate this question by comparing the reasoning efficiency of FLMs and discrete diffusion models, measured by solution accuracy under matched denoising steps. Unlike discrete diffusion, which passes categorical states between denoising steps, FLMs evolve a continuous sequence representation throughout denoising and decodes it into discrete tokens only at the end. Our theoretical analysis shows, from a superposition perspective, how information retained in these continuous states can benefit reasoning. Intermediate-state interventions provide further empirical support for this theoretical account, showing that removing information about alternative candidates reduces subsequent solution recovery. Together, these findings show that FLMs allow evidence for multiple candidates to persist and inform subsequent reasoning before a discrete answer is produced. Furthermore, our experiments on maze planning and Sudoku tasks show that FLMs achieve greater reasoning efficiency in the few-step regime: FLMs achieves higher sequence accuracy than discrete diffusion baselines at matched model sizes and small denoising steps. On maze planning tasks, FLMs can also achieve comparable accuracy with smaller models. For example, on Maze15, FLM reaches the 95% accuracy target at 64 denoising steps with 36.5% fewer parameters than MDLM. These findings point to continuous state spaces as a promising foundation for reasoning models that require fewer refinement steps.
Hanru Bai, Faissal Izermine, Oscar Davis +1
Max Planck Institute for Intelligent Systems · ETH Zurich · ELLIS Institute Tübingen +3