Organizations: School of Computer Science and Engineering, Northeastern University, China · Hong Kong Polytechnic University, China · NiuTrans Research, Shenyang, China
Continuous language flows generate text by denoising all positions of a target canvas together. The natural way to add reasoning to such a model is to write a trace ahead of the answer, but the trace length changes from question to question. The answer start is therefore unknown during denoising, and the model has to decide the trace length, the place of every trace token, and the answer at the same time. We propose Plan Canvas to fix the boundary between the trace and the answer. A plan region of fixed capacity holds a compact trace, supervised padding fills its unused positions, and the answer starts at a fixed position. The fixed regions also allow separate denoising clocks for the plan and for the answer. With the trace text, backbone, and canvas length of the free-trace baseline held fixed, Plan Canvas improves accuracy on ProsQA and on Deep ProsQA, a graph benchmark with longer proofs. On Deep ProsQA, accuracy rises from 73.0% to 87.0%, the share of questions answered with a valid path rises from 30.8% to 59.1%, and the gain is largest on the longest proofs.
Figures & tables
Figure 1: Training and sampling use the same ELF backbone with fixed plan and answer regions. Each training example selects either flow matching with region corruption or token cross-entropy with separate decoder corruption. The branch probabilities are 0.8 for flow matching and 0.2 for token cross-entropy in the reported configuration. During sampling, the prompt stays fixed while region clocks control Euler updates of every target position. The model predicts clean states, which define velocity, and switches to token decoding once both regions reach time one.
ProsQA
Deep ProsQA
Layout
Full canvas S
Readout G
Full canvas S
Readout G
Direct answer
720
16
784
16
Full trace
800
96
960
192
Free trace / answer then trace
752
48
848
80
Plan Canvas, either clock interface
752
48
848
80
Table 1: Sequence budgets in token positions, shared by training and evaluation within each target layout. Prompt caps are 704 positions for ProsQA and 768 positions for Deep ProsQA across all layouts. The readout window is computed from these fixed caps and never from a reference answer’s content or length.
ProsQA
Deep ProsQA
Target layout
Accuracy
Valid path
Correct / invalid
Accuracy
Valid path
Correct / invalid
Direct answer
87.6
–
–
57.7
–
–
Full trace
91.0
79.8
12.2
73.5
58.5
20.3
Free trace
91.4
75.0
17.0
73.0
30.8
42.5
Answer then trace
88.2
79.4
8.8
80.2
38.5
41.7
Plan Canvas, single clock
96.8
90.6
6.2
82.1
55.0
27.1
Table 2: Answer accuracy and path checks on the test splits, in percent of all test questions. A valid path starts at a fact about the entity, follows directed edges of the question graph, and ends at the predicted concept. Correct / invalid counts correct answers with an invalid or missing path. Free trace, answer then trace, and both Plan Canvas variants share the trace text, canvas length, and readout window, and shaded rows use the fixed plan region.
fimpus nulpus molpus slilpus rupus chepus, then padding
chepus
Segment clocks
fimpus nulpus molpus slilpus rupus chepus, then padding
chepus
Table 3: Emitted plans for Deep ProsQA test question 450, which asks whether Eva is a chepus or a pripus among 40 stated facts and rules. The underlined pair repeats a concept and then makes a hop that is not an edge of the question graph.
Figure 2: Shape of the emitted paths on the Deep ProsQA test split, in percent of all test questions. Panel (a) counts paths that contain the same concept twice, grouped by shortest proof length. Panel (b) gives the distribution of the emitted path length minus the shortest proof length.
Figure 3: Deep ProsQA test accuracy under three changes. Panel (a) groups the test split by shortest proof length, with 111 or 112 questions per length. Panel (b) evaluates the same checkpoints with 32, 64, 96, and 128 synchronous Euler iterations, counting two guided backbone calls per iteration and one decoder call. Panel (c) trains the segment-clock canvas with different plan capacities, and the dashed line marks the free trace.
Configuration
Fixed
Marker
Prefix
Time
Deep ProsQA
Free trace
–
–
–
–
73.0
Fixed plan region
✓
–
–
–
82.1
+ segment markers
✓
✓
–
–
84.0
+ time-prefix groups
✓
✓
✓
–
85.9
+ position time
✓
✓
✓
✓
87.0
Table 4: Deep ProsQA accuracy in percent as the parts of the segment-clock interface are added under tied training times. Fixed denotes the plan region, and Marker, Prefix, and Time denote segment markers, time-prefix groups, and the position-dependent time term.
Train times
Schedule
(op,oa)
Joint
Calls
ProsQA
Deep
Tied
Synchronous
(0,0)
64
129
95.2
87.0
Tied
Plan-first
(0,64)
128
257
96.0
87.6
Tied
Answer-first
(64,0)
128
257
94.6
85.4
Mixed
Synchronous
(0,0)
64
129
95.2
87.9
Mixed
Plan-lead
(0,32)
96
193
95.4
89.0
Mixed
Plan-first
(0,64)
128
257
95.4
88.7
Table 5: Accuracy in percent under different offsets with the same weights and 64 active steps per region. Joint counts Euler iterations, Calls counts backbone evaluations including the decoder, and mixed training shares region times for half the examples.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Layout
P valid
P test [95% CI]
D valid
D test [95% CI]
Direct answer
88.0
87.6 [84.4, 90.2]
58.2
57.7 [54.6, 60.7]
Full trace
90.3
91.0 [88.2, 93.2]
78.6
73.5 [70.7, 76.1]
Free trace
89.3
91.4 [88.6, 93.5]
71.0
73.0 [70.2, 75.7]
Answer then trace
90.3
88.2 [85.1, 90.7]
78.2
80.2 [77.6, 82.5]
Plan Canvas (single clock)
94.0
96.8 [94.9, 98.0]
82.0
82.1 [79.6, 84.4]
Plan Canvas (segment clocks)
94.3
95.2 [93.0, 96.8]
86.8
87.0 [84.8, 88.9]
Appendix
Table 6: Validation and test answer accuracy in percent, with P denoting ProsQA and D denoting Deep ProsQA. Bracketed values give marginal 95% Wilson intervals for the complete test splits under each fixed checkpoint.
Set
First / second
Δ (pp)
Paired 95% CI
W/L
p
P
Canvas single / Free trace
+5.4
[+3.0, +8.0]
34/7
2.53×10−5
P
Canvas single / Full trace
+5.8
[+3.2, +8.4]
37/8
1.54×10−5
P
Canvas single / Answer then trace
+8.6
[+5.6, +11.6]
52/9
1.80×10−8
P
Segment tied, sync / Canvas single
-1.6
[-3.0, -0.2]
3/11
0.057
P
Segment mixed, sync / Segment tied, sync
+0.0
[-1.4, +1.4]
7/7
1.000
P
Tied, Plan-first / Segment tied, sync
+0.8
[+0.0, +1.8]
5/1
0.219
Appendix
Table 7: Paired test contrasts with first minus second accuracy differences, with P denoting ProsQA and D denoting Deep ProsQA. Differences and interval endpoints use percentage points, while W/L counts discordant questions won or lost by the first model. Canvas single denotes the single-clock canvas, segment variants specify tied or mixed training, and sync denotes synchronous generation. The percentile bootstrap interval and exact discrete test need not cross their respective thresholds at the same contrast. The last three rows are the component steps of Table 4 , whose bootstrap uses 10,000 replicates with seed 0.