Planning requires a transition model that predicts how each action changes the current state. When a large language model (LLM) plays this role, every next state is generated token by token, which makes searching over many possible futures slow and expensive. Existing alternatives either still query an LLM at every step or require a symbolic model of the domain. We propose EmbedPlan, a transition model built on frozen text embeddings: it embeds natural language descriptions of the state and the action with a frozen LLM, predicts the embedding of the next state with a lightweight learned network, and returns the closest real state. Because this network can be trained on top of any encoder, EmbedPlan also provides a controlled way to compare text representations for learning transitions. We evaluate it on 9 classical planning domains, under six settings that hold out progressively more of the data, from transitions to entire domains, and against baselines ranging from predicting no change to learning symbolic action rules. On planning problems seen during training, EmbedPlan almost always ranks the true next state among its top five guesses, still does so for most queries even when every observed state is a candidate, and retains 92-99% of its single-step accuracy when predicting several steps ahead from its own outputs. Given the same candidate states as GPT-5.4, it picks the true next state more often while taking about 0.17 ms per transition with cached embeddings. Accuracy is lower on unseen problems and near chance on unseen domains, and the controlled comparison traces this limit to the state representation rather than to the learned transition.
Figures & tables
Figure 1: EmbedPlan . A frozen LLM encoder E embeds the state s and the action a (Blocksworld, pick-up(C) ), and learned heads πs,πa project them into a latent space where the transition is computed: a small network Tθ predicts the next-state embedding h^s′ , and the nearest real state is retrieved as the successor s^′ . Fed back as the next input, it supports multi-step rollout. Search itself is external. Latency is per transition with state embeddings cached (Appendix B.5 ).
Exposure
Protocol
Trained on → tested on
Distractors from
Observed problems
Interpolation
80%→20% of transitions
the whole domain
Plan-Variant
some → other optimal plans
alternative-plan successors
Unseen problems
Extrapolation
∼80%→20% of problems
the query’s problem
Multi-Domain
the same, all nine domains at once
the query’s problem
Unseen domains
Cross-Domain
one domain → another
the target domain
Leave-One-Out
eight domains → the ninth
the target domain
Table 1: Evaluation protocols, grouped by exposure: what the test data share with training. Every query is ranked against 128 states, the true successor and 127 distractors (definitions in Appendix A ).
Table 3Figure 4
Regime
Ferry
Logistics
Blocksworld
True state at each step
95.3±0.5
74.9±2.2
80.3
Own prediction, snapped
94.2±0.7
69.0±2.1
78.4
Own prediction, not snapped
68.3±2.0
40.0±1.9
59.6
Retention (snapped)
99%
92%
98%
Table 4: Replacing each state prediction with the nearest real state before the next step ( snapped ) keeps multi-step rollout within 92 – 99% of the accuracy obtained with the true state at each step. Step-wise Hit@1 (%), Interpolation, 1,000 candidate states, mean ± SE over seeds (Blocksworld: single run). Retention: snapped over true-state accuracy.
Table 6
Hit@5
Hit@1
Method
Ferry
Logistics
Goldminer
Mean
Mean
Given the symbolic facts of each state
Lifted STRIPS induction (oracle)
100.0±0.0
100.0±0.0
99.3±0.3
99.8
99.8
Bag-of-literals
51.9±5.7
81.4±0.7
95.7±2.0
76.3
37.1
Sparse lexical features of the same text
Char 3–5 grams
71.1±1.7
69.8±7.4
97.3±0.7
79.4
41.9
Table 7: With the head, data, and pool fixed, the state representation decides transfer to unseen problems. Extrapolation, Hit@5 (%) per domain (mean ± SE over 3 seeds) and three-domain means. The learned rows train the identical head on a different state representation (lifted STRIPS induction instead induces symbolic effects), and the first group requires the symbolic facts of each state. On these templated states an action edits only a few facts, which favors sparse and symbolic representations (Section 4.3 ). Bold: best without symbolic facts.
Appendix figures & tables36 assets
Supplementary material from the paper’s appendix.
Appendix
Protocol
Measures
Shared
Shared
Shared
Difficulty
domain
problems
plans
Interpolation
Interpolation
Yes
Yes
Yes
Lowest
Plan-Variant
Plan generalization
Yes
Yes
No
Low
Extrapolation
Extrapolation
Yes
No
No
Medium
Multi-Domain
Capacity sharing
Yes
No
No
Medium
Cross-Domain
Zero-shot transfer
No
No
No
High
Appendix
Table 8: Evaluation protocol hierarchy by generalization difficulty.
Figure 4: Complete EmbedPlan architecture. State and action descriptions are encoded by a frozen LLM encoder E into high-dimensional embeddings zs,za . Learned projection heads πs,πa reduce dimensionality to a shared 128-d space. The transition network Tθ (with a residual connection from hs ) predicts the next-state embedding h^s′ , trained via InfoNCE to maximize similarity to the ground-truth embedding. At inference, the model retrieves the most similar state from a candidate pool.
Hyperparameter
Values explored
Architecture
Model type
MLP , hypernetwork
Hidden size
128 , 256
Number of layers
2 , 4
Dropout
0.0 , 0.5
Layer normalization
Yes
Appendix
Table 9: Hyperparameter search space. Bold values indicate the final configuration.
Domain
Metric
128
512
2,048
8,192
Full
Infl.
Ferry 46,205 states
Hit@1
98.96±0.09
97.21±0.27
90.90±0.55
76.58±0.92
35.23±1.14
2.81×
Hit@5
99.97±0.03
99.85±0.08
99.38±0.09
97.97±0.25
86.93±1.36
1.15×
Hit@10
99.99±0.01
99.94±0.03
99.73±0.08
99.03±0.14
94.96±0.63
1.05×
Logistics 13,373 states
Hit@1
96.29±0.23
88.21±0.65
70.12±1.83
47.20±1.60
32.55±1.69
2.96×
Hit@5
99.85±0.08
99.37±0.12
96.73±0.31
83.13±0.93
74.89±1.71
1.33×
Hit@10
99.90±0.05
99.72±0.09
98.84±0.11
92.99±0.50
88.01±0.92
1.14×
Appendix
Table 10: Effect of candidate pool size under Interpolation, Llama-3.3-70B, mean ± SE over 3 seeds. “Full” retrieves against every observed state in the domain. The inflation column is the 128-way value divided by the full-pool value.
EmbedPlan
Metric
Full pipeline
Batched/cached
LLM API
Latency
18.6 ms
0.168 ms
1,888.2 ms
Speedup vs. LLM
101 ×
11,215 ×
1 ×
Compute per transition
∼ 286M FLOPs †
∼ 286M FLOPs †
198 tokens
Appendix
Table 11: Wall-clock latency and computational cost comparison. The full pipeline includes text-to-vector encoding time with BGE-M3. EmbedPlan eliminates the sequential token bottleneck of autoregressive generation. † Projection heads, transition network, and scoring of 128 candidates with cached encoder outputs (Llama-3.3-70B dimensions). The full-pipeline latency additionally includes encoding the query texts.
Domain
Problems
States
Transitions
Actions
Blocksworld
5
43,551
43,065
4
Depot
7
5,795
13,256
5
Ferry
10
46,205
225,300
3
Floortile
6
33,608
166,565
6
Goldminer
7
12,237
52,023
7
Grid
5
8,671
664,346
5
Appendix
Table 12: Dataset statistics by domain.
Figure 5: Representative transition from Blocksworld.
Figure 6: Representative transition from Ferry.
Figure 7: Representative transition from Logistics.
Figure 8: Representative transition from Rovers.
Qwen2.5-7B
Llama-3.3-70B
Domain
Interp.
Extrap.
Interp.
Extrap.
Blocksworld
100.0
41.6 ± 8.7
100.0
49.1 ± 10
Depot
98.2
24.8 ± 5.2
98.8
25.9 ± 6
Ferry
99.9
36.7 ± 0.6
100.0
40.6 ± 3
Floortile
99.4
55.2 ± 11
99.6
68.8 ± 16
Goldminer
99.9
74.4 ± 5.8
100.0
76.2 ± 5
Appendix
Table 13: Per-domain Hit@5 (%) for the Interpolation and Extrapolation splits.
Next state prediction (%)
Action disambiguation (%)
Domain
Hit@1
Hit@5
Hit@10
Acc@1
Acc@5
Acc@10
Blocksworld
94.6±0.4
100.0±0.0
100.0±0.0
25.9±0.9
88.6±0.2
99.5±0.1
Depot
76.9±1.2
98.8±0.1
99.4±0.1
29.8±2.3
82.2±2.1
93.3±1.4
Ferry
98.5±0.2
100.0±0.0
100.0±0.0
43.1±0.8
96.8±0.1
99.8±0.0
Floortile
97.9±0.4
99.6±0.0
99.8±0.0
37.1±2.8
86.5±1.9
97.3±0.8
Goldminer
93.8±0.8
100.0±0.0
100.0±0.0
17.0±0.5
73.3±0.6
97.4±0.2
Appendix
Table 14: Single-step transition prediction (Interpolation split, Llama-3.3-70B). Next state prediction : whether the ground-truth next state is in the top- k retrieved candidates. Action disambiguation : whether, among all actions applied to s , the correct action a produces the prediction closest to ground-truth s′ . Mean Acc@5 is 85.3%.
State prediction
Action accuracy
Domain
Hit@1
Hit@5
Hit@10
Acc@1
Acc@5
Acc@10
Blocksworld
17.6 ± 5.2
49.1 ± 9.8
64.6 ± 7.5
0.7 ± 0.1
7.9 ± 1.2
24.0 ± 2.3
Depot
4.7 ± 1.2
25.9 ± 6.4
41.2 ± 8.1
0.7 ± 0.2
7.8 ± 1.7
21.8 ± 2.3
Ferry
12.0 ± 1.7
40.6 ± 3.5
58.1 ± 4.0
1.1 ± 0.4
10.3 ± 1.2
23.7 ± 2.3
Floortile
37.6 ± 13
68.8 ± 16
78.9 ± 13
3.2 ± 0.6
28.3 ± 8.7
52.3 ± 14
Goldminer
35.8 ± 6.4
76.2 ± 5.2
88.0 ± 1.7
3.0 ± 0.2
16.5 ± 1.2
45.7 ± 3.5
Appendix
Table 15: Full metrics (Llama-3.3-70B, Extrapolation split). Action accuracy measures whether the correct action is identified given (s,s′) .
Block
Depot
Ferry
Floor
Gold
Grid
Logis
Rover
Satel
Mean
Blocksworld
5.7
5.5
6.7
6.1
6.5
9.2
5.1
5.3
6.3
Depot
4.3
5.3
4.8
5.0
4.9
6.0
7.5
5.4
5.4
Ferry
5.3
8.7
9.3
5.2
9.0
22.3
6.6
5.6
9.0
Floortile
5.0
7.3
7.0
6.1
6.5
7.5
10.0
7.9
7.2
Goldminer
4.2
5.4
5.7
5.0
7.5
5.4
5.9
4.9
5.5
Grid
5.4
6.6
8.6
9.0
10.4
14.3
5.9
5.4
8.2
Appendix
Table 16: Cross-Domain transfer (Llama-3.3-70B). Hit@5 (%) training on the row domain and testing on the column domain. Chance: 3.9% (5 of 128). Untrained model: 4.2%.
Held-out domain
Hit@1
Hit@5
Hit@10
Logistics
3.3 ± 0.4
15.8 ± 2.2
27.6 ± 3.3
Grid
2.3 ± 0.2
12.3 ± 0.8
22.9 ± 1.4
Rovers
2.6 ± 0.3
12.6 ± 1.1
22.6 ± 1.6
Ferry
1.8 ± 0.1
9.0 ± 0.5
17.2 ± 0.8
Satellite
1.6 ± 0.1
8.4 ± 0.8
16.1 ± 1.5
Floortile
1.7 ± 0.3
7.9 ± 1.1
14.9 ± 2.0
Appendix
Table 17: Leave-One-Out results (Llama-3.3-70B). Train on eight domains, test on the held-out domain.
Domain
Hit@5
Domain
Hit@5
Floortile
52.0 ± 10
Blocksworld
36.9 ± 13
Rovers
51.9 ± 5
Satellite
34.3 ± 7
Goldminer
46.0 ± 8
Grid
32.6 ± 1
Ferry
37.7 ± 14
Logistics
24.8 ± 10
Depot
18.9 ± 12
Mean : 37.2 ± 3.8 (vs. 54.6 single-domain)
Appendix
Table 18: Multi-Domain unified model (Llama-3.3-70B, Extrapolation split).
Table 20: Plan-level evaluation (Llama-3.3-70B). Trajectory metrics under the Plan-Variant and Extrapolation splits.
BGE-M3
MPNet
Plan-Variant
Mean Hit@5
Exact Hit@5
Mean Hit@5
Exact Hit@5
Blocksworld
23.7±1.5
0.2±0.1
16.7±0.3
0.3±0.1
Depot
45.5±1.0
20.9±0.1
31.1±1.6
10.8±1.1
Ferry
63.3±1.4
27.8±2.1
60.9±2.0
29.1±2.3
Floortile
49.2±1.0
13.1±1.1
9.2±0.5
3.0±0.4
Goldminer
67.9±2.6
35.0±3.2
26.8±0.8
10.5±0.5
Appendix
Table 21: Plan-level evaluation across encoders. Top : Plan-Variant trajectory metrics for BGE-M3 and MPNet (Llama-3.3-70B in Table 20 ). Bottom : Extrapolation, mean ± SE over the six domains with an independent run for every encoder (Blocksworld, Depot, Ferry, Floortile, Goldminer, Grid).
Depth
n
True state at each step
Snapped
Not snapped
1
300
94.3±0.7
94.3±0.7
94.3±0.7
2
300
96.3±0.5
94.4±0.7
52.4±3.4
3
57
94.2±1.9
92.4±0.5
20.9±2.4
4
7
96.7±4.7
93.3±9.4
10.3±7.4
Appendix
Table 22: Closed-loop rollout by depth on Ferry. Step-wise Hit@1 in % (mean ± SE), where n is the number of test trajectories reaching that depth. Same models, trajectories, and pool as Table 4 , and the depth rows weighted by n reproduce the aggregate to rounding.
Statistic
Value
Interpolation (mean ± SE)
99.5% ± 0.2%
Extrapolation (mean ± SE)
47.7% ± 4.9%
Gap
51.8 pp
Paired t -test
t(8)=10.58
p -value
5.57×10−6
95% CI
[40.7,62.9] pp
Appendix
Table 23: Statistical analysis of the generalization gap (Qwen2.5-7B, paired across the nine domains).
Comparison
Encoder
Δ
Cohen’s d
p
Interpolation vs. Extrapolation
Qwen2.5-7B
− 51.8 pp
5.25
<10−5
Extrapolation vs. Cross-Domain
Llama-3.3-70B
− 48.0 pp
4.32
<10−5
Llama vs. MPNet (Extrap.)
both
+ 27.8 pp
2.53
<0.001
Single vs. Multi (Extrap.)
Llama-3.3-70B
+ 17.4 pp
1.42
0.028
Appendix
Table 24: Effect sizes for major findings. The encoder column is given explicitly because the comparisons are drawn from different encoder runs.
Domain
Gap (pp)
t -statistic
p -value
Depot
73.4
12.87
1.17×10−4
Ferry
63.3
65.49
3.24×10−6
Satellite
59.9
153.39
4.11×10−8
Blocksworld
58.4
5.95
2.70×10−3
Logistics
55.0
4.28
8.28×10−3
Rovers
50.5
20.22
8.90×10−6
Appendix
Table 25: Independent t -tests comparing interpolation and extrapolation per domain (Qwen2.5-7B).
Figure 9: Embedding space fragmentation across scales. PCA of frozen state embeddings, colored by domain. (A) MPNet places each domain in its own tight, isolated cluster. (B) Llama-3.3-70B, despite being roughly 640 × larger, shows the same fragmentation. Pre-trained embeddings thus group states by domain-specific surface features rather than by abstract planning roles, regardless of model scale.
Domain
Actions
Predicates
Gap
Depot
5
8
73.4
Logistics
6
6
55.0
Satellite
4
8
59.9
Blocksworld
4
5
58.4
Ferry
3
5
63.3
Rovers
9
26
50.5
Appendix
Table 26: Domain complexity metrics and generalization gaps (Qwen2.5-7B). Action counts are action schemas as defined in the PDDL domain file, matching Table 12 .
Encoder
Parameters
MLP
Hypernetwork
Llama-3.3-70B
70B
54.6 ± 5.5
53.8 ± 5.5
Qwen2.5-7B
7B
47.7 ± 4.9
46.2 ± 5.0
Appendix
Table 27: Architecture comparison for the two largest encoders (Extrapolation split, Hit@5 %).
Hit@5 (%)
Action Acc@5 (%)
Domain
λ=0
λ=2
λ=0
λ=2
Blocksworld
30.2 ± 8
49.1 ± 10
1.8 ± 0.4
7.9 ± 2
Depot
13.4 ± 4
25.9 ± 6
1.6 ± 0.3
7.8 ± 3
Ferry
24.8 ± 3
40.6 ± 3
2.5 ± 0.5
10.3 ± 2
Floortile
45.3 ± 14
68.8 ± 16
8.2 ± 3
28.3 ± 15
Goldminer
55.7 ± 6
76.2 ± 5
4.9 ± 1.2
16.5 ± 2
Appendix
Table 28: Ablation of the action disambiguation loss. Comparison of models trained without ( λ=0 ) and with ( λ=2 ) action disambiguation under Extrapolation evaluation (Llama-3.3-70B). The action loss yields gains in both state prediction ( + 19.3 pp) and action accuracy ( 3.4× ).
Domain
Ground truth s′
Retrieved s^′
Blocksworld
Block A is clear, arm holds B , C is on table
Block A is clear, arm holds C , B is on table
Ferry
Car c1 at l0, ferry empty, c2 at l1
Car c1 at l0, ferry empty, c2 at l0
Logistics
Package p1 in truck t0 , t0 at l1-0
Package p1 at l1-0 , t0 at l1-0
Appendix
Table 29: Representative Hit@5 errors. Retrieved states differ from the ground truth by one or two predicates (in italics), typically involving object locations or holdings within the same problem instance.
Domain
Split
Hit@1
Hit@5
Hit@10
Blocksworld
Interpolation
94.3 ± 0.2
100.0 ± 0.0
100.0 ± 0.0
Extrapolation
14.2 ± 4.0
41.6 ± 8.8
56.6 ± 8.3
Depot
Interpolation
76.7 ± 1.0
98.2 ± 0.1
99.0 ± 0.0
Extrapolation
4.9 ± 1.2
24.8 ± 4.9
42.2 ± 7.8
Ferry
Interpolation
98.7 ± 0.1
99.9 ± 0.0
100.0 ± 0.0
Extrapolation
10.7 ± 0.8
36.7 ± 0.7
52.8 ± 1.1
Appendix
Table 30: Complete metrics for Qwen2.5-7B across all domains and splits.
Domain
Split
Hit@1
Hit@5
Hit@10
Blocksworld
Interpolation
96.2 ± 0.2
100.0 ± 0.0
100.0 ± 0.0
Extrapolation
17.6 ± 5.2
49.1 ± 9.8
64.6 ± 7.5
Depot
Interpolation
79.5 ± 0.7
98.8 ± 0.1
99.2 ± 0.1
Extrapolation
4.7 ± 1.2
25.9 ± 6.4
41.2 ± 8.1
Ferry
Interpolation
99.1 ± 0.1
100.0 ± 0.0
100.0 ± 0.0
Extrapolation
12.0 ± 1.7
40.6 ± 3.5
58.1 ± 4.0
Appendix
Table 31: Complete metrics for Llama-3.3-70B across all domains and splits.
Method
Ferry
Logistics
Goldminer
Mean
Lifted STRIPS induction (oracle)
100.0±0.0
100.0±0.0
99.3±0.3
99.8
Char 3–5 grams
71.1±1.7
69.8±7.4
97.3±0.7
79.4
Bag-of-words
70.8±2.7
67.3±4.8
99.3±0.2
79.1
Bag-of-literals
51.9±5.7
81.4±0.7
95.7±2.0
76.3
TF-IDF bigrams
59.5±2.1
50.0±4.3
97.2±1.2
68.9
EmbedPlan (Llama-3.3-70B)
40.6±3.5
53.7±9.8
76.2±5.2
56.8
Appendix
Table 32: Extrapolation, Hit@5 (%), mean ± SE over 3 seeds. The Llama-3.3-70B row is reproduced from Table 13 , and all other rows were computed for this comparison under the identical protocol.
Method
Ferry
Logistics
Goldminer
Mean
Lifted STRIPS induction (oracle)
100.0±0.0
100.0±0.0
99.3±0.3
99.8
Bag-of-words
31.5±3.3
25.4±4.2
78.0±1.9
45.0
Char 3–5 grams
31.0±2.4
24.2±2.4
70.6±3.5
41.9
TF-IDF bigrams
24.7±1.5
14.2±2.9
72.8±2.5
37.2
Bag-of-literals
14.2±0.9
30.7±0.8
66.5±8.7
37.1
EmbedPlan (Llama-3.3-70B)
12.0±1.7
16.5±3.5
35.8±6.4
21.4
Appendix
Table 33: Extrapolation, Hit@1 (%), mean ± SE over 3 seeds. The separation between EmbedPlan and the non-learned floors is wider here than at Hit@5.
Method
Ferry
Logistics
Goldminer
Mean
Hit@5
Bag-of-literals
100.0±0.0
100.0±0.0
100.0±0.1
100.0
Bag-of-words
99.9±0.1
100.0±0.0
100.0±0.0
100.0
EmbedPlan (Llama-3.3-70B)
100.0±0.0
99.9±0.0
100.0±0.0
100.0
TF-IDF bigrams
99.9±0.0
100.0±0.0
100.0±0.0
100.0
Char 3–5 grams
99.9±0.0
99.9±0.1
100.0±0.0
100.0
Appendix
Table 34: Interpolation, Hit@5 and Hit@1 (%), mean ± SE over 3 seeds. Six methods lie within 0.2 points of one another on Hit@5.
Encoder treatment
Extrapolation
Interpolation
Frozen
3.8
25.6
LoRA, cold start
6.2
39.6
LoRA, warm start
9.3
41.7
Appendix
Table 35: Full-pool Hit@1 (%), mean over three domains and three seeds. The full pool contains every state in the domain ( 12,237 – 46,205 ) rather than 128 candidates.
Domain
Extrapolation
Interpolation
Goldminer
+13.3
+10.5
Ferry
−0.1
+1.6
Logistics
−3.9
−5.9
Appendix
Table 36: Effect of the warm start, in points of full-pool Hit@1 relative to the cold start. The sign is consistent per domain across two independent splits.
Domain
cos(s,s′)
cos(s,σ(s))
dorder/daction
Ferry
0.99672
0.99654
1.06
Goldminer
0.99881
0.99862
1.16
Logistics
0.99928
0.99915
1.18
Mean
1.13
Appendix
Table 37: Embedding displacement from a semantically null reordering ( dorder=1−cos(s,σ(s)) , with σ a random clause order) against that from a real transition ( daction=1−cos(s,s′) ), BGE-M3.
Chain-of-Thought (CoT) empowers Large Language Models (LLMs) to tackle complex problems, but remains constrained by the computational cost and early token commitments in discrete reasoning traces. Recent latent reasoning approaches attempt to optimize efficiency by performing reasoning within continuous hidden states. However, many such methods optimize latent states end to end without a trained interface for intermediate textual readout, and several representative configurations use a pre-defined number of latent steps during inference. In this work, we introduce \textbf{PLaT} (\textbf{P}lanning with \textbf{La}tent \textbf{T}houghts), a framework that decouples latent planning from verbalization. The Planner deterministically evolves latent planning states, while an independent Decoder provides textual readouts when needed. Answer-aware textual stopping allows the latent rollout to use a problem-dependent number of groups rather than a pre-specified chain length. PLaT achieves competitive coverage at larger k in several mathematical settings, with lower Pass@1: on Llama-1B GSM8K, it reaches 80.59% Pass@128 versus CODI's 72.37%. These results support PLaT as a candidate-generation interface supplying multiple textual readouts for downstream verification or reranking.
Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on large generative models unsuited for the high-sampling nature of model-based planning. To address these challenges, we introduce Latent Goal Prediction from Language (LAGO), a framework that predicts both sequences of intermediate goal states from language instructions and action-conditioned rollouts, all within the same latent space. Rather than optimizing toward a single global objective, LAGO dynamically decomposes instructions into explicitly predicted, locally tractable latent subgoals. By updating these subgoals online and using a soft minimum trajectory cost during planning, LAGO enables an agent to follow coherent latent trajectories over long horizons. Evaluation across multiple environments planning horizons shows that LAGO avoids the sharp degradation of prior methods. By achieving robust and precise long-horizon planning purely from language, LAGO bridges the precision of visual goals with the flexibility of text-guided control.
Samuel Barbeau, Simon Roy, Giovanni Beltrame +2
1École de Technologie Supérieure, Montréal · 2International Laboratory on Learning Systems (ILLS) · 3Polytechnique Montréal +4
We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during the forward pass, and whether they causally drive generation. Using rhyming-couplet completion as a clean test of forward-looking constraint, we apply two lightweight methods (linear probing and activation patching) across Qwen3, Gemma-3, and Llama-3 at more than ten scales. Probing shows that future-rhyme information is linearly decodable at the line boundary, with signal that strengthens with scale in all three families. Activation patching reveals that only Gemma-3-27B causally relies on this encoding, exhibiting a handoff in which the causal driver migrates from the rhyme word to the line boundary around layer 30. Every other model we test conditions on the rhyme word throughout generation, with near-zero causal effect at the line boundary despite strong probe signal. We localize the Gemma-3-27B handoff to five attention heads through two-stage path patching that recover ~90% of the rhyme-routing capacity at the newline.