AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundamental insight, we frame "understanding" as learning constraints and propagating their consequences. We construct learning tasks on monoid worlds, sets of states connected by action transitions, where observed training transitions and an unseen constraint jointly determine held-out transitions. Measuring generalization tests whether models can learn global constraints from local transitions and propagate their consequences. We consider inverse, commutativity, composition, and periodicity constraints relevant to spatial and semantic structure. Across attention, recurrent, and state-space architectures, next-state training fits the data but fails to propagate non-trivial constraints. Compositional training, which uses identical paths but hides intermediate states from the input, achieves 96% accuracy on inverse, commutativity, and composition constraints across architectures, yields corresponding improvements in geometric generalization of world models trained on embodied environments and relational generalization in Wikidata-finetuned LLMs. How far do models propagate constraints when inferring an unseen fact may depend on first inferring others? We define proof depth d of a held-out transition, measuring the minimum number of inference rounds to infer the transition, and find that model generalization decreases sharply with proof depth. Increasing compositional path length T improves generalization. These results provide a formal way to investigate global understanding in language and world models and demonstrate that compositional training promotes information propagation and integration.
Figures & tables
Figure 1: Toy instantiation of an inverse monoid world with training and evaluation setup. a) A four-state world with actions L and R satisfying inverse constraint LR=RL=1 , seen and unseen transitions, and sample training path. b) Constraint propagation uses seen transitions and the constraint to infer unseen transitions. c) Next-state training shows intermediate states, while compositional training does not. Evaluation is generalization over unseen transitions. d) 1) Next-state training fails to propagate inverse constraint, while compositional training succeeds ( ≥93% ). 2) Training curves for T=2 show sharp grokking-like generalization. Plots averaged over 5 seeds.
Figure 2: Across monoid worlds and architectures, compositional training generalizes better than next-state training with equal optimizer steps. Compositional objective propagates inverse, commutative, and composition constraints with average accuracy of 96% , and copy at 100% . Next-state training only succeeds on copy and fails on other tested families. Neither propagates periodicity. Points averaged over 5 seeds and bars show standard deviations.
Figure 3: Compositional objectives improve constraint propagation over next-state-style training in vision and language a) 1) AI2-THOR virtual 3D environment, where an embodied agent translates between observed states. Inverse LR=RL=1 and commutative FR=RF constraints hold approximately due to collisions. 2) Training and evaluation follow Figure 1 c with image observations replacing abstract states. 3) Next-observation training marginally outperforms pixel-shifting baseline ( +8% ) while compositional training improves by ( +46% ) (observation). Training to predict state, minimizing visual cues, collapses next-state generalization to chance, while compositional training generalizes with 96% (state). Results averaged over three seeds. b) 1) Wikidata knowledge graph world, where entities are states and relations are multi-valued actions. 2) LLMs are finetuned with next-token prediction, the objective changing the sentence phrasing. 3) Compositional, not next-state, phrasing propagates inverse constraint at average accuracy of 45% and composition at 73% .
Figure 4: Section of a commutative monoid world with unseen transitions of proof depths d=1 and d=2 . Proof depth measures propagation complexity and predicts generalization and training dynamics. a) Section of a larger commutative world containing two unseen transitions. b) A constraint cascade iteratively applies a simple derivation rule (here, commutative equations) to infer unseen transitions. Proof depth is the minimal number of iterations. The d=1 transition is a premise for deriving d=2 c) Generalization decreases geometrically with proof depth, but increasing composition length improves propagation to higher depths. d) Compositional training propagates constraints sequentially in proof-depth order, suggesting an internal constraint cascade.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Width
Parameters (M)
States ( N )
Steps
Seen (%)
Unseen (%)
Model size
64
0.33
1024
150k
100.0
0.4 ± 0.5
128
1.05
1024
150k
100.0
0.4 ± 0.5
256
3.67
1024
150k
100.0
0.2 ± 0.4
320
5.57
1024
150k
100.0
0.4 ± 0.5
World size
320
5.57
1024
150k
100.0
0.4 ± 0.5
320
7.54
4096
150k
100.0
0.1 ± 0.1
Appendix
Table 1: The next-state objective does not propagate constraints even after scaling model, world, and compute. Transformers trained on inverse world at T=2 . Every row is next-state training except the last, which trains the same world, training paths, and model but compositionally. Accuracies are average ± standard deviation over five seeds.
Figure 5: Models systematically fail by recovering unseen transitions with another action’s target. When models fail to recover an unseen transition (s,A,A(s)) , they instead often predict B(s) for another action B . Points averaged over 5 seeds and bars show standard deviations.
Figure 6: Compositionally trained transformers recover unseen transitions in proportion to constraint evidence. Transformers trained on the composition world N=1024 with T=2 paths. Points averaged over 5 seeds and bars show standard deviations.
Figure 7: Composition length has minimal effect on propagating immediately derivable transitions. Compositional training at T=4 propagates inverse, commutative, and composition constraints with an average accuracy of 96% , the same as at T=2 , and copy at 100% . Increasing composition length does not rescue periodicity. Points averaged over 5 seeds and bars show standard deviations.
Figure 8: Converse constraint propagation decreases sharply with branching. Compositional training propagates the inverse constraint ( β=0 ) at 99% , but only a 10% chance of a second child ( β=0.1 ) drops generalization to 63% , and for β≥0.75 accuracy drops to ≤3% . Next-state training fails at every β . Points averaged over 3 seeds and bars show standard deviations.
Objective
Shortcut paths
Seen (%)
Unseen (%)
Next-state
with
100
0 ± 0
without
100
0 ± 0
Compositional
with
100
92 ± 6
without
100
50 ± 10
Appendix
Table 2: Compositional training propagates composition without shortcut paths. Transformers trained on the composition world N=1024 at T=2 with and without the shortcut path sAA(s)BC(s) of each held-out transition. Accuracies are average ± standard deviation over five seeds.
Unseen (%) at path length T
Constraint
1
2
3
4
5
6
7
8
A3=1
0.4 ± 0.5
0.4 ± 0.5
0.2 ± 0.4
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
A4=1
0.6 ± 0.5
0.2 ± 0.4
0.2 ± 0.4
0.2 ± 0.4
0.0 ± 0.0
0.2 ± 0.4
0.0 ± 0.0
0.2 ± 0.4
A5=1
0.6 ± 0.5
0.2 ± 0.4
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
0.0 ± 0.0
Appendix
Table 3: Periodicity is not propagated across composition lengths. All transformers trained compositionally on N=1020 periodicity worlds with Ak=1 for k=3,4,5 and composition lengths 1≤T≤8 fail to propagate the periodicity constraint. Accuracies are averaged over five seeds, with standard deviations of at most 0.5% .
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection interventions, the sequential update induces more robustness than separate policy RL when removing the world model's leading input directions, suggesting that it has learned alternative input pathways. Behaviorally, we find the sequentially trained agent explores a wider range of states and actions. Based on this, we ask: does policy training preserve world knowledge as well as it could? We probe this with training-free merging built on the geometrically motivated input basis plus an online world-model loss during policy RL, and show both improve over the untreated baseline. Our findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Whether language models can systematically generalize remains actively debated. Yet empirical performance is jointly shaped by multiple factors such as training data, training paradigms, and inference-time strategies, making failures difficult to interpret. We introduce a controlled synthetic environment based on shortest-path planning, a canonical composable sequential optimization problem. The setup enables clean separation of these factors and supports two orthogonal axes of generalization: spatial transfer to unseen maps and length scaling to longer-horizon problems. We find that models exhibit strong spatial transfer but consistently fail under length scaling due to recursive instability. We further analyze how distinct stages of the learning pipeline influence systematic problem-solving: for example, data coverage sets capability limits; reinforcement learning improves training stability but does not expand those limits; and inference-time scaling enhances performance but cannot rescue length-scaling failures.
Yao Tong, Jiayuan Ye, Anastasia Borovykh +1
National University of Singapore · Capital Fund Management · Google Research
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling based on language models can further push the boundaries of general agents. (i) We first focus on building foundation models for agentic environment simulation. We introduce Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, the first language world models capable of simulating agentic environments covering 7 domains via long chain-of-thought reasoning. Leveraging more than 10M environment interaction trajectories of 7 domains in real-world environments, we develop Qwen-AgentWorld through a three-stage training pipeline: CPT injects general-purpose world modeling capabilities from the state transition dynamics and augmented professional corpora, SFT activates next-state-prediction reasoning, and RL sharpens simulation fidelity through a tailored framework with hybrid rubric-and-rule rewards. To evaluate language world models, we present AgentWorldBench, a comprehensive benchmark constructed from real-world interactions of 5 frontier models on 9 established benchmarks. Empirical results demonstrate that Qwen-AgentWorld significantly outperforms existing frontier models. (ii) Beyond foundation models, we further investigate two complementary paradigms through which world modeling enhances general agents. First, as a decoupled environment simulator, Qwen-AgentWorld supports scalable and controllable simulation of thousands of real-world environments for agentic RL, yielding gains that surpass real-environment training alone. Second, as a unified agent foundation model, world-model training acts as a highly effective warm-up that improves downstream performance across 7 agentic benchmarks. Code: https://github.com/QwenLM/Qwen-AgentWorld