EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
Authors: Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, +3 more
Organizations: Shanghai Jiaotong University · Eastern Institute of Technology, Ningbo · Zhongguancun Academy, Beijing, China · Hong Kong Polytechnic University
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.
Figures & tables
Figure 1: From prediction to preference. (a) World-model approaches predict action consequences before choosing an action. (b) Under single-goal supervision, a policy can fit habits that fail to transfer. (c) Evoke instead supervises action preferences at a fixed state and history under different goals.
Figure 2: The Evoke training loop. At collected states, actions are executed and assessed under alternative goals with the state, history, and available actions fixed. The policy learns by contrastive ranking on aggregated preferences; actions can switch between positive and competing across goals.
Table 3
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg
Qwen2.5-3B-Instruct
Prompting
ReAct
17.4
6.7
8.8
7.4
9.1
0.0
8.2
Policy optimization and distillation
GRPO
73.9
60.0
82.4
59.3
72.7
76.9
70.9
SDAR
71.0
77.5
71.8
73.9
73.2
57.6
70.8
Table 3: ALFWorld Unseen results (%). The best and second-best results are highlighted within each backbone.
Method
Alt. goals
Objective
Seen ↑
Unseen ↑
Avg Steps ↓
πboot
—
Walkthrough SFT
69.3
60.4
25.9
Source only
No
Ranking
82.9
82.8
16.0
Source replay
No
Ranking
90.4
85.9
15.3
Positive SFT
Yes
SFT
91.4
84.3
15.0
Evoke
Yes
Ranking
92.1
91.8
11.4
Table 4: Decision-supervision ablations. Micro success (%) and average unseen steps, averaged over training seeds. All variants start from πboot . The best and second-best results are highlighted.
Figure 6
Figure 5: From knowledge to goal-directed decisions on unseen games. Consequences are decodable before any ALFWorld training, and Evoke makes the fewest habitual errors. (a) Balanced accuracy of linear probes that predict an action’s consequence from hidden states before the action is executed, on actions whose text leads to different outcomes; Qwen2.5-3B is the backbone before any ALFWorld training. (b) Outcome of each decision on 812 goal pairs that require different actions. A habitual error chooses the action that is correct for the other goal of the pair.
Figure 6: Execution on unseen games. Evoke succeeds earlier and wastes fewer actions. (a) Cumulative success over interaction steps. (b) Invalid actions per game. (c) Revisits to an already visited location per successful game.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Round
States
Contexts
Groups
Games
Epochs
Learning rate
Updates
1
344
687
706
193
2
10−5
178
2 (aggregate)
767
1,482
1,524
312
1
5×10−6
191
Appendix
Table 5: Training schedule of the analyses in Section 5 . Round 2 trains on the aggregate of both rounds.
Supervised decisions
% of πboot
Games
Seen
Unseen
πboot
16,247
100.0
2,700
69.3
60.4
+ Round 1
687
4.2
193
88.6
79.1
+ Round 2 (aggregate)
1,482
9.1
312
92.1
91.8
Appendix
Table 6: Supervision and success. πboot imitates one action per demonstration step; Evoke ranks the candidate actions of each context. Micro success (%).
Data
20%
40%
60%
80%
100%
States
154
307
460
613
767
Contexts
301
592
889
1,185
1,482
Appendix
Table 7: Data subsets. States and goal contexts after aggregation.
Seen
Unseen
Data
Evoke
SFT
Evoke
SFT
Δ
20%
77.4 ± 1.6
75.7 ± 0.7
72.6 ± 0.4
67.9 ± 0.0
+4.7
40%
82.4 ± 1.5
78.6 ± 0.0
82.3 ± 0.4
68.2 ± 1.6
+14.2
60%
89.3 ± 0.7
84.3 ± 0.7
87.6 ± 1.1
78.4 ± 1.3
+9.2
80%
92.1 ± 1.9
85.7 ± 0.0
90.5 ± 1.1
82.1 ± 2.0
+8.5
100%
93.6 ± 1.4
89.5 ± 0.8
92.5 ± 2.0
84.8 ± 1.1
+7.7
Appendix
Table 8: Data efficiency. Micro success (%), mean ± s.d. over three training orders; all budgets, including 100%, are retrained on nested subsets, independently of Table 4 . Δ : Unseen Evoke minus SFT.
Backbone
πboot
Round 1
Round 2
1.7B
80.0 / 65.7
87.1 / 74.6
92.1 / 85.8
3B
69.3 / 60.4
88.6 / 79.1
92.1 / 91.8
7B
81.4 / 72.4
87.1 / 83.6
95.7 / 94.0
Appendix
Table 9: Iterative training. Seen / Unseen micro success (%) after each round.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Micro
Macro
Seen
πboot
74.3
61.5
77.8
62.5
60.0
70.8
69.3
67.8
Source only
91.4
84.6
81.5
81.3
68.0
87.5
82.9
82.4
Source replay
95.4
95.9
90.9
92.5
73.9
95.3
90.4
90.6
Positive SFT
97.1
84.6
81.5
87.5
96.0
95.8
91.4
90.4
Evoke
91.4
76.9
96.3
100.0
92.0
91.7
92.1
91.4
Appendix
Table 10: Taskwise results of the decision-supervision ablations. Success (%) on Seen and Unseen games; averaged over training seeds.
Goals
Seen
Unseen
Original goals only
84.6
78.8
Half alternative goals
86.1
84.5
Appendix
Table 11: Alternative goals under a fixed budget. Micro success (%).
Ranked against
Seen
Unseen
Steps
Policy-favored actions ( Evoke )
92.1
91.8
11.4
Other goals’ positives
85.7
79.9
16.7
Random actions
78.6
74.6
18.2
Least likely actions
72.1
67.2
21.8
Appendix
Table 12: Competitors in the ranking. Micro success (%) and average unseen steps.
Goal
EVOKE
Positive SFT
Clean a bowl; place in cabinet
Takes the bowl (3), cleans it at the sink (5), and places it in the cabinet (7).
Repeatedly visits countertops; reaches the 50-action limit.
Put a hot cup in cabinet
Takes the cup out of the cabinet (6), heats it in the microwave (8), and returns it to the cabinet (10).
Revisits a countertop and repeats ineffective moves; reaches the 50-action limit.
Put two keychains in safe
Places the first keychain (5), finds and takes a second keychain (14), and places it in the same safe (16).
Shares actions 1–5, then revisits locations and picks up a watch (21); reaches the 50-action limit.
Appendix
Table 13: Examples of goal-directed execution from paired unseen trajectories; parentheses give action indices.
Hot
Held
Cool
Clean
Contained
Input
All
Same
All
Same
All
Same
All
Same
All
Same
Action text
64.9
52.5
68.5
52.6
54.4
47.3
56.4
46.6
70.1
49.5
Pre-action state
86.9
85.6
72.7
77.0
83.4
82.6
85.1
85.1
60.4
66.4
Qwen2.5-3B
99.5
99.6
99.1
97.7
98.8
98.6
99.5
99.6
99.3
98.4
πboot
99.4
98.8
99.5
98.8
98.4
98.1
99.8
100.0
99.3
98.4
Positive SFT
99.7
99.8
99.1
98.3
98.4
98.3
99.8
100.0
99.1
98.1
Appendix
Table 14: Consequence probes on unseen games. Balanced accuracy (%) on all test transitions and on the same-action subset.
Held
Contained
Input
All
Same
All
Same
Action text
68.5
52.6
70.1
49.5
History text
80.9
77.1
75.7
71.3
Random weights
66.1
65.0
69.2
68.0
Qwen2.5-3B
99.1
97.7
99.3
98.4
Qwen2.5-3B, control task
60.5
56.9
66.7
60.2
Appendix
Table 15: Probe controls on unseen games. Balanced accuracy (%).
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection interventions, the sequential update induces more robustness than separate policy RL when removing the world model's leading input directions, suggesting that it has learned alternative input pathways. Behaviorally, we find the sequentially trained agent explores a wider range of states and actions. Based on this, we ask: does policy training preserve world knowledge as well as it could? We probe this with training-free merging built on the geometrically motivated input basis plus an online world-model loss during policy RL, and show both improve over the untreated baseline. Our findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with long-horizon tasks, especially under partial observability. World models are a promising way to enhance policy performance, both during training and inference. During inference, agents currently use world models to simulate the consequences of candidate actions before choosing an action, which can improve decision-making. However, we argue that simulation alone is an incomplete interface for decision-making under partial observability: simulation does not adequately capture uncertainty about the current state, which agents may need for accurate decision-making. We address this limitation with Belief-Based World Models (BB-WMs), which maintain a belief that LLMs can query to access information on what is known and uncertain about the current state. Before developing methods to learn accurate BB-WMs, this paper focuses on a more fundamental question: does exposing a world model's belief directly to an LLM policy improve decision-making? Our results show that giving LLM agents access to beliefs improves task performance under partial observability, while remaining complementary to existing simulation-based world models. Code: https://github.com/skumar-ml/belief-world-models.
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agent's current decision. To bridge this gap, we propose Agent-Authored World Modeling (AAWM), a training procedure that constructs supervision from the policy's own decision needs. Specifically, at each state, the agent identifies what it needs to understand about the environment before acting. These needs drive the retrieval of relevant transition evidence across trajectories, which is then synthesized into training targets that capture decision-oriented dynamics instead of reconstructing the next observation. This aligns the training objective with the dynamics the policy needs before acting, not with the contents of the next observation. Experimental results validate the effectiveness of AAWM across multiple environments and training settings. These results show that decision-aware world-model targets provide a more effective learning signal than next-observation prediction.