Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training
Authors: Xinyu Che, Hang Yan, Yanchen Liu, Haochen Liu, Ruifeng Li, Anran Shi, Heng Wang, Jun Liu
Organizations: Xi’an Jiaotong University · University of Southern California · University of the Chinese Academy of Sciences · East China Normal University
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observation training methods help the agent to learn the environmental content. However, they additionally involve an optimization process, which may introduce several effects other than learning to predict the world. Consequently, where the performance gain comes from during the training process remains an open question. We answer this research question through replacing true next-observation targets with in-distribution mismatched observations during the training process. Across two interactive text environments, mismatched targets lower prediction accuracy by 15.3-61.6% relative to ground-truth targets, yet retain substantial task gains over the base model. Compared with the base model, trained models consider more candidate actions and exhibit less looping. We also introduce a setting that replaces prediction-based rewards with independent random signals. This training expands task coverage (pass@64) even when the reward carries no environment information. We also generalize this finding to VisualWebArena, where random-reward training raises pass@64 by 14.3% relative to the base model, without observation-matching rewards or an external multimodal teacher for reward construction.
Figures & tables
Figure 1: Overview of the experimental design and main finding. (a) World-model post-training uses agreement between predicted and observed next states as a reward. (b) We conduct our analysis experiments in four conditions that isolate the role of correct prediction targets. (c) In ALFWorld, higher prediction accuracy does not necessarily imply better task performance. Agents trained with mismatched targets retain substantial task gains despite less accurate predictions.
Condition
Target
Reward
Purpose
BASE
none
none
untrained reference
GT
ground truth
embedding similarity
RWML method
MIS
mismatched
embedding similarity
content placebo
COIN
unused
Bernoulli( 0.5 )
zero-information control
Table 1: Core RL conditions and their control roles.
ALFWorld
Condition
Domain score: pass@1 (%) ↑
pass@1 (%)
pass@64 (%)
Pred. (%)
AF (%)
LR (%)
Turns
Pick Clean Heat Cool Look Pick2
List regime
BASE
30.48 4.93 6.89 5.64 17.04 0.15
11.48
56.20
45.75
63.28
24.69
28.96
GRPO-GT
70.90 23.14 32.13 30.26 58.77 5.56
37.30
83.94
57.00
19.35
3.26
23.41
GRPO-MIS
63.45 29.15 31.37 22.62 50.76 4.73
34.55
83.21
27.25
14.27
2.13
24.38
Table 2: Task-type pass@1 and overall results by evaluation regime and training method. AF is action failure rate and LR is loop ratio. Turns is the mean number of interaction steps per episode. RWML reports three-run ALFWorld success of 13.0±1.3% for ReAct and 32.6±2.1% for GRPO-GT using Qwen2.5-7B-Instruct ( Yu et al., 2026b ) .
Figure 2: Pass@ k curves across training conditions on ALFWorld in the list ( left ) and no-list ( middle ) regimes, and ScienceWorld in the no-list regime ( right ). Shaded regions show task-level bootstrap 95% intervals.
Figure 3: Prediction preferences ( left ) and action-distribution overlap without CoT ( middle ) and with CoT ( right ). Error bars show 95% bootstrap intervals. Dashed lines mark group means, and the arrow shows their difference.
Figure 4: An MIS decision with action revision. Gray Shading highlights reasoning about candidate actions, Red Underlining marks the incorrect prediction, and Green marks the final correct action. Circled numbers indicate the order of candidate mentions.
Condition
List
No-list
Tokens
Entity (%)
Revision (%)
Candidates (%)
Tokens
Entity (%)
Revision (%)
Candidates (%)
BASE
113
45.6
6.0
7.7
98
7.5
0.7
1.0
GRPO-GT
122
63.6
11.7
12.3
99
43.3
2.3
2.0
GRPO-MIS
132
57.7
9.3
9.7
112
40.2
4.0
3.7
GRPO-COIN
150
53.0
11.7
12.7
131
27.9
5.7
8.3
OPSD-GT
137
64.5
11.7
12.7
113
39.7
4.7
5.0
Table 3: Chain-of-thought characteristics on ALFWorld. Tokens is the mean output length per step.
Figure 5: VisualWebArena reward construction ( left ) and pass@ k for BASE and COIN ( right ).
Website
Condition
pass@1 (%)
pass@64 (%)
LR (%)
Tokens/step
Turns
Classifieds
BASE
11.59
33.33
62.33
94.85
12.80
COIN
12.45
37.50
51.81
108.83
11.20
Reddit
BASE
2.53
16.22
50.80
94.14
8.39
COIN
2.91
16.22
39.68
105.92
7.15
Shopping
BASE
16.20
47.06
46.45
93.53
7.56
COIN
18.24
55.88
36.42
106.75
7.11
Table 4: Task performance and interaction statistics on VisualWebArena.
BASE
COIN
Termination
Episodes
%
Episodes
%
Agent stop command
5,444
42.32
6,822
53.03
Repeated-action limit
5,583
43.40
4,703
36.56
Turn budget exhausted
1,370
10.65
962
7.48
Consecutive parsing failures
401
3.12
310
2.41
Table 5: Episode termination types.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
GRPO
OPSD
Parameter
ALFWorld
ScienceWorld
VisualWebArena
ALFWorld
Optimizer
AdamW
AdamW
AdamW
AdamW
Learning rate
10−6
10−6
10−6
10−6
Adam β1,β2
0.9, 0.999
0.9, 0.999
0.9, 0.999
0.9, 0.999
Weight decay
0.01
0.01
0.01
0.01
Learning rate schedule
Constant
Constant
Constant
Constant
Appendix
Table 6: Training hyperparameters. Batch and mini-batch sizes count prompts, rollout n counts responses per prompt, and micro-batch size counts responses per GPU. Training samples are counted before length filtering.
ALFWorld
Condition
Domain score: pass@1 (%) ↑
pass@1 (%)
pass@64 (%)
Pred. (%)
AF (%)
LR (%)
Turns
Pick Clean Heat Cool Look Pick2
List regime
GT
70.90 23.14 32.13 30.26 58.77 5.56
37.30
83.94
57.00
19.35
3.26
23.41
GT-PERM
57.87 21.36 20.99 19.57 41.53 3.16
28.43
71.53
52.00
23.25
4.90
24.86
Δ
-13.03 -1.78 -11.14 -10.70 -17.24 -2.40
-8.87
-12.41
-5.00
+3.90
+1.64
+1.45
Appendix
Table 7: Task-type pass@1 and overall results before and after reward permutation on ALFWorld. Δ is PERM minus the original condition, computed before rounding. Rate changes are in %, and Turns changes are steps per episode.
Figure 6: Trained agents compare or rule out candidate actions, including under mismatched supervision and random rewards.
Recent studies on world modeling for Large Language Model (LLM) agents typically formulate the learning objective as next-observation prediction. However, this objective ties supervision to what a transition happens to reveal, which may omit the dynamics most relevant to the agent's current decision. To bridge this gap, we propose Agent-Authored World Modeling (AAWM), a training procedure that constructs supervision from the policy's own decision needs. Specifically, at each state, the agent identifies what it needs to understand about the environment before acting. These needs drive the retrieval of relevant transition evidence across trajectories, which is then synthesized into training targets that capture decision-oriented dynamics instead of reconstructing the next observation. This aligns the training objective with the dynamics the policy needs before acting, not with the contents of the next observation. Experimental results validate the effectiveness of AAWM across multiple environments and training settings. These results show that decision-aware world-model targets provide a more effective learning signal than next-observation prediction.
Live future prediction refers to the task of making predictions about real-world events before they unfold. This task is increasingly studied using large language model-based agent systems, and it is important for building agents that can continually learn from the real world. It can provide a large number of prediction questions grounded in diverse real-world events, while preventing answer leakage. To leverage the advantages of future prediction, we present FutureWorld, a live agentic reinforcement learning environment that closes the training loop between prediction, outcome realization, and parameter updates. Specifically, we modify and extend verl-tool, resulting in a new framework that we call verl-tool-future. Unlike standard reinforcement learning training frameworks that rely on immediate rewards, verl-tool-future stores prediction-time rollouts, backfills rewards after real-world outcomes become available, and then replays the completed trajectories for policy update. Across three open-source agents, successive FutureWorld training rounds lead to consistent improvements in prediction accuracy, probabilistic scoring, and calibration, demonstrating that delayed real-world outcome feedback can serve as an effective reinforcement learning signal.
Zhixin Han, Yanzhi Zhang, Chuyang Wei +11
College of Software, Nankai University · Zhongguancun Academy, Beijing, China · Academy of Mathematics and Systems Science, Chinese Academy of Sciences +3
Text-agent environments are typically modeled as partially observable Markov decision processes (POMDPs), assuming that the simulator's latent state and transition dynamics are hidden from the agent. Yet little work has examined whether executable code can be induced to serve as a world model for prediction and planning under partial observability. We introduce PatchWorld, a gradient-free framework that turns offline trajectories into executable Python world models through counterexample-guided code repair. Instead of predicting the next observation with a black-box model, PatchWorld induces symbolic belief-state programs whose action updates can be inspected, replayed, and locally patched. Across seven AgentGym environments, PatchWorld-Simple achieves the highest code-based planning score among evaluated methods, reaching 76.4% macro success in live one-step lookahead while invoking no LLM calls inside the world-model prediction module itself. We further find that a human-specified residual-memory bias improves surface observation fidelity but weakens decision utility. This exposes a tradeoff in executable world models, since improving observation fidelity can come at the expense of action-discriminative dynamics, and vice versa. Code is available at https://github.com/HKBU-KnowComp/PatchWorld.
Jiaxin Bai, Yue Guo, Yifei Dong +13
1Hong Kong Baptist University · 2Independent Researcher · 3HKUST +4