Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training
Authors: Xinyu Che, Hang Yan, Yanchen Liu, Haochen Liu, Ruifeng Li, Anran Shi, Heng Wang, Jun Liu
Organizations: Xi’an Jiaotong University · University of Southern California · University of the Chinese Academy of Sciences · East China Normal University
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. Existing next-observation training methods help the agent to learn the environmental content. However, they additionally involve an optimization process, which may introduce several effects other than learning to predict the world. Consequently, where the performance gain comes from during the training process remains an open question. We answer this research question through replacing true next-observation targets with in-distribution mismatched observations during the training process. Across two interactive text environments, mismatched targets lower prediction accuracy by 15.3-61.6% relative to ground-truth targets, yet retain substantial task gains over the base model. Compared with the base model, trained models consider more candidate actions and exhibit less looping. We also introduce a setting that replaces prediction-based rewards with independent random signals. This training expands task coverage (pass@64) even when the reward carries no environment information. We also generalize this finding to VisualWebArena, where random-reward training raises pass@64 by 14.3% relative to the base model, without observation-matching rewards or an external multimodal teacher for reward construction.
Figures & tables
Figure 1: Overview of the experimental design and main finding. (a) World-model post-training uses agreement between predicted and observed next states as a reward. (b) We conduct our analysis experiments in four conditions that isolate the role of correct prediction targets. (c) In ALFWorld, higher prediction accuracy does not necessarily imply better task performance. Agents trained with mismatched targets retain substantial task gains despite less accurate predictions.
Condition
Target
Reward
Purpose
BASE
none
none
untrained reference
GT
ground truth
embedding similarity
RWML method
MIS
mismatched
embedding similarity
content placebo
COIN
unused
Bernoulli( 0.5 )
zero-information control
Table 1: Core RL conditions and their control roles.
ALFWorld
Condition
Domain score: pass@1 (%) ↑
pass@1 (%)
pass@64 (%)
Pred. (%)
AF (%)
LR (%)
Turns
Pick Clean Heat Cool Look Pick2
List regime
BASE
30.48 4.93 6.89 5.64 17.04 0.15
11.48
56.20
45.75
63.28
24.69
28.96
GRPO-GT
70.90 23.14 32.13 30.26 58.77 5.56
37.30
83.94
57.00
19.35
3.26
23.41
GRPO-MIS
63.45 29.15 31.37 22.62 50.76 4.73
34.55
83.21
27.25
14.27
2.13
24.38
Table 2: Task-type pass@1 and overall results by evaluation regime and training method. AF is action failure rate and LR is loop ratio. Turns is the mean number of interaction steps per episode. RWML reports three-run ALFWorld success of 13.0±1.3% for ReAct and 32.6±2.1% for GRPO-GT using Qwen2.5-7B-Instruct ( Yu et al., 2026b ) .
Figure 2: Pass@ k curves across training conditions on ALFWorld in the list ( left ) and no-list ( middle ) regimes, and ScienceWorld in the no-list regime ( right ). Shaded regions show task-level bootstrap 95% intervals.
Figure 3: Prediction preferences ( left ) and action-distribution overlap without CoT ( middle ) and with CoT ( right ). Error bars show 95% bootstrap intervals. Dashed lines mark group means, and the arrow shows their difference.
Figure 4: An MIS decision with action revision. Gray Shading highlights reasoning about candidate actions, Red Underlining marks the incorrect prediction, and Green marks the final correct action. Circled numbers indicate the order of candidate mentions.
Condition
List
No-list
Tokens
Entity (%)
Revision (%)
Candidates (%)
Tokens
Entity (%)
Revision (%)
Candidates (%)
BASE
113
45.6
6.0
7.7
98
7.5
0.7
1.0
GRPO-GT
122
63.6
11.7
12.3
99
43.3
2.3
2.0
GRPO-MIS
132
57.7
9.3
9.7
112
40.2
4.0
3.7
GRPO-COIN
150
53.0
11.7
12.7
131
27.9
5.7
8.3
OPSD-GT
137
64.5
11.7
12.7
113
39.7
4.7
5.0
Table 3: Chain-of-thought characteristics on ALFWorld. Tokens is the mean output length per step.
Figure 5: VisualWebArena reward construction ( left ) and pass@ k for BASE and COIN ( right ).
Website
Condition
pass@1 (%)
pass@64 (%)
LR (%)
Tokens/step
Turns
Classifieds
BASE
11.59
33.33
62.33
94.85
12.80
COIN
12.45
37.50
51.81
108.83
11.20
Reddit
BASE
2.53
16.22
50.80
94.14
8.39
COIN
2.91
16.22
39.68
105.92
7.15
Shopping
BASE
16.20
47.06
46.45
93.53
7.56
COIN
18.24
55.88
36.42
106.75
7.11
Table 4: Task performance and interaction statistics on VisualWebArena.
BASE
COIN
Termination
Episodes
%
Episodes
%
Agent stop command
5,444
42.32
6,822
53.03
Repeated-action limit
5,583
43.40
4,703
36.56
Turn budget exhausted
1,370
10.65
962
7.48
Consecutive parsing failures
401
3.12
310
2.41
Table 5: Episode termination types.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
GRPO
OPSD
Parameter
ALFWorld
ScienceWorld
VisualWebArena
ALFWorld
Optimizer
AdamW
AdamW
AdamW
AdamW
Learning rate
10−6
10−6
10−6
10−6
Adam β1,β2
0.9, 0.999
0.9, 0.999
0.9, 0.999
0.9, 0.999
Weight decay
0.01
0.01
0.01
0.01
Learning rate schedule
Constant
Constant
Constant
Constant
Appendix
Table 6: Training hyperparameters. Batch and mini-batch sizes count prompts, rollout n counts responses per prompt, and micro-batch size counts responses per GPU. Training samples are counted before length filtering.
ALFWorld
Condition
Domain score: pass@1 (%) ↑
pass@1 (%)
pass@64 (%)
Pred. (%)
AF (%)
LR (%)
Turns
Pick Clean Heat Cool Look Pick2
List regime
GT
70.90 23.14 32.13 30.26 58.77 5.56
37.30
83.94
57.00
19.35
3.26
23.41
GT-PERM
57.87 21.36 20.99 19.57 41.53 3.16
28.43
71.53
52.00
23.25
4.90
24.86
Δ
-13.03 -1.78 -11.14 -10.70 -17.24 -2.40
-8.87
-12.41
-5.00
+3.90
+1.64
+1.45
Appendix
Table 7: Task-type pass@1 and overall results before and after reward permutation on ALFWorld. Δ is PERM minus the original condition, computed before rounding. Rate changes are in %, and Turns changes are steps per episode.
Figure 6: Trained agents compare or rule out candidate actions, including under mismatched supervision and random rewards.
College of Software, Nankai University · Zhongguancun Academy, Beijing, China · Academy of Mathematics and Systems Science, Chinese Academy of Sciences +3