Authors: Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
Organizations: The Chinese University of Hong Kong · Tianjin University · Shanxi University · Independent Researcher · CAIR, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences · Nanjing University
Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.
Figures & tables
Figure 1: Jev judges what its input describes. It fails when a prediction must be made first, and recovers when code supplies it. Top: one ALFWorld decision in a kitchen scene. (a) Jev sees only the commands and picks the drawer, which the task mentions. (b) Code runs every command in a copy of the game and adds the result to the option, which reveals why the sink matters. Bottom: accuracy in one-shot games (left) and at ALFWorld steps (middle), and ALFWorld games solved out of 134 as the lookahead grows (right). Seen and unseen are two evaluation splits defined by ALFWorld. The middle panel tests single decisions: at each step of the 116 seen games that ALFWorld’s hand-coded expert solves, Jev chooses the next command given the expert’s history so far. The dashed line in the left panel marks chance (1/3). The hatched bar replaces Jev with a simple rule that picks the option sharing the most words with the task, given the same lookahead.
Four options
Published, open-ended answers ( Hagendorff et al., 2023 )
Jev
Humans (n=455)
GPT-3 davinci-003
ChatGPT-3.5
ChatGPT-4
Lure questions, correct
0.99
0.38
0.05
0.59
0.96
Lure questions, intuitive
0.01
0.55
0.80
0.15
0.00
Control versions, correct
0.96
–
–
–
–
Table 1: Jev avoids the intuitive answer on the CRT. Fractions over the 150 questions, each in four option orders. In the control versions, the intuitive answer is correct. Published results are open-ended answers, shown for reference only. The 95% interval of Jev’s correct rate on lure questions is [0.97, 1.00].
Figure 2: One-shot games: failures follow the surface shortcut, not depth. (a) An example game. (b) Accuracy is how often Jev chooses the equilibrium action (the right answer). Error bars are 95% bootstrap intervals over games. In aligned games the surface shortcut picks the equilibrium, and in misleading games it does not. Plain, reason and told are the three conditions of Section 4.1 . Reason does not help, while told restores accuracy at every depth.
Figure 3: ALFWorld: Jev is right when the surface shortcut is right. (a) Jev’s top choice at five decisions of one game, with probabilities averaged over two option orders. (b) Accuracy is how often Jev chooses an acceptable command. Error bars are 95% bootstrap intervals over games, and dashes mark uniform random choice. The aligned/misleading split was defined on the unseen games and tested on the seen games.
Predict
Decide
Decisions
Missing z
z alone
no z given
correct z given
own z given
Games, misleading
Column’s action
0.90
0.33
0.98
0.94
ALFWorld, ordered form
next subgoal
1.00
0.68
0.86
0.86
ALFWorld, goal form
next subgoal
0.33
0.06
0.92
0.33
Table 2: Jev can predict in a separate call and use a given prediction, but not both in one call. z is the missing prediction. Own z is Jev’s answer from the separate call. Rows are the 90 misleading games and the 64 misleading ALFWorld steps of the seen games from task types that have both an ordered and a goal form.
Figure 4: Stacker: Jev follows a new instruction by choosing among simulated skills. Each frame shows the scene after one decision, with the chosen skill and Jev’s probability for it. Code executes and simulates every skill, and Jev never outputs a torque. On the same scene, the hand-written controller carries box0 to the red target marker at its fourth decision (right).
Raw torques, no simulation
Skills simulated by code
Controller
Native task
Native task
5 new instructions
Jev
0/20 (0%)
9/20 (45%)
22/40 (55%)
Hand-written controller
–
8/20 (40%)
3/40 (8%)
Random choice
0/20 (0%)
1/20 (5%)
0/40 (0%)
Table 3: Stacker without and with physics simulation performed by code. Success counts over episodes not already solved at reset. Random chooses uniformly among the same actions. The hand-written controller is the state machine for the native task, run unchanged.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
plain
told
Decision rule
KL
choice
λ
KL
choice
λ
Level-1: average own payoff
0.068
0.77
5.5
0.771
0.57
6.2
Maxmax: best own payoff
0.137
0.66
4.2
0.919
0.50
3.4
Maxmin: worst own payoff
0.172
0.66
3.0
0.889
0.55
3.3
Level-2: best reply to a level-1 opponent
0.192
0.65
1.9
0.485
0.89
7.5
Equilibrium action
0.203
0.65
0.8
0.076
0.99
4.2
Appendix
Table 4: Jev plays as if the opponent acted at random (the logit level-1 rule). One-parameter models P(a)∝exp(λv(a)) are fitted to Jev’s probabilities over 630 calls per condition. The columns give the KL divergence from Jev’s distribution (lower is better), the fraction of Jev’s choices the model predicts, and the fitted λ per 100 payoff units. Bold marks the best fit per condition.
Unseen (shortcut defined)
Seen (shortcut tested)
Decision type
Jev
random
Jev
random
E (take, put or clean)
0.96 [0.94, 0.98]
0.04
0.93 [0.90, 0.96]
0.04
S-known, aligned
0.95 [0.89, 1.00]
0.08
0.99 [0.97, 1.00]
0.07
S-known, misleading
0.43 [0.33, 0.53]
0.04
0.36 [0.26, 0.48]
0.04
S-search, aligned
0.79 [0.72, 0.85]
0.66
0.79 [0.74, 0.85]
0.74
S-search, misleading
0.34 [0.24, 0.43]
0.55
0.50 [0.36, 0.64]
0.55
Appendix
Table 5: ALFWorld decisions on both splits. The decisions are replayed from the expert’s successful trajectories. The table gives the fraction of acceptable choices, with 95% bootstrap intervals over games, and the rate under uniform random choice. The surface shortcut was defined on the unseen split and tested on the seen split.
Condition
Solved
95% CI
Run 2
Pick
Look
Clean
Heat
Cool
Two
Stuck
Tokens
Jev alone
42/134
[0.24, 0.40]
39/134
0.58
0.56
0.19
0.09
0.05
0.53
54/92
10.6M
Jev + lookahead v1 (next observation)
79/134
[0.51, 0.67]
84/134
0.96
0.72
0.42
0.39
0.19
1.00
45/55
9.2M
Jev + lookahead v2 (+ new commands)
116/134
[0.81, 0.92]
117/134
0.96
0.94
0.77
0.78
0.81
1.00
14/18
5.9M
Word-overlap rule + lookahead v2
18/134
[0.08, 0.19]
–
0.17
0.44
0.13
0.04
0.05
0.00
1/116
–
Word-overlap rule + lookahead v1
7/134
[0.01, 0.09]
–
0.12
0.22
0.00
0.00
0.00
0.00
1/127
–
Random
1/134
[0.00, 0.02]
–
0.00
0.06
0.00
0.00
0.00
0.00
0/133
–
Appendix
Table 6: ALFWorld games solved in closed loop. 134 unseen games, at most 50 steps. The columns from Pick to Two give the fraction solved for each task type in the first run, and Run 2 repeats the Jev conditions. Stuck counts failed games whose last ten commands contain at most three distinct commands. Tokens are Jev’s input tokens. The game copy shows what each command would reveal, so these numbers are not comparable to agents without a simulator.
Figure 5: Jev knows what comes first only when the task mentions it first. Misleading ALFWorld steps of both splits under three task templates. Code swaps the template and changes nothing else. Left: Jev predicts the next subgoal in a separate call. Right: Jev chooses a command.
Figure 6: When Jev is wrong, it picks the lure. The decisions with a lure in each domain, scored as correct, lure or other error ( Hagendorff et al., 2023 ) . The panels cover the 150 CRT questions, the 90 misleading games with D≥1 and the 70 misleading S-known decisions of the ALFWorld seen split. The lure is the intuitive answer in the CRT and the action with the highest average payoff in games. In ALFWorld it is any wrong command that mentions an object or place named in the task description. The color of the correct segment marks the condition, as in the other figures. Light aqua marks partial help from code. Here code writes back Jev’s own answer to a separate question. Rates are averaged over items and option orders.
Experiment
Input tokens
CRT
0.73M
Games
1.18M
Game probes
4.1M
ALFWorld subgoal probe
1.3M
ALFWorld template swaps
1.1M + 0.7M
ALFWorld decisions: unseen / seen / seen with lookahead