Authors: Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang
Organizations: The Chinese University of Hong Kong · Tianjin University · Shanxi University · Independent Researcher · CAIR, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences · Nanjing University
Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.
Figures & tables
Figure 1: Jev judges what its input describes. It fails when a prediction must be made first, and recovers when code supplies it. Top: one ALFWorld decision in a kitchen scene. (a) Jev sees only the commands and picks the drawer, which the task mentions. (b) Code runs every command in a copy of the game and adds the result to the option, which reveals why the sink matters. Bottom: accuracy in one-shot games (left) and at ALFWorld steps (middle), and ALFWorld games solved out of 134 as the lookahead grows (right). Seen and unseen are two evaluation splits defined by ALFWorld. The middle panel tests single decisions: at each step of the 116 seen games that ALFWorld’s hand-coded expert solves, Jev chooses the next command given the expert’s history so far. The dashed line in the left panel marks chance (1/3). The hatched bar replaces Jev with a simple rule that picks the option sharing the most words with the task, given the same lookahead.
Four options
Published, open-ended answers ( Hagendorff et al., 2023 )
Jev
Humans (n=455)
GPT-3 davinci-003
ChatGPT-3.5
ChatGPT-4
Lure questions, correct
0.99
0.38
0.05
0.59
0.96
Lure questions, intuitive
0.01
0.55
0.80
0.15
0.00
Control versions, correct
0.96
–
–
–
–
Table 1: Jev avoids the intuitive answer on the CRT. Fractions over the 150 questions, each in four option orders. In the control versions, the intuitive answer is correct. Published results are open-ended answers, shown for reference only. The 95% interval of Jev’s correct rate on lure questions is [0.97, 1.00].
Figure 2: One-shot games: failures follow the surface shortcut, not depth. (a) An example game. (b) Accuracy is how often Jev chooses the equilibrium action (the right answer). Error bars are 95% bootstrap intervals over games. In aligned games the surface shortcut picks the equilibrium, and in misleading games it does not. Plain, reason and told are the three conditions of Section 4.1 . Reason does not help, while told restores accuracy at every depth.
Figure 3: ALFWorld: Jev is right when the surface shortcut is right. (a) Jev’s top choice at five decisions of one game, with probabilities averaged over two option orders. (b) Accuracy is how often Jev chooses an acceptable command. Error bars are 95% bootstrap intervals over games, and dashes mark uniform random choice. The aligned/misleading split was defined on the unseen games and tested on the seen games.
Predict
Decide
Decisions
Missing z
z alone
no z given
correct z given
own z given
Games, misleading
Column’s action
0.90
0.33
0.98
0.94
ALFWorld, ordered form
next subgoal
1.00
0.68
0.86
0.86
ALFWorld, goal form
next subgoal
0.33
0.06
0.92
0.33
Table 2: Jev can predict in a separate call and use a given prediction, but not both in one call. z is the missing prediction. Own z is Jev’s answer from the separate call. Rows are the 90 misleading games and the 64 misleading ALFWorld steps of the seen games from task types that have both an ordered and a goal form.
Figure 4: Stacker: Jev follows a new instruction by choosing among simulated skills. Each frame shows the scene after one decision, with the chosen skill and Jev’s probability for it. Code executes and simulates every skill, and Jev never outputs a torque. On the same scene, the hand-written controller carries box0 to the red target marker at its fourth decision (right).
Raw torques, no simulation
Skills simulated by code
Controller
Native task
Native task
5 new instructions
Jev
0/20 (0%)
9/20 (45%)
22/40 (55%)
Hand-written controller
–
8/20 (40%)
3/40 (8%)
Random choice
0/20 (0%)
1/20 (5%)
0/40 (0%)
Table 3: Stacker without and with physics simulation performed by code. Success counts over episodes not already solved at reset. Random chooses uniformly among the same actions. The hand-written controller is the state machine for the native task, run unchanged.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
plain
told
Decision rule
KL
choice
λ
KL
choice
λ
Level-1: average own payoff
0.068
0.77
5.5
0.771
0.57
6.2
Maxmax: best own payoff
0.137
0.66
4.2
0.919
0.50
3.4
Maxmin: worst own payoff
0.172
0.66
3.0
0.889
0.55
3.3
Level-2: best reply to a level-1 opponent
0.192
0.65
1.9
0.485
0.89
7.5
Equilibrium action
0.203
0.65
0.8
0.076
0.99
4.2
Appendix
Table 4: Jev plays as if the opponent acted at random (the logit level-1 rule). One-parameter models P(a)∝exp(λv(a)) are fitted to Jev’s probabilities over 630 calls per condition. The columns give the KL divergence from Jev’s distribution (lower is better), the fraction of Jev’s choices the model predicts, and the fitted λ per 100 payoff units. Bold marks the best fit per condition.
Unseen (shortcut defined)
Seen (shortcut tested)
Decision type
Jev
random
Jev
random
E (take, put or clean)
0.96 [0.94, 0.98]
0.04
0.93 [0.90, 0.96]
0.04
S-known, aligned
0.95 [0.89, 1.00]
0.08
0.99 [0.97, 1.00]
0.07
S-known, misleading
0.43 [0.33, 0.53]
0.04
0.36 [0.26, 0.48]
0.04
S-search, aligned
0.79 [0.72, 0.85]
0.66
0.79 [0.74, 0.85]
0.74
S-search, misleading
0.34 [0.24, 0.43]
0.55
0.50 [0.36, 0.64]
0.55
Appendix
Table 5: ALFWorld decisions on both splits. The decisions are replayed from the expert’s successful trajectories. The table gives the fraction of acceptable choices, with 95% bootstrap intervals over games, and the rate under uniform random choice. The surface shortcut was defined on the unseen split and tested on the seen split.
Condition
Solved
95% CI
Run 2
Pick
Look
Clean
Heat
Cool
Two
Stuck
Tokens
Jev alone
42/134
[0.24, 0.40]
39/134
0.58
0.56
0.19
0.09
0.05
0.53
54/92
10.6M
Jev + lookahead v1 (next observation)
79/134
[0.51, 0.67]
84/134
0.96
0.72
0.42
0.39
0.19
1.00
45/55
9.2M
Jev + lookahead v2 (+ new commands)
116/134
[0.81, 0.92]
117/134
0.96
0.94
0.77
0.78
0.81
1.00
14/18
5.9M
Word-overlap rule + lookahead v2
18/134
[0.08, 0.19]
–
0.17
0.44
0.13
0.04
0.05
0.00
1/116
–
Word-overlap rule + lookahead v1
7/134
[0.01, 0.09]
–
0.12
0.22
0.00
0.00
0.00
0.00
1/127
–
Random
1/134
[0.00, 0.02]
–
0.00
0.06
0.00
0.00
0.00
0.00
0/133
–
Appendix
Table 6: ALFWorld games solved in closed loop. 134 unseen games, at most 50 steps. The columns from Pick to Two give the fraction solved for each task type in the first run, and Run 2 repeats the Jev conditions. Stuck counts failed games whose last ten commands contain at most three distinct commands. Tokens are Jev’s input tokens. The game copy shows what each command would reveal, so these numbers are not comparable to agents without a simulator.
Figure 5: Jev knows what comes first only when the task mentions it first. Misleading ALFWorld steps of both splits under three task templates. Code swaps the template and changes nothing else. Left: Jev predicts the next subgoal in a separate call. Right: Jev chooses a command.
Figure 6: When Jev is wrong, it picks the lure. The decisions with a lure in each domain, scored as correct, lure or other error ( Hagendorff et al., 2023 ) . The panels cover the 150 CRT questions, the 90 misleading games with D≥1 and the 70 misleading S-known decisions of the ALFWorld seen split. The lure is the intuitive answer in the CRT and the action with the highest average payoff in games. In ALFWorld it is any wrong command that mentions an object or place named in the task description. The color of the correct segment marks the condition, as in the other figures. Light aqua marks partial help from code. Here code writes back Jev’s own answer to a separate question. Rates are averaged over items and option orders.
Experiment
Input tokens
CRT
0.73M
Games
1.18M
Game probes
4.1M
ALFWorld subgoal probe
1.3M
ALFWorld template swaps
1.1M + 0.7M
ALFWorld decisions: unseen / seen / seen with lookahead
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access. In one case study, two frontier LLM judges scored a plausible agent response 0.85 and higher. But the trace told a different story: the agent had never retrieved the artifact its answer depended on, yielding a GroundEval score of 0.000. We introduce GroundEval, a judge-free framework for evaluating agents against grounded, time-bounded, and access-controlled evidence. GroundEval uses a domain configuration to generate questions, lets the agent choose how to answer, and then scores both the final answer and the recorded trajectory that produced it. The benchmark targets three failures that LLM-as-judge evaluation struggles to detect: whether an agent checked before claiming absence, reasoned only from evidence available to the actor at the relevant time, and used the correct causal mechanism rather than a plausible one. These correspond to three tracks: Silence, Perspective, and Counterfactual. GroundEval exposes when plausible answers rest on invalid evidence paths, and produces structured per-question diagnostics that pair tool activity with the agent's turn-level narration, making each score inspectable rather than merely reported. Our case studies suggest this failure mode is common rather than exceptional, one that final-answer and judge-based evaluation cannot detect by construction.
Dedicated decision models such as Jev map unstructured language to probability distributions over finite choices, allowing their outputs to directly route requests, select tools, and trigger actions. Yet real-world inputs rarely arrive in isolation: they come with background details and surrounding context. We find that short additions that fit naturally into this context can nevertheless redirect an otherwise correct decision, even when the correct answer remains unchanged. To study this behavior, we fix a wrong target option for each initially correct item and use the model's option probabilities to refine fluent context additions while preserving the source, question, choices, and gold answer. Within 64 accepted target evaluations, the optimizer identifies contexts that redirect Jev on 312 of 508 initially correct decisions (61.4%); in 229 cases, Jev assigns at least 0.7 probability to the fixed wrong option. Across seven datasets, three additional decision systems show targeted flip rates of 64.9%-73.2% on decisions they initially answer correctly. Taken together, these results expose a pronounced fragility in current decision models: short, ordinary-looking context can shift a correct choice to a high-confidence wrong one. Because these models turn language directly into downstream choices, this sensitivity raises concerns about treating their probability outputs as reliable decision interfaces.