AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
Organizations: ServiceNow Research · McGill University · Mila – Quebec AI Institute · ÉTS Montréal · Seoul National University · Polytechnique Montréal · Canada CIFAR AI Chair · Université Laval
Abstract
Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect. To identify these errors, a judge needs to examine the trajectory with respect to the user's instruction. To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems. By recording trajectories for closely related instructions, we can construct negative tasks by swapping the instructions. This paired design evaluates judges on their ability to distinguish a successful trajectory from one that completed a similar (but incompatible) request. We release the benchmark under three splits: a frontier split, AgentHorizon (AH), a simplified split, AgentHorizon-Simple (AH-S), and a development split, AgentHorizon-Development (AH-D). We further evaluate eleven judges by (1) directly passing the full trajectory (with up to 300 screenshots and actions), and (2) using them as coding agents across five agent harnesses. We find that our best agentic judge, GPT-5.5, achieves 80.9% balanced accuracy on the AH subset. We find that tool-use improves certain models but results in worse performance for open-weight models, and that judges differ drastically in their ability to accept a valid trajectory and reject failed ones. Our findings highlight the need for judges that are capable of locating and verifying often hidden evidence that a task was properly completed inside long interaction histories.
Figures & tables
| AH-D | AH | AH-S | |
| Judging tasks | 162 | 528 | 683 |
| Positives | 62 | 227 | 234 |
| Adversarial negatives | 100 | 301 | 449 |
| Mistake types among negatives | |||
| Critical Mistake | 32 | 48 | 193 |
| Bad Side Effect | 26 | 107 | 99 |
| Model | Open? | Interface | AH Bal. | AH Pos | AH Neg | AH-S Bal. |
|---|---|---|---|---|---|---|
| Agentic judges | ||||||
| GPT-5.5 | Codex | 80.9 | 71.8 | 90.0 | 92.6 | |
| Gemini 3.1 Pro | Gemini CLI | 77.7 | 65.6 | 89.7 | 86.3 | |
| Claude Opus 4.7 | Claude Code | 76.0 | 76.7 | 75.4 | 94.6 | |
| Qwen 3.6 27B | ✓ | OpenCode | 70.4 | 75.3 | 65.4 | 93.6 |
| GPT-5.4 mini | Codex | 69.2 | 67.4 | 71.1 | 91.2 | |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Hours/task | Total hours |
|---|---|---|
| Annotation | ||
| Review L1 | ||
| Mistake-category review | ||
| Second-stage review | ||
| Total |
| Stage | Positive | Negative | Total |
|---|---|---|---|
| Paired construction pool | 850 | 850 | 1,700 |
| Reviewed benchmark | 523 | 850 | 1,373 |
| AH-D | 62 | 100 | 162 |
| AH | 227 | 301 | 528 |
| AH-S | 234 | 449 | 683 |
| Type | Boundary | Typical recovery |
|---|---|---|
| Critical Mistake | The core action or target is not achieved. | Restart or repeat the task; the original outcome cannot satisfy the request. |
| Bad Side Effect | The goal is achieved, but an unrequested action creates cost, risk, exposure, or cleanup. | Reverse an external action, coordinate with others, or spend substantial effort undoing it. |
| Misunderstanding | The execution is coherent but misreads a parameter or constraint without external harm. | A local edit or short correction is sufficient. |
| Model | Interface | Input tokens | Tools | Images | $/trajectory |
|---|---|---|---|---|---|
| GPT-5.5 | Codex | 513,556 | 18.6 | 11.9 | 0.4035 |
| Gemini 3.1 Pro | Gemini CLI | 1,393,271 | 24.2 | 30.2 | 0.6172 |
| Claude Opus 4.7 | Claude Code | 554,387 | 9.2 | 7.8 | 0.6592 |
| Qwen 3.6 27B | OpenCode | 243,440 | 8.1 | 6.0 | 0 |
| GPT-5.4 mini | Codex | 323,808 | 20.0 | 8.4 | 0.1066 |
| Qwen 3.6 35B-A3B | OpenCode | 190,550 | 8.1 | 5.6 | 0 |
| Model | Interface | Critical | Side effect | Misunderstanding |
|---|---|---|---|---|
| GPT-5.5 | Codex | 77.7 | 24.1 | 72.9 |
| Gemini 3.1 Pro | Gemini CLI | 57.5 | 7.3 | 87.3 |
| Claude Opus 4.7 | Claude Code | 75.8 | 12.1 | 71.4 |
| Qwen 3.6 27B | OpenCode | 60.4 | 5.6 | 74.3 |
| GPT-5.4 mini | Codex | 60.8 | 2.6 | 74.6 |
| Qwen 3.6 35B-A3B | OpenCode | 63.0 | 6.0 | 53.1 |
| Model | Interface | AH Bal. | AH Pos | AH Neg | AH-S Bal. |
|---|---|---|---|---|---|
| GPT-5.5 | Codex † | 80.9 | 71.8 | 90.0 | 92.6 |
| OpenCode | 77.2 | 70.9 | 83.4 | 92.2 | |
| Qwen 3.6 27B | Codex | 71.6 | 70.5 | 72.8 | 94.5 |
| OpenCode † | 70.4 | 75.3 | 65.4 | 93.6 | |
| OpenHands | 51.2 | 58.1 | 44.2 | 89.0 | |
| Gemini 3.1 Flash Lite | Gemini CLI † | 52.1 | 58.6 | 45.5 | 75.2 |
| Instruction | Screenshots | Action log | TP | FN | FP | TN | BA |
|---|---|---|---|---|---|---|---|
| ✗ | ✓ all | ✓ | 246 | 5 | 273 | 0 | 49.0% |
| ✓ | — final only | ✗ | 104 | 147 | 118 | 155 | 49.1% |
| ✗ | — final only | ✗ | 215 | 35 | 245 | 26 | 47.8% |
| ✓ | ✗ | ✓ | 205 | 46 | 181 | 92 | 57.7% |
| ✗ | ✗ | ✓ | 249 | 2 | 272 | 1 | 49.8% |