What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents
Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Kuaishou Technology
Abstract
As agents take on long-horizon tasks, users shift from making individual decisions to overseeing autonomous execution. Yet the volume of agent activity and the fragmentation of supporting evidence make it difficult to determine which decisions warrant user verification. We study monitors that identify consequential decisions and locate evidence to help users assess their implications. We introduce AgentMonBench, a software-engineering benchmark comprising three subsets that cover two complementary dimensions: alignment between requirements and behavior, and awareness of consequential autonomous decisions for verification. To support these judgments, we propose the Evidence-Grounded Behavior Graph (EBG), a training-free method that groups source-linked evidence into behaviors and organizes their relationships into a graph. EBG presents task-oriented views of this graph to help monitors interpret behavior in context. Experiments across eight models show that EBG improves decision identification and evidence localization in most settings compared with direct access to the original context. Further experiments show that EBG's evidence-localization gains persist across input scales and hyperparameter settings, while real-world applications illustrate its practical value for human oversight.
Figures & tables
| Monitor | Method | SpecGAP | SilentSwap | FeedbackTrace | ||||||
| SF1 | LF1 | LC | CC | LS | VPA | ESS | EHR | |||
| Commercial models | ||||||||||
| Base | 0.250 | 0.627 | 0.305 | 0.330 | 0.268 | 0.403 | 0.460 | 0.545 | 0.530 | |
| GPT-5.6 Luna | RepoGraph | 0.241 | 0.587 | 0.270 | 0.208 | 0.193 | 0.293 | - | ||
| EBG | 0.326 | 0.722 | 0.418 | 0.333 | 0.288 | 0.455 | 0.485 | 0.560 | 0.570 | |
| Case | Task scenario | Monitor actions and outcomes |
| Ambiguous Optimization Objective | Metrics: accuracy and recall; the plan gives no optimization priority. | Checked the plan. Disclosed the missing priority and asked the user. The user chose accuracy. |
| Changed Randomization Protocol | A serial fallback after failure changes per-problem seeding. | Found changed randomization in the fallback code. Disclosed the change to the user and restored per-problem seeds. |
| Validation Data Leakage | Validation-derived samples enter training, yielding an apparent score gain. | Traced added samples to validation data. Disclosed the leakage and retained the baseline. |
| Unverified API Connectivity | API checks pass using cached responses, without making a live API request. | Found only a reused response in the records. Disclosed that live API connectivity remained unverified. |
| Unequal Search Budgets | The candidate scores higher than the baseline under unequal search budgets. | Found unequal budgets in search records. Disclosed that the gain could not be attributed solely to the optimizer. |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Domain | Evaluation target | DA | ES | UF |
| EvoArena ( 2026b ) | Multi-domain | Environment adaptation | – | ||
| PaperBench ( 2025 ) | Research | Research replication | – | ||
| ScienceAgentBench ( 2025 ) | Science | Scientific task execution | – | ||
| SWE-bench ( 2024 ) | Software | Issue resolution | – | ||
| Ambig-SWE ( 2026 ) | Software | Clarification and repair | Sim. | ||
| -Bench ( 2026 ) | Assistant | Hidden-intent handling | Sim. |
| Dataset | Annotations | Input tokens | |||||
| Benchmark | Examples | Count | Types | Avg. | Low | Medium | High |
| SpecGAP | 100 | 484 | 7 | 22,870.32 | 10,964 | 20,363 | 32,643 |
| SilentSwap | 100 | 500 | 8 | 22,855.91 | 12,677 | 17,956 | 31,861.5 |
| FeedbackTrace | 100 | 100 | 4 | 53,942.98 | 9,468 | 28,936 | 86,138.5 |
| Benchmark | Agreements / | Agreement |
| SpecGAP | 447/484 | 92.4% |
| SilentSwap | 468/500 | 93.6% |
| FeedbackTrace | 93/100 | 93.0% |
| Benchmark | Review outcome | Count / | Retention |
| SpecGAP | Model mapping retained after human review | 322/484 | 66.5% |
| SilentSwap | Draft retained without semantic revision | 486/500 | 97.2% |
| FeedbackTrace | Evidence set retained | 64/100 | 64.0% |
| Feedback type retained | 74/100 | 74.0% | |
| Criticality label retained | 75/100 | 75.0% |
| Name | Role | Model identifier / version | Reasoning setting |
| Luna | Monitor | gpt-5-6-luna | none |
| Terra | Monitor | gpt-5-6-terra | none |
| Sol | Monitor | gpt-5-6-sol | none |
| Flash | Monitor | deepseek-v4.1-flash | disabled |
| Pro | Monitor | deepseek-v4-pro | disabled |
| Kimi-K3 | Monitor | kimi-k3 | low |
| Setting | Base | RepoGraph |
| Reading unit | Contiguous source block, up to 400 lines | One-hop symbol subgraph around a function or class (Python) |
| Units per read | Up to 6 source blocks | Up to 32 symbol subgraphs |
| Search scope | File paths and source text | Symbol names, qualified names, and paths |
| Benchmark | Model Prediction |
| SpecGAP | A list of findings . Each finding contains a finding ID , claim , verification question , why important , downstream impact , and code evidence . Each evidence item records the source path, line range, owning Symbol, and an explanation. |
| SilentSwap | Exactly five swaps . Each swap contains a source target , swap type , code change , trigger condition , and the Before / After behavioral effects. |
| FeedbackTrace | One verification point , one or two supporting evidence IDs , and a criticality value of either Must-Disclose or Worth-Disclose . |
| Monitor | Method | SpecGAP | SilentSwap | FeedbackTrace | ||||||
| SF1 | LF1 | LC | CC | LS | VPA | ESS | EHR | |||
| Luna | Base | 0.265 | 0.636 | 0.305 | 0.385 | 0.245 | 0.403 | 0.510 | 0.545 | 0.530 |
| RepoGraph | 0.264 | 0.613 | 0.270 | 0.295 | 0.195 | 0.293 | - | |||
| EBG | 0.352 | 0.713 | 0.418 | 0.355 | 0.305 | 0.455 | 0.525 | 0.550 | 0.570 | |
| Flash | Base | 0.401 | 0.829 | 0.459 | 0.230 | 0.130 | 0.262 | 0.440 | 0.445 | 0.540 |
| RepoGraph | 0.362 | 0.718 | 0.388 | 0.240 | 0.165 | 0.344 | - | |||
| Monitor | Method | SpecGAP | SilentSwap | FeedbackTrace | ||||||
| SF1 | LF1 | LC | CC | LS | VPA | ESS | EHR | |||
| Luna | Base | 0.234 | 0.618 | 0.305 | 0.275 | 0.290 | 0.403 | 0.410 | 0.545 | 0.530 |
| RepoGraph | 0.217 | 0.561 | 0.270 | 0.120 | 0.190 | 0.293 | - | |||
| EBG | 0.300 | 0.730 | 0.418 | 0.310 | 0.270 | 0.455 | 0.445 | 0.570 | 0.570 | |
| Flash | Base | 0.350 | 0.834 | 0.459 | 0.100 | 0.120 | 0.262 | 0.390 | 0.460 | 0.540 |
| RepoGraph | 0.309 | 0.714 | 0.388 | 0.200 | 0.190 | 0.344 | - | |||
| Component | Metric | Mean abs. diff. | Max abs. diff. | Same direction | Spearman |
| SpecGAP | SF1 | 0.0336 | 0.0750 | 6/8 | 1.0000 |
| SpecGAP | 0.0219 | 0.0870 | 7/8 | 1.0000 | |
| SilentSwap | LC | 0.0969 | 0.2700 | 5/8 | 0.8929 |
| SilentSwap | CC | 0.0222 | 0.0450 | 6/8 | 0.9286 |
| FeedbackTrace | VPA | 0.0953 | 0.1550 | 8/8 | 0.8929 |
| FeedbackTrace | ESS | 0.0109 | 0.0300 | 8/8 | 0.9643 |
| Monitor | Removed | SpecGAP | SilentSwap | FeedbackTrace | ||||||
| SF1 | LF1 | LC | CC | LS | VPA | ESS | EHR | |||
| Luna | Behavior | -0.102 | -0.147 | -0.136 | -0.018 | -0.013 | -0.049 | -0.008 | -0.040 | -0.040 |
| Graph | -0.201 | -0.412 | -0.249 | -0.090 | -0.038 | -0.112 | -0.020 | -0.050 | -0.050 | |
| Guidance | -0.107 | -0.213 | -0.140 | -0.108 | -0.050 | -0.150 | – | – | – | |
| Kimi-K3 | Behavior | -0.043 | +0.005 | -0.057 | -0.048 | -0.183 | -0.090 | -0.080 | -0.115 | -0.120 |
| Graph | -0.088 | -0.068 | -0.106 | -0.083 | +0.005 | -0.120 | -0.080 | -0.128 | -0.100 | |
| Component | Group | Min. tokens | Median tokens | Max. tokens | |
| SpecGAP | Low | 33 | 4,908 | 10,964.0 | 14,735 |
| SpecGAP | Medium | 33 | 14,870 | 20,363.0 | 25,051 |
| SpecGAP | High | 34 | 25,417 | 32,643.0 | 76,478 |
| SilentSwap | Low | 33 | 4,041 | 12,677.0 | 15,025 |
| SilentSwap | Medium | 33 | 15,059 | 17,956.0 | 23,002 |
| SilentSwap | High | 34 | 24,312 | 31,861.5 | 76,256 |
| Monitor | Method | SpecGAP | SilentSwap | FeedbackTrace | |||
| TPR | FPR | TPR | FPR | TPR | FPR | ||
| GPT-5.6 Luna | Base | 64 | 28 | 84 | 52 | 90 | 88 |
| EBG | 76 | 22 | 86 | 58 | 94 | 84 | |
| Kimi K3 | Base | 76 | 22 | 86 | 38 | 50 | 40 |
| EBG | 84 | 20 | 88 | 22 | 58 | 30 | |
| GLM-5.3 | Base | 70 | 24 | 72 | 24 | 46 | 36 |
| Component | Monitor | Units | Failures | File | Evidence | Judgment |
| SpecGAP | Luna | 484 | 362 | 31 (8.6%) | 37 (10.2%) | 294 (81.2%) |
| Kimi-K3 | 484 | 258 | 31 (12.0%) | 31 (12.0%) | 196 (76.0%) | |
| GLM-5.3 | 484 | 273 | 29 (10.6%) | 42 (15.4%) | 202 (74.0%) | |
| SilentSwap | Luna | 500 | 221 | 19 (8.6%) | 32 (14.5%) | 170 (76.9%) |
| Kimi-K3 | 500 | 139 | 11 (7.9%) | 26 (18.7%) | 102 (73.4%) | |
| GLM-5.3 | 500 | 111 | 7 (6.3%) | 27 (24.3%) | 77 (69.4%) |
| Component | Monitor | Units | Failures | File | View | Judgment |
| SpecGAP | Kimi-K3 | 484 | 265 | 31 (11.7%) | 30 (11.3%) | 204 (77.0%) |
| SpecGAP | Luna | 484 | 377 | 32 (8.5%) | 40 (10.6%) | 305 (80.9%) |
| SpecGAP | GLM-5.3 | 484 | 288 | 29 (10.1%) | 43 (14.9%) | 216 (75.0%) |
| SilentSwap | Kimi-K3 | 500 | 134 | 11 (8.2%) | 26 (19.4%) | 97 (72.4%) |
| SilentSwap | Luna | 500 | 243 | 19 (7.8%) | 32 (13.2%) | 192 (79.0%) |
| SilentSwap | GLM-5.3 | 500 | 118 | 7 (5.9%) | 26 (22.0%) | 85 (72.0%) |