Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or specialized checks to block errors, but deciding before execution whether intervention will benefit eventual task completion remains challenging. To address this challenge, we propose HiSentinel, a hindsight-distillation framework that trains lightweight 0.6B and 1.7B sentinels to select pre-execution interventions aimed at improving task completion rather than correcting every imperfect action. A privileged teacher uses recorded execution outcomes as evidence for intervention judgments, which are distilled into a causal student that receives only the pre-action context and proposed action. Beyond identifying whether and when to intervene, the sentinel must also provide actionable feedback that helps the coding agent recover or obtain necessary human input. To support these capabilities, we introduce SWE-Intervene, an action-level dataset constructed from software-engineering trajectories that annotates whether an action should be allowed, autonomously redirected, or paused for human assistance, together with corresponding intervention feedback. Across SWE-bench Verified Mini and Ask or Assume, HiSentinel consistently improves task completion across Sentinel scales and coding-agent families, with gains of up to 14% and 10%, respectively, while maintaining competitive token consumption. These results demonstrate that lightweight pre-execution intervention can effectively prevent error propagation and improve the reliability of autonomous coding agents.
Figures & tables
Figure 1: Pre-execution intervention for coding agents. The Sentinel evaluates each proposed action before execution and selects Allow , Redirect , or Hard-Pause .
Figure 2: Overview of HiSentinel . Privileged hindsight is used only during training, while the deployed causal sentinel selects an intervention before the proposed action is executed.
Action
Operational definition
Allow
Execute the proposed action without intervention because there is insufficient evidence that intervening would improve task completion.
Redirect
Withhold the proposed action and provide corrective feedback, prompting the agent to revise its plan and propose a new action autonomously.
Hard-Pause
Suspend execution and request missing information, authorization, or preferences from a human before the agent continues.
Table 1: Intervention actions used by HiSentinel .
Agent
Backbone
Method
SWE-V Mini
Ask or Assume
Res. ↑
Cost ↓
Res. ↑
Cost ↓
Sonnet- 4.6
–
Base
60.00
0.26
58.00
0.47
Qwen3-1.7B
HiSentinel
66.00
0.41
61.00
0.82
Qwen3- Coder
–
Base
30.00
0.29
23.00
0.26
Reflexion
34.00
0.31
20.00
0.36
SDS-4B
SDS
28.00
0.57
24.00
0.44
Table 2: End-to-end results. SWE-V Mini denotes SWE-bench Verified Mini; SDS denotes Steer, Don’t Solve . Res.: resolved (%). Cost: average input and output tokens per trajectory (millions).
Backbone
Method
In-Domain
OOD
SWE-Intervene
RootSE
R-Judge
M-F1 ↑
I-F1 ↑
Cost ↓
T@0 ↑
T@1 ↑
Cost ↓
BAcc ↑
M-F1 ↑
Cost ↓
Qwen3-0.6B
Step-by-Step
27.23
21.21
8.08K
3.92
10.78
0.12M
60.68
60.61
0.58K
All-at-Once
28.88
29.79
8.08K
3.92
6.86
0.07M
60.02
60.77
0.66K
TrajAudit
30.23
19.13
10.16K
0.00
3.92
0.10M
50.66
34.92
1.80K
RCTA
28.61
19.75
11.28K
2.94
10.78
0.03M
50.00
32.09
1.30K
Table 3: Results on static intervention-recognition benchmarks. No data from RootSE or R-Judge are used for training. M-F1 and I-F1 denote Macro-F1 and Intervention F1. BAcc denote Balanced accuracy. For RootSE, T@ k counts predictions within k steps of the annotated earliest decisive error. Cost denotes the average total number of input and output tokens per trajectory.
Figure 3: A representative causal intervention. The displayed Redirect is one local episode in the trajectory; the final resolved/unresolved labels are produced by the official SWE-bench evaluator.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Instances
Proportion
Open-SWE-Traces
3,451
49.85%
SWE-Hero
1,089
15.73%
SWE-chat
2,383
34.42%
Total
6,923
100%
Appendix
Table 5: Composition of SWE-Intervene .
Audit outcome
Instances
Provisional intervention candidates
2,252
Retained after audit
1,682
Rejected after audit
570
Rejection rate
25.3%
Appendix
Table 6: Results of the intervention-candidate audit.
Hyperparameter
Value
General configuration
Backbones
Qwen3-0.6B and Qwen3-1.7B
Future-view teacher
Frozen Qwen3-Coder-30B-A3B-Instruct
Numerical precision
BF16
Optimizer
AdamW, β1=0.9 , β2=0.95
Weight decay
0.01
Appendix
Table 7: Training hyperparameters of HiSentinel . The settings are shared by the Qwen3-0.6B and Qwen3-1.7B variants. The feedback learning rate decays to its floor after the first epoch.
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals, adding constraints, and correcting mistakes over multiple turns. We introduce SWE-Together, a multi-turn benchmark reconstructed from real user-agent coding sessions. To make real interactions verifiable, we curate 109 repository-level tasks from 11,260 recorded sessions, selecting sessions with recoverable repository states, clear user goals, and observable outcomes. To replay these interactions across agents, we build a reactive LLM-based user simulator that preserves the original users' intents and provides feedback when the coding agent's progress requires it. To evaluate agents as collaborators, we measure both final repository correctness and the number of corrective feedback turns required during the interaction. Experiments with frontier coding agents show that stronger agents generally achieve higher final success rates while requiring fewer interventions, suggesting an improved user experience.
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: structure-aware retrieval, agent-synthesized skills, and developer-designed skills, over 10k trajectories on held-out SWE-bench Verified and Pro tasks. Our main findings are: (1) The three behaviors affect 79.00%--98.00% of coding tasks and account for up to 22.75% of task cost. (2) Structure-aware retrieval can introduce retrieval overhead and alter agent delegation, causing inconsistent improvements in retrieval efficiency and cost increases of up to 28.14%. (3) Agent-synthesized skills tend to produce low-level, trace-specific guidance, limiting their effectiveness and generality. (4) In contrast, developer-designed skills provide high-level, trace-agnostic guidance, reducing cost by up to 41.73%, roughly twice the maximum gain from agent-synthesized skills.
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time. Code is publicly available at https://github.com/MSR-Orchard/execcritic.
Leitian Tao, Baolin Peng, Haorui Wang +7
University of Wisconsin–Madison · Microsoft Research · Georgia Tech