Coding agents solve repository-level tasks through sequences of actions, where a single erroneous action can misdirect subsequent decisions and increase recovery costs. Existing approaches use execution feedback for recovery or specialized checks to block errors, but deciding before execution whether intervention will benefit eventual task completion remains challenging. To address this challenge, we propose HiSentinel, a hindsight-distillation framework that trains lightweight 0.6B and 1.7B sentinels to select pre-execution interventions aimed at improving task completion rather than correcting every imperfect action. A privileged teacher uses recorded execution outcomes as evidence for intervention judgments, which are distilled into a causal student that receives only the pre-action context and proposed action. Beyond identifying whether and when to intervene, the sentinel must also provide actionable feedback that helps the coding agent recover or obtain necessary human input. To support these capabilities, we introduce SWE-Intervene, an action-level dataset constructed from software-engineering trajectories that annotates whether an action should be allowed, autonomously redirected, or paused for human assistance, together with corresponding intervention feedback. Across SWE-bench Verified Mini and Ask or Assume, HiSentinel consistently improves task completion across Sentinel scales and coding-agent families, with gains of up to 14% and 10%, respectively, while maintaining competitive token consumption. These results demonstrate that lightweight pre-execution intervention can effectively prevent error propagation and improve the reliability of autonomous coding agents.
Figures & tables
Figure 1: Pre-execution intervention for coding agents. The Sentinel evaluates each proposed action before execution and selects Allow , Redirect , or Hard-Pause .
Figure 2: Overview of HiSentinel . Privileged hindsight is used only during training, while the deployed causal sentinel selects an intervention before the proposed action is executed.
Action
Operational definition
Allow
Execute the proposed action without intervention because there is insufficient evidence that intervening would improve task completion.
Redirect
Withhold the proposed action and provide corrective feedback, prompting the agent to revise its plan and propose a new action autonomously.
Hard-Pause
Suspend execution and request missing information, authorization, or preferences from a human before the agent continues.
Table 1: Intervention actions used by HiSentinel .
Agent
Backbone
Method
SWE-V Mini
Ask or Assume
Res. ↑
Cost ↓
Res. ↑
Cost ↓
Sonnet- 4.6
–
Base
60.00
0.26
58.00
0.47
Qwen3-1.7B
HiSentinel
66.00
0.41
61.00
0.82
Qwen3- Coder
–
Base
30.00
0.29
23.00
0.26
Reflexion
34.00
0.31
20.00
0.36
SDS-4B
SDS
28.00
0.57
24.00
0.44
Table 2: End-to-end results. SWE-V Mini denotes SWE-bench Verified Mini; SDS denotes Steer, Don’t Solve . Res.: resolved (%). Cost: average input and output tokens per trajectory (millions).
Backbone
Method
In-Domain
OOD
SWE-Intervene
RootSE
R-Judge
M-F1 ↑
I-F1 ↑
Cost ↓
T@0 ↑
T@1 ↑
Cost ↓
BAcc ↑
M-F1 ↑
Cost ↓
Qwen3-0.6B
Step-by-Step
27.23
21.21
8.08K
3.92
10.78
0.12M
60.68
60.61
0.58K
All-at-Once
28.88
29.79
8.08K
3.92
6.86
0.07M
60.02
60.77
0.66K
TrajAudit
30.23
19.13
10.16K
0.00
3.92
0.10M
50.66
34.92
1.80K
RCTA
28.61
19.75
11.28K
2.94
10.78
0.03M
50.00
32.09
1.30K
Table 3: Results on static intervention-recognition benchmarks. No data from RootSE or R-Judge are used for training. M-F1 and I-F1 denote Macro-F1 and Intervention F1. BAcc denote Balanced accuracy. For RootSE, T@ k counts predictions within k steps of the annotated earliest decisive error. Cost denotes the average total number of input and output tokens per trajectory.
Figure 3: A representative causal intervention. The displayed Redirect is one local episode in the trajectory; the final resolved/unresolved labels are produced by the official SWE-bench evaluator.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Instances
Proportion
Open-SWE-Traces
3,451
49.85%
SWE-Hero
1,089
15.73%
SWE-chat
2,383
34.42%
Total
6,923
100%
Appendix
Table 5: Composition of SWE-Intervene .
Audit outcome
Instances
Provisional intervention candidates
2,252
Retained after audit
1,682
Rejected after audit
570
Rejection rate
25.3%
Appendix
Table 6: Results of the intervention-candidate audit.
Hyperparameter
Value
General configuration
Backbones
Qwen3-0.6B and Qwen3-1.7B
Future-view teacher
Frozen Qwen3-Coder-30B-A3B-Instruct
Numerical precision
BF16
Optimizer
AdamW, β1=0.9 , β2=0.95
Weight decay
0.01
Appendix
Table 7: Training hyperparameters of HiSentinel . The settings are shared by the Qwen3-0.6B and Qwen3-1.7B variants. The feedback learning rate decays to its floor after the first epoch.