Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.
Figures & tables
Figure 1 : Illustration of Opera . Left: Opera discovers and tracks issues at different stages (search, edit, test, etc.) until it is resolved or refined in the long-horizon turns. Right: Opera improves task resolve rate across benchmarks, harnesses, and policy models.
Figure 2 : Overview of Opera . ❶ Reviews are triggered at fixed intervals or by execution events. ❷ The operator critic diagnoses one issue with a typed operator, and an audit decides whether the feedback is delivered. ❸ Adherence and resolution are tracked separately; resolved notes are closed, unresolved notes are refined, and note history conditions the next review.
Figure 3 : Core mechanisms of Opera . (a) Nine operators, grouped by repair stage and specified by contracts, map an observed issue to a targeted fix. (b) Admission audits gate note creation and updates, rejecting unsupported or redundant feedback; the release audit closes a note only when its fixed criterion is met.
Method
TB-2.1
SWE-Pro
DeepSWE
Non-critic
65.9 ± 1.7
75.3 ± 1.2
32.4 ± 1.0
SWE-PRM
68.9 ± 3.2
77.7 ± 2.5
36.6 ± 3.1
SWE-Search
70.0 ± 4.7
77.3 ± 3.8
38.6 ± 4.4
LLM-as-verifier
71.2 ± 2.6
78.7 ± 2.3
38.9 ± 2.3
Agentic Rubrics
71.9 ± 4.5
78.3 ± 3.2
39.5 ± 2.8
Opera (ours)
73.8 ± 2.3
79.3 ± 3.5
41.3 ± 3.1
Table 1 : Comparison with critic baselines (policy: Qwen3.8-27B). Resolve rate (%, mean ± std) on Terminal-Bench 2.1 (TB-2.1, 89 tasks), a 100-task subset of SWE-Bench Pro (SWE-Pro), and DeepSWE v1.1 (113 tasks). Green bold / blue underline: best / second-best.
Figure 4 : Task resolve rate of Opera across policy models (critic: GPT-5.6-Sol). Gains (pp) are generally larger for weaker backbones and on benchmarks with more headroom. Qwen3.5-9B resolves no DeepSWE task with or without the critic. Error bars denote sample standard deviation.
Table 2: Task distribution of the SFT training split and the held-out OOD evaluation split of SWE-Bench Pro. TS/JS: TypeScript/JavaScript.
Training data
SWE-Pro
TB2.1
None (base)
28.7±1.9
23.6±1.9
Qwen3.8-27B rollouts
39.5±1.4
7.9±1.9
Opera -guided (ours)
38.9±2.4
25.8±2.2
Table 3: Resolve rate (%) of Qwen3.5-9B trained on different trajectory sources, evaluated without a critic on held-out SWE-Bench Pro tasks and Terminal-Bench 2.1. Mean ± std over three runs.
Figure 5 : Distribution of admitted operators. Edit-stage operators dominate while the leading one shifts from requirement contract to state transition as model capability increases.
Figure 6 : Task outcomes on trials receiving notes. (a) Observed Opera success versus expected no-critic success on the same tasks. (b) Rescue rate among tasks never solved without the critic and regression rate among tasks always solved without it; the two rates use different denominators.
Variant
TB2.1
SWE-Pro
DeepSWE
Non-critic
65.9±1.7
75.3±1.2
32.4±1.0
Opera (full)
73.8±2.3 (+7.9)
79.3±3.5 (+4.0)
41.3±3.1 (+8.9)
Edit-only
71.9±1.6 (+6.0)
79.0±3.5 (+3.7)
40.1±1.8 (+7.7)
Non-edit
70.2±0.8 (+4.3)
78.3±0.6 (+3.0)
39.5±2.7 (+7.1)
Table 4: Operator ablation with Qwen3.8-27B. Resolve rate (%), mean ± std over three runs; parentheses show gains over non-critic (pp).
Figure 7 : Task resolve rate with different critic models: the policy model itself (Self), Claude, and GPT. Opera has consistent improvement regardless of the critic. Self-critique is effective where the policy is competent, while stronger critics bring large gains where the policy struggles.
Variant
TB2.1
SWE-Pro
DeepSWE
Non-critic
65.9±1.7
75.3±1.2
32.4±1.0
Opera (w/ audit)
73.8±2.3 (+7.9)
79.3±3.5 (+4.0)
41.3±3.1 (+8.9)
Opera (w/o audit)
68.5±2.2 (+2.6)
78.3±2.1 (+3.0)
36.9±1.8 (+4.5)
Table 5 : Audit ablation with Qwen3.8-27B as the policy and GPT-5.6-Sol as the critic. Resolve rate (%), mean ± std over three runs; parentheses show gains over the non-critic agent (pp). Bold: best mean per benchmark.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Shared settings
Student model
Qwen3.5-9B, full-parameter, bf16
Framework
ms-swift, DeepSpeed ZeRO-3
Maximum sequence length
128k tokens
Loss
Target turn only
Optimizer
AdamW, β2=0.95 , wd 0.1, clip 1.0
Learning rate
2×10−6 , cosine, 3% warm-up
Appendix
Table 6: Training configuration. Shared settings apply to both trajectory sources; per-source statistics are listed at the bottom.
Figure 8 : Training dynamics in SFT using Opera guided self-reflection trajectories and Qwen3.8-27B trajectories.
Figure 9 : Agent responses to common operators. Bars show the percentages of delivered notes followed by an edit or a test within three turns, and the percentage eventually closed under audit.
Figure 10 : Task-success gains relative to baseline expectations, grouped by the first admitted operator. Common operator groups are shown. Gains reflect subsequent execution as a whole, not the isolated effect of the first operator.
Figure 11 : Task resolve-rate gains (pp) relative to no-critic expectations, grouped by policy model, first admitted operator, and note-closure status. Differences between open-note and closed-note cohorts describe associations rather than the causal effect of closure.
Figure 12 : Intervention candidate outcomes across critic models. Left: delivered, audit-rejected, pre-audit-filtered, and unclassified candidates. Right: rejection rates for individual pre-audit gates. All percentages use all intervention candidates for the corresponding critic as the denominator.
Figure 13 : Proposal and delivery rates across critic and policy models. Each point represents a policy–critic setting within a benchmark: candidates per review on the x-axis and delivered notes per review on the y-axis. Colors denote critics, shapes denote policies, and lines connect the same policy across critics. Delivery rates are estimated from rounded aggregate statistics.