Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is delivered. We present Opera, a verbal critic framework that treats each correction as a persistent note, followed until the diagnosed problem is resolved. Opera decides when to review through periodic and event-driven triggers, diagnoses issues with typed operators, audits feedback against visible evidence before delivery, and tracks the agent's subsequent actions to distinguish mere compliance from actual resolution. As a test-time critic, Opera improves the resolve rate of non-critic agents by up to 12.4, 15.0, and 8.9 percentage points on Terminal-Bench 2.1, a SWE-Bench Pro subset, and DeepSWE v1.1, respectively, across four policy models, and achieves the highest mean resolve rate among competitive critic baselines on all three benchmarks, and also improves policy models when the policy critiques itself. Beyond inference, Opera-guided rollouts provide approximately on-policy training data: fine-tuning Qwen3.5-9B on them improves its resolve rate on held-out SWE-Bench Pro repositories by 10.2 percentage points without a critic at inference time, matching fine-tuning on rollouts from a stronger model, while preserving its performance when switching harness, i.e., from Openhands to Terminus-2, which the latter substantially degrades.
Figures & tables
Figure 1 : Illustration of Opera . Left: Opera discovers and tracks issues at different stages (search, edit, test, etc.) until it is resolved or refined in the long-horizon turns. Right: Opera improves task resolve rate across benchmarks, harnesses, and policy models.
Figure 2 : Overview of Opera . ❶ Reviews are triggered at fixed intervals or by execution events. ❷ The operator critic diagnoses one issue with a typed operator, and an audit decides whether the feedback is delivered. ❸ Adherence and resolution are tracked separately; resolved notes are closed, unresolved notes are refined, and note history conditions the next review.
Figure 3 : Core mechanisms of Opera . (a) Nine operators, grouped by repair stage and specified by contracts, map an observed issue to a targeted fix. (b) Admission audits gate note creation and updates, rejecting unsupported or redundant feedback; the release audit closes a note only when its fixed criterion is met.
Method
TB-2.1
SWE-Pro
DeepSWE
Non-critic
65.9 ± 1.7
75.3 ± 1.2
32.4 ± 1.0
SWE-PRM
68.9 ± 3.2
77.7 ± 2.5
36.6 ± 3.1
SWE-Search
70.0 ± 4.7
77.3 ± 3.8
38.6 ± 4.4
LLM-as-verifier
71.2 ± 2.6
78.7 ± 2.3
38.9 ± 2.3
Agentic Rubrics
71.9 ± 4.5
78.3 ± 3.2
39.5 ± 2.8
Opera (ours)
73.8 ± 2.3
79.3 ± 3.5
41.3 ± 3.1
Table 1 : Comparison with critic baselines (policy: Qwen3.8-27B). Resolve rate (%, mean ± std) on Terminal-Bench 2.1 (TB-2.1, 89 tasks), a 100-task subset of SWE-Bench Pro (SWE-Pro), and DeepSWE v1.1 (113 tasks). Green bold / blue underline: best / second-best.
Figure 4 : Task resolve rate of Opera across policy models (critic: GPT-5.6-Sol). Gains (pp) are generally larger for weaker backbones and on benchmarks with more headroom. Qwen3.5-9B resolves no DeepSWE task with or without the critic. Error bars denote sample standard deviation.
Table 2: Task distribution of the SFT training split and the held-out OOD evaluation split of SWE-Bench Pro. TS/JS: TypeScript/JavaScript.
Training data
SWE-Pro
TB2.1
None (base)
28.7±1.9
23.6±1.9
Qwen3.8-27B rollouts
39.5±1.4
7.9±1.9
Opera -guided (ours)
38.9±2.4
25.8±2.2
Table 3: Resolve rate (%) of Qwen3.5-9B trained on different trajectory sources, evaluated without a critic on held-out SWE-Bench Pro tasks and Terminal-Bench 2.1. Mean ± std over three runs.
Figure 5 : Distribution of admitted operators. Edit-stage operators dominate while the leading one shifts from requirement contract to state transition as model capability increases.
Figure 6 : Task outcomes on trials receiving notes. (a) Observed Opera success versus expected no-critic success on the same tasks. (b) Rescue rate among tasks never solved without the critic and regression rate among tasks always solved without it; the two rates use different denominators.
Variant
TB2.1
SWE-Pro
DeepSWE
Non-critic
65.9±1.7
75.3±1.2
32.4±1.0
Opera (full)
73.8±2.3 (+7.9)
79.3±3.5 (+4.0)
41.3±3.1 (+8.9)
Edit-only
71.9±1.6 (+6.0)
79.0±3.5 (+3.7)
40.1±1.8 (+7.7)
Non-edit
70.2±0.8 (+4.3)
78.3±0.6 (+3.0)
39.5±2.7 (+7.1)
Table 4: Operator ablation with Qwen3.8-27B. Resolve rate (%), mean ± std over three runs; parentheses show gains over non-critic (pp).
Figure 7 : Task resolve rate with different critic models: the policy model itself (Self), Claude, and GPT. Opera has consistent improvement regardless of the critic. Self-critique is effective where the policy is competent, while stronger critics bring large gains where the policy struggles.
Variant
TB2.1
SWE-Pro
DeepSWE
Non-critic
65.9±1.7
75.3±1.2
32.4±1.0
Opera (w/ audit)
73.8±2.3 (+7.9)
79.3±3.5 (+4.0)
41.3±3.1 (+8.9)
Opera (w/o audit)
68.5±2.2 (+2.6)
78.3±2.1 (+3.0)
36.9±1.8 (+4.5)
Table 5 : Audit ablation with Qwen3.8-27B as the policy and GPT-5.6-Sol as the critic. Resolve rate (%), mean ± std over three runs; parentheses show gains over the non-critic agent (pp). Bold: best mean per benchmark.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Shared settings
Student model
Qwen3.5-9B, full-parameter, bf16
Framework
ms-swift, DeepSpeed ZeRO-3
Maximum sequence length
128k tokens
Loss
Target turn only
Optimizer
AdamW, β2=0.95 , wd 0.1, clip 1.0
Learning rate
2×10−6 , cosine, 3% warm-up
Appendix
Table 6: Training configuration. Shared settings apply to both trajectory sources; per-source statistics are listed at the bottom.
Figure 8 : Training dynamics in SFT using Opera guided self-reflection trajectories and Qwen3.8-27B trajectories.
Figure 9 : Agent responses to common operators. Bars show the percentages of delivered notes followed by an edit or a test within three turns, and the percentage eventually closed under audit.
Figure 10 : Task-success gains relative to baseline expectations, grouped by the first admitted operator. Common operator groups are shown. Gains reflect subsequent execution as a whole, not the isolated effect of the first operator.
Figure 11 : Task resolve-rate gains (pp) relative to no-critic expectations, grouped by policy model, first admitted operator, and note-closure status. Differences between open-note and closed-note cohorts describe associations rather than the causal effect of closure.
Figure 12 : Intervention candidate outcomes across critic models. Left: delivered, audit-rejected, pre-audit-filtered, and unclassified candidates. Right: rejection rates for individual pre-audit gates. All percentages use all intervention candidates for the corresponding critic as the denominator.
Figure 13 : Proposal and delivery rates across critic and policy models. Each point represents a policy–critic setting within a benchmark: candidates per review on the x-axis and delivered notes per review on the y-axis. Colors denote critics, shapes denote policies, and lines connect the same policy across critics. Delivery rates are estimated from rounded aggregate statistics.
Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation. While coding agents are optimized for the joint capabilities, individual capabilities such as high-level planning may have different optima and remain a major bottleneck. To address this challenge, we train a separate critic model that is specialized in high-level planning to steer the coding agent in inference. We construct SFT and DPO data to train the critic model to identify errors made by the coding agent and provide correct and clear high-level guidance without generating concrete actions. Experiments show that our fine-tuned 4B and 8B critic models significantly improve the performance of 6 larger coding agents (e.g., improving the resolved rates of GLM-4.7-Flash-30B-A3B and GPT-OSS-120B by 16.0% and 14.4% on SWE-Bench Verified). The critic model also reduces the total inference costs for some coding agents by solving tasks in fewer steps (e.g., reducing the per-example inference cost for GPT-OSS-20B from $0.07 to $0.03). Code: https://github.com/shubhamrgandhi/critic-training
Execution feedback can guide coding agents toward correct repository repairs, but only when the tests capture the behavior requested by the issue. Agent-generated tests can encode incomplete or incorrect behavioral targets; when the same trajectory writes both the patch and the test, their errors can agree and create false confidence. We introduce ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it. The scaffold separates test construction from source-code repair: a Test agent independently generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from their execution feedback without changing the tests. Both roles use Qwen-3.5-35B-A3B as the backbone and are trained separately. In Learn to Test, the Test agent learns to produce behaviorally valid tests that distinguish correct from incorrect patches. In Test to Improve, the Repair agent learns both direct task resolution and feedback-guided revision. On SWE-bench Verified, test quality determines whether feedback helps: holding the base Repair agent fixed, tests from the base Test agent reduce resolved rate from a no-test baseline of 61.2% to 57.3%, whereas tests from GPT-5.6-sol raise it to 65.3%. Role-specific post-training raises the Qwen Test agent's Base-to-Gold success from 22.2% to 62.2%; composing the two post-trained Qwen agents reaches 72.6%, an 11.4-point gain over the original no-test baseline without stronger-model or Oracle feedback at evaluation time. Code is publicly available at https://github.com/MSR-Orchard/execcritic.
Leitian Tao, Baolin Peng, Haorui Wang +7
University of Wisconsin–Madison · Microsoft Research · Georgia Tech
Test-time scaling has become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that can be directly compared, ranked or refined. Long-horizon coding agents violate this premise: each attempt produces an extended trajectory of actions, observations, errors, and partial progress taken by the agent. In this setting, the main challenge is no longer generating more attempts, but representing prior experience in a form that can be effectively selected from and reused. We propose a test-time scaling framework for agentic coding based on compact representations of rollout trajectories. Our framework converts each rollout into a structured summary that preserves its salient hypotheses, progress, and failure modes while discarding low-signal trace details. This representation enables two complementary forms of inference-time scaling. For parallel scaling, we introduce Recursive Tournament Voting (RTV), which recursively narrows a population of rollout summaries through small-group comparisons. For sequential scaling, we adapt Parallel-Distill-Refine (PDR) to the agentic setting by conditioning new rollouts on summaries distilled from prior attempts. Our method consistently improves the performance of frontier coding agents across SWE-Bench Verified and Terminal-Bench v2.0. For example, by using our method Claude-4.5-Opus improves from 70.9% to 77.6% on SWE-Bench Verified (mini-SWE-agent) and 46.9% to 59.1% on Terminal-Bench v2.0 (Terminus 1). Our results suggest that test-time scaling for long-horizon agents is fundamentally a problem of representation, selection, and reuse.
Joongwon Kim, Wannan Yang, Kelvin Niu +13
1Meta Superintelligence Labs · University of Washington · 3New York University +4