As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.
Figures & tables
Figure 1: We train an idea-level critic model from ML agent evolution trajectories through SFT on soft-rubric critiques followed by GRPO. The trained critic model then supports (1) inference-time evolution through idea selection and task-specific test-time alignment, and (2) policy training by serving as a learned reward model while routing uncertain cases to empirical verification.
Figure 2: Overview of our idea-level screening pipeline. Given an ML task, a parent solution, and a candidate idea, the critic model predicts whether the idea will improve upon the parent. Only ideas predicted to improve are empirically verified, reducing costly verification.
Model
Training
Acc
Dir-3
F1
5-Ladder
Gemini-3.1-Pro
Zero-shot
54.3
46.6
59.8
28.2
Qwen-3.5-9B
Base
49.9
41.3
58.3
22.4
Qwen-3.5-35B
Base
49.9
36.4
51.8
20.1
Deepseek-GRM-16B
Ckpt from Liu et al. (2025)
58.5
57.8
73.3
25.3
Deepseek-GRM-27B
Ckpt from Liu et al. (2025)
58.4
58.1
73.6
25.3
Qwen-3.5-9B
SFT
62.4
55.6
70.9
31.5
Table 1: Idea-critic performance on the 1.5K-example evaluation set. Split-breakdown (Sampled/revised/guided) results are in Table 12 . Best results are shown in bold.
Domain
None
Gemini
Qwen-3.5-9B
Qwen-3.5-35B
3.1-Pro
3.5-Flash
Base
Critic
Base
Critic
Language Models
18.04
18.85
17.66
32.48
32.57
32.54
31.47
Robotics
23.18
29.14
28.03
25.52
29.75
26.71
28.29
Vision and Generation
16.11
15.04
15.08
15.48
17.38
14.86
17.84
ML Systems and Efficient ML
64.00
64.01
63.73
63.11
65.54
63.59
63.80
AI for Science
33.55
34.52
32.38
33.94
31.67
32.68
33.09
Table 2: Openevolve with critic model integrated result with MLS-Lite-Bench with Gemini-3.1-Pro as idea proposer. All domains use 10 evolution rounds. The best result in each row is shown in bold .
Task
Frozen Critic
TT-GRPO
TT-SDFT
9B
35B
9B
35B
9B
35B
ML Dimensionality Reduction
49.62
38.43
50.06
42.15
50.43
51.77
TS Exogenous Forecast
48.43
49.57
48.90
47.02
49.11
51.17
AI4Bio Mutation Effect Prediction
48.01
49.88
47.34
49.47
51.66
50.34
Graph Generation
70.69
73.30
72.58
73.05
73.51
76.64
Table 3: Task-specific test-time continual learning result. TT stands for test time continual learning alignment. Scores higher than no continual learning baseline are shown in bold .
Figure 3: Trained idea-proposing policies performance on MLS-Lite-Bench. For Qwen-3.5-9B (left) and Qwen-3.5-35B (right), we compare the base policy, verify-all GRPO, and critic-model-in-the-loop GRPO across 11 domains. For each domain, scores are normalized against the best of the six arms.
Figure 4: Uncertainty analysis. Increasing the number of critic samples yields diminishing AUROC gains beyond k=16 (left). At k=16 , an uncertainty threshold of 0.25 recalls around 70% of incorrect predictions while routing about 30% of all ideas to empirical verification (middle and right).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Task and parent program
Research Question
Design an activation function for deep convolutional neural networks that improves test accuracy across different architectures (ResNet, VGG) and datasets (CIFAR-10, CIFAR-100, FashionMNIST), while keeping the model definitions, optimizer, initialization, and data pipeline fixed.
Background
Activation functions introduce nonlinearity into neural networks and critically affect training dynamics, gradient flow, sparsity, and generalization. Classic and modern choices include:
- ReLU : max (0, x ) — simple, sparse, but zero gradient for negative inputs .
- GELU : x * Phi ( x ) where Phi is the standard Gaussian CDF.
Appendix
Table 4: Running example: task context and parent program for task activation-function design.
Stage 1: idea proposal
System prompt
You are an expert ML researcher proposing ONE concrete improvement to an existing solution. You propose the IDEA only — you do NOT write code in this step. Be specific and technical: name the modeling/training/feature change and why it should move the target metric. Propose a single coherent idea, not a menu of options.
User prompt
Task: mls__dl-activation-function
Shared input: insert the complete task description from Table 4 .
Metric: combined_score (higher is better)
Appendix
Table 5: Idea proposal for the activation-function example. The system instruction, stage-specific user instruction, and recorded approach are shown. Repeated task and parent inputs are referenced by table number.
Stage 2: implementation
System prompt
You are an expert ML researcher. Implement the committed idea by rewriting ONLY the editable band of the file. Emit the replacement band in a single “ python “ block and nothing else after it.
User prompt
Committed idea (implement EXACTLY this)
You have already decided on the following idea. Implement EXACTLY this idea. Do NOT change the approach, do NOT substitute a different method, do NOT add unrelated changes — just turn this idea into correct, runnable code.
Shared input: insert the complete proposed approach from Table 5 .
Appendix
Table 6: Implementation of the committed Mish-ACON-C idea. The replacement code is produced before critic model screening and is subsequently reused for verification. Repeated inputs and the full fixed code context are referenced by table number.
Stage 3: solution summarization
System prompt
You are an expert ML solver. You are given a task, a candidate solution’s code, and its short approach. Write a NEUTRAL, purely descriptive summary of what the solution does. Do NOT evaluate, rank, judge, compare, or predict scores.
User prompt
Task: mls__dl-activation-function
Shared input: insert the complete task description from Table 4 .
Metric: combined_score (higher is better)
Appendix
Table 7: Code-grounded summarization of the implemented child. The summarizer receives the task, proposed approach, and implemented code; its recorded output forms the child description supplied to the critic model.
Task-specific context for the critic model
Task understanding
Modality: Computer Vision (Image Classification)
Metric: Test accuracy (higher-is-better)
Overview
The task requires implementing a custom activation function as a PyTorch nn . Module to serve as a drop-in replacement for standard activations in ResNet, VGG, and MobileNetV2 architectures. The submission format is Python code modifying the forward method and optionally the init method of the CustomActivation class within an 18-line editable region in pytorch - vision / custom_activation . py . The implemented function must perform an element-wise or channel-wise operation that preserves the exact shape of the input tensor.
Data
Appendix
Table 8: Precomputed task understanding supplied to the critic model for activation-function design. This context is generated once per task and reused across candidates; it is shown here in its recorded wording.
Component
Prompt content
Role
You are an expert ML solution critic. Given a parent solution and a candidate child idea, predict whether implementing the child idea will improve the task metric over the parent.
Decision stages
Stage 1: Improvement. Determine whether the child will improve , lose , or tie relative to the parent. Stage 2: Win tier. If the Stage-1 prediction is improve , determine whether the expected gain is big , medium , or small . Otherwise, output NA .
Task input
{task_description}
# Parent (baseline)
{parent_idea_summary}
# Child (candidate)
{child_idea_summary}
Task context
{neutral_understanding_md}
Appendix
Table 9: Prompt template for the idea-level critic. The provided task-specific criteria must all be considered, while the critic may freely determine their relative importance and aggregation.
Freestyle Reasoning
Soft Rubric-Based Reasoning
Analysis
The child proposes a custom U-Net trained from scratch on a subset of the available data, with images downsampled to 128×128 . Although the pipeline is computationally feasible, training a CNN from scratch on a small medical-imaging dataset without pretrained weights or data augmentation is likely to cause severe overfitting. The parent instead relies on statistical priors, such as mean masks and spatial probabilities. These priors are robust on small datasets and can exploit the relatively consistent anatomical locations of the target organs. The freestyle critic therefore predicts that the child will generalize worse than the parent.
1. Criteria considered. MRI handling and feature extraction: whether the method captures spatial structure, intensity patterns, and patient-specific anatomical variation. Dice and Hausdorff performance: whether the method can produce accurate masks and boundaries under the task metric. … 2. Criterion importance. Feature extraction receives high importance ( 40% ) because medical segmentation depends on learning spatial and intensity features. Metric alignment also receives high importance ( 30% ), particularly because Hausdorff distance is sensitive to boundary errors and outlier pixels. … 3. Parent–child comparison. For feature extraction, the child is substantially better: the parent’s statistical templates cannot capture patient-specific variation, whereas U-Net can learn nonlinear hierarchical features directly from the images. For metric alignment, the child is also better: U-Net’s skip connections can produce more accurate and continuous segmentation boundaries than fixed statistical priors. …. 4. Aggregation. After weighting the criterion-level comparisons, the expected benefits of learned spatial features and improved boundaries outweigh the child’s increased risk of overfitting.
Verdict
[IMPROVEMENT] lose The child is likely to overfit because it trains a U-Net from scratch on only a subset of a small dataset. [WINTIER] NA The predicted outcome is not an improvement.
[IMPROVEMENT] improve The child replaces a rigid statistical prior with a model capable of learning task-specific spatial and intensity features. [WINTIER] big Moving from a non-learned statistical baseline to U-Net is expected to produce a substantial improvement in both mask accuracy and boundary quality.
Appendix
Table 10: Qualitative comparison between freestyle and soft rubric-based critic reasoning. The example is abridged for readability. Freestyle reasoning focuses primarily on the risk of overfitting, whereas soft rubric reasoning explicitly compares the parent and child under task-specific criteria before aggregating the judgments.
Model
Training
Acc
Dir-3
F1
5-Ladder
Qwen-3.5-9B
Base
49.9
41.3
58.3
22.4
Qwen-3.5-35B
Base
49.9
36.4
51.8
20.1
Qwen-3.5-9B
SFT, freestyle
61.0
54.6
71.5
31.9
Qwen-3.5-9B
SFT, soft-rubric
62.4
55.6
70.9
31.5
Qwen-3.5-9B
SFT+GRPO, freestyle
49.1
40.3
48.6
26.8
Qwen-3.5-9B
SFT+GRPO, soft-rubric
66.7
60.7
75.6
41.5
Appendix
Table 11: Idea-critic performance on the 1.5K-example evaluation set. SFT substantially improves both model sizes, while criterion-guided GRPO provides consistent additional gains. Best results are shown in bold.
Model
Training
Improve Acc.
F1
Direction-3
5-Ladder
Sampled
Revised
Guided
Sampled
Revised
Guided
Sampled
Revised
Guided
Sampled
Revised
Guided
Qwen-3.5-9B
Base
54.2
50.6
45.0
54.2
59.2
54.5
46.8
42.6
34.6
34.4
20.0
12.0
Qwen-3.5-35B
Base
58.0
47.4
44.4
58.0
51.7
46.5
48.0
35.4
25.8
29.8
20.0
8.6
Gemini-3.1-Pro
Zero-shot
74.8
46.6
41.4
77.2
55.3
47.4
72.4
39.8
27.6
60.6
17.8
4.6
Qwen-3.5-9B
SFT
72.0
61.2
54.0
77.6
71.1
64.2
65.2
55.2
46.4
48.0
27.2
19.4
Qwen-3.5-9B
SFT+GRPO
72.4
65.4
61.4
77.5
74.2
72.7
69.0
63.6
52.8
56.0
37.8
28.0
Appendix
Table 12: Critic performance (%) across the Sampled, Revised, and Guided subsets. These subsets correspond to increasing levels of teacher assistance during critique collection. The best result in each column is shown in bold.
Table 13: The 30 tasks in MLS-Bench-Lite, grouped by their 12 official domains. Tasks marked with † were included in the task environments used for SFT data collection, while tasks marked with ‡ were introduced only during critic model GRPO training. Unmarked tasks were held out from critic model training. We retain all 30 tasks to preserve the complete MLS-Bench-Lite evaluation suite.
Task: LLM post-training quantization (INT4; WikiText-2 perplexity, lower is better).
Parent idea: Apply GPTQ with Hessian-based error compensation, dynamically computed group quantization parameters, and damped Cholesky decomposition.
Child idea: Integrate activation-aware channel scaling into GPTQ, use asymmetric group quantization, and add adaptive dampening while retaining GPTQ error compensation.
Empirical outcome: Perplexity decreases from 5.0428 to 4.9906 ( 1.04% relative reduction); label: Improve .
Student: without outcome hint
Teacher: with outcome hint
Trained critic with the SDFT LoRA adapter enabled.
Frozen critic; receives the measured parent/child metrics and outcome label.
Quantization reasoning
Quantization reasoning
Appendix
Table 14: Student and outcome-conditioned teacher critiques on an LLM PTQ example. Idea descriptions are summarized; critique excerpts are quoted verbatim, with emphasis added. The critiques were regenerated from an archived parent–child pair for qualitative inspection.
Domain
Qwen-3.5-9B
Qwen-3.5-35B
Base
Verify-all
Critic-in-loop
Base
Verify-all
Critic-in-loop
Language Models
9.24
12.17
12.04
11.94
12.90
12.62
Robotics
7.91
18.83
15.87
16.83
21.20
23.46
Vision and Generation
11.66
6.61
13.23
10.83
11.03
14.07
ML Systems and Efficient ML
21.88
23.99
36.16
11.39
16.87
30.99
AI for Science
12.49
14.17
13.86
11.09
8.34
12.61
Appendix
Table 15: Policy GRPO under a fixed budget of 10,000 empirical verifications. We compare the base policy, verify-all GRPO, and critic-model-in-the-loop GRPO for Qwen-3.5-9B and Qwen-3.5-35B. Scores are averaged within each of the 11 retained MLS-Lite-Bench domains, and the best result in each row is shown in bold .
Figure 5: Critic-model-only reward ablation. The left panels vary the number of policy-training steps, while the right panels vary the number of critic model rollouts k with training fixed to 40 steps.
Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including τ3 and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.
Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation. While coding agents are optimized for the joint capabilities, individual capabilities such as high-level planning may have different optima and remain a major bottleneck. To address this challenge, we train a separate critic model that is specialized in high-level planning to steer the coding agent in inference. We construct SFT and DPO data to train the critic model to identify errors made by the coding agent and provide correct and clear high-level guidance without generating concrete actions. Experiments show that our fine-tuned 4B and 8B critic models significantly improve the performance of 6 larger coding agents (e.g., improving the resolved rates of GLM-4.7-Flash-30B-A3B and GPT-OSS-120B by 16.0% and 14.4% on SWE-Bench Verified). The critic model also reduces the total inference costs for some coding agents by solving tasks in fewer steps (e.g., reducing the per-example inference cost for GPT-OSS-20B from $0.07 to $0.03). Code: https://github.com/shubhamrgandhi/critic-training
As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Caiqi Zhang, Rujun Han, Zifeng Wang +4
University of Cambridge · Google Cloud AI Research