Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
Organizations: Penn State University · Meta
Abstract
As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.
Figures & tables
| Model | Training | Acc | Dir-3 | F1 | 5-Ladder |
|---|---|---|---|---|---|
| Gemini-3.1-Pro | Zero-shot | 54.3 | 46.6 | 59.8 | 28.2 |
| Qwen-3.5-9B | Base | 49.9 | 41.3 | 58.3 | 22.4 |
| Qwen-3.5-35B | Base | 49.9 | 36.4 | 51.8 | 20.1 |
| Deepseek-GRM-16B | Ckpt from Liu et al. (2025) | 58.5 | 57.8 | 73.3 | 25.3 |
| Deepseek-GRM-27B | Ckpt from Liu et al. (2025) | 58.4 | 58.1 | 73.6 | 25.3 |
| Qwen-3.5-9B | SFT | 62.4 | 55.6 | 70.9 | 31.5 |
| Domain | None | Gemini | Qwen-3.5-9B | Qwen-3.5-35B | |||
| 3.1-Pro | 3.5-Flash | Base | Critic | Base | Critic | ||
| Language Models | 18.04 | 18.85 | 17.66 | 32.48 | 32.57 | 32.54 | 31.47 |
| Robotics | 23.18 | 29.14 | 28.03 | 25.52 | 29.75 | 26.71 | 28.29 |
| Vision and Generation | 16.11 | 15.04 | 15.08 | 15.48 | 17.38 | 14.86 | 17.84 |
| ML Systems and Efficient ML | 64.00 | 64.01 | 63.73 | 63.11 | 65.54 | 63.59 | 63.80 |
| AI for Science | 33.55 | 34.52 | 32.38 | 33.94 | 31.67 | 32.68 | 33.09 |
| Task | Frozen Critic | TT-GRPO | TT-SDFT | |||
|---|---|---|---|---|---|---|
| 9B | 35B | 9B | 35B | 9B | 35B | |
| ML Dimensionality Reduction | 49.62 | 38.43 | 50.06 | 42.15 | 50.43 | 51.77 |
| TS Exogenous Forecast | 48.43 | 49.57 | 48.90 | 47.02 | 49.11 | 51.17 |
| AI4Bio Mutation Effect Prediction | 48.01 | 49.88 | 47.34 | 49.47 | 51.66 | 50.34 |
| Graph Generation | 70.69 | 73.30 | 72.58 | 73.05 | 73.51 | 76.64 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Task and parent program |
|---|
| Research Question |
| Design an activation function for deep convolutional neural networks that improves test accuracy across different architectures (ResNet, VGG) and datasets (CIFAR-10, CIFAR-100, FashionMNIST), while keeping the model definitions, optimizer, initialization, and data pipeline fixed. |
| Background |
| Activation functions introduce nonlinearity into neural networks and critically affect training dynamics, gradient flow, sparsity, and generalization. Classic and modern choices include: |
| - ReLU : max (0, x ) — simple, sparse, but zero gradient for negative inputs . |
| - GELU : x * Phi ( x ) where Phi is the standard Gaussian CDF. |
| Stage 1: idea proposal |
|---|
| System prompt |
| You are an expert ML researcher proposing ONE concrete improvement to an existing solution. You propose the IDEA only — you do NOT write code in this step. Be specific and technical: name the modeling/training/feature change and why it should move the target metric. Propose a single coherent idea, not a menu of options. |
| User prompt |
| Task: mls__dl-activation-function |
| Shared input: insert the complete task description from Table 4 . |
| Metric: combined_score (higher is better) |
| Stage 2: implementation |
|---|
| System prompt |
| You are an expert ML researcher. Implement the committed idea by rewriting ONLY the editable band of the file. Emit the replacement band in a single “ python “ block and nothing else after it. |
| User prompt |
| Committed idea (implement EXACTLY this) |
| You have already decided on the following idea. Implement EXACTLY this idea. Do NOT change the approach, do NOT substitute a different method, do NOT add unrelated changes — just turn this idea into correct, runnable code. |
| Shared input: insert the complete proposed approach from Table 5 . |
| Stage 3: solution summarization |
|---|
| System prompt |
| You are an expert ML solver. You are given a task, a candidate solution’s code, and its short approach. Write a NEUTRAL, purely descriptive summary of what the solution does. Do NOT evaluate, rank, judge, compare, or predict scores. |
| User prompt |
| Task: mls__dl-activation-function |
| Shared input: insert the complete task description from Table 4 . |
| Metric: combined_score (higher is better) |
| Task-specific context for the critic model |
|---|
| Task understanding |
| Modality: Computer Vision (Image Classification) |
| Metric: Test accuracy (higher-is-better) |
| Overview |
| The task requires implementing a custom activation function as a PyTorch nn . Module to serve as a drop-in replacement for standard activations in ResNet, VGG, and MobileNetV2 architectures. The submission format is Python code modifying the forward method and optionally the init method of the CustomActivation class within an 18-line editable region in pytorch - vision / custom_activation . py . The implemented function must perform an element-wise or channel-wise operation that preserves the exact shape of the input tensor. |
| Data |
| Component | Prompt content |
|---|---|
| Role | You are an expert ML solution critic. Given a parent solution and a candidate child idea, predict whether implementing the child idea will improve the task metric over the parent. |
| Decision stages | Stage 1: Improvement. Determine whether the child will improve , lose , or tie relative to the parent. Stage 2: Win tier. If the Stage-1 prediction is improve , determine whether the expected gain is big , medium , or small . Otherwise, output NA . |
| Task input | {task_description} |
| # Parent (baseline) | {parent_idea_summary} |
| # Child (candidate) | {child_idea_summary} |
| Task context | {neutral_understanding_md} |
| Freestyle Reasoning | Soft Rubric-Based Reasoning | |
|---|---|---|
| Analysis | The child proposes a custom U-Net trained from scratch on a subset of the available data, with images downsampled to . Although the pipeline is computationally feasible, training a CNN from scratch on a small medical-imaging dataset without pretrained weights or data augmentation is likely to cause severe overfitting. The parent instead relies on statistical priors, such as mean masks and spatial probabilities. These priors are robust on small datasets and can exploit the relatively consistent anatomical locations of the target organs. The freestyle critic therefore predicts that the child will generalize worse than the parent. | 1. Criteria considered. MRI handling and feature extraction: whether the method captures spatial structure, intensity patterns, and patient-specific anatomical variation. Dice and Hausdorff performance: whether the method can produce accurate masks and boundaries under the task metric. … 2. Criterion importance. Feature extraction receives high importance ( ) because medical segmentation depends on learning spatial and intensity features. Metric alignment also receives high importance ( ), particularly because Hausdorff distance is sensitive to boundary errors and outlier pixels. … 3. Parent–child comparison. For feature extraction, the child is substantially better: the parent’s statistical templates cannot capture patient-specific variation, whereas U-Net can learn nonlinear hierarchical features directly from the images. For metric alignment, the child is also better: U-Net’s skip connections can produce more accurate and continuous segmentation boundaries than fixed statistical priors. …. 4. Aggregation. After weighting the criterion-level comparisons, the expected benefits of learned spatial features and improved boundaries outweigh the child’s increased risk of overfitting. |
| Verdict | [IMPROVEMENT] lose The child is likely to overfit because it trains a U-Net from scratch on only a subset of a small dataset. [WINTIER] NA The predicted outcome is not an improvement. | [IMPROVEMENT] improve The child replaces a rigid statistical prior with a model capable of learning task-specific spatial and intensity features. [WINTIER] big Moving from a non-learned statistical baseline to U-Net is expected to produce a substantial improvement in both mask accuracy and boundary quality. |
| Model | Training | Acc | Dir-3 | F1 | 5-Ladder |
|---|---|---|---|---|---|
| Qwen-3.5-9B | Base | 49.9 | 41.3 | 58.3 | 22.4 |
| Qwen-3.5-35B | Base | 49.9 | 36.4 | 51.8 | 20.1 |
| Qwen-3.5-9B | SFT, freestyle | 61.0 | 54.6 | 71.5 | 31.9 |
| Qwen-3.5-9B | SFT, soft-rubric | 62.4 | 55.6 | 70.9 | 31.5 |
| Qwen-3.5-9B | SFT+GRPO, freestyle | 49.1 | 40.3 | 48.6 | 26.8 |
| Qwen-3.5-9B | SFT+GRPO, soft-rubric | 66.7 | 60.7 | 75.6 | 41.5 |
| Model | Training | Improve Acc. | F1 | Direction-3 | 5-Ladder | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sampled | Revised | Guided | Sampled | Revised | Guided | Sampled | Revised | Guided | Sampled | Revised | Guided | ||
| Qwen-3.5-9B | Base | ||||||||||||
| Qwen-3.5-35B | Base | ||||||||||||
| Gemini-3.1-Pro | Zero-shot | 72.4 | 60.6 | ||||||||||
| Qwen-3.5-9B | SFT | ||||||||||||
| Qwen-3.5-9B | SFT+GRPO | 72.7 | 63.6 | ||||||||||
| Domain | # | Tasks |
|---|---|---|
| Language Models | 3 | llm-dllm-demask-strategy ; llm-pretrain-optimizer ; llm-rl-importance-sampling |
| Robotics | 5 | jepa-planning ; robo-diffusion-guidance ; robo-diffusion-policy ; robo-humanoid-sim2real-algo ; robomimic-bc-loss |
| Vision and Generation | 3 | cv-3dgs-densification ; cv-dbm-sampler ; cv-vae-loss |
| Reinforcement Learning | 1 | rl-value-discrete |
| ML Systems and Efficient ML | 3 | llm-ptq-algorithm ; llm-qat-algorithm ; mlsys-sparse-attention-inference |
| AI for Science | 3 | ai4bio-mutation-effect-prediction ; ai4sci-inverse-diffusion-algo ; ai4sci-pla-binding-affinity |
| Task: LLM post-training quantization (INT4; WikiText-2 perplexity, lower is better). | |
| Parent idea: Apply GPTQ with Hessian-based error compensation, dynamically computed group quantization parameters, and damped Cholesky decomposition. | |
| Child idea: Integrate activation-aware channel scaling into GPTQ, use asymmetric group quantization, and add adaptive dampening while retaining GPTQ error compensation. | |
| Empirical outcome: Perplexity decreases from to ( relative reduction); label: Improve . | |
| Student: without outcome hint | Teacher: with outcome hint |
| Trained critic with the SDFT LoRA adapter enabled. | Frozen critic; receives the measured parent/child metrics and outcome label. |
| Quantization reasoning | Quantization reasoning |
| Domain | Qwen-3.5-9B | Qwen-3.5-35B | ||||
|---|---|---|---|---|---|---|
| Base | Verify-all | Critic-in-loop | Base | Verify-all | Critic-in-loop | |
| Language Models | 9.24 | 12.17 | 12.04 | 11.94 | 12.90 | 12.62 |
| Robotics | 7.91 | 18.83 | 15.87 | 16.83 | 21.20 | 23.46 |
| Vision and Generation | 11.66 | 6.61 | 13.23 | 10.83 | 11.03 | 14.07 |
| ML Systems and Efficient ML | 21.88 | 23.99 | 36.16 | 11.39 | 16.87 | 30.99 |
| AI for Science | 12.49 | 14.17 | 13.86 | 11.09 | 8.34 | 12.61 |