cs.AIOct 6, 2026

Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents

Authors: Jiamu Bai, Lizhu Zhang, Xin Yu, Yanhong Wu, Zellux Wang, Serena Li, Weiwei Li, Zhuokai Zhao, +4 more

Organizations: Penn State University · Meta

Abstract

As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea selectors. We address this gap with specialized idea-level critic models that predict whether a proposed ML modification will improve upon the current solution, allowing agents to screen ideas and concentrate verification resources on the most promising candidates. We train the critic models through supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to further improve their predictive accuracy. Empirically, our critic models outperform Gemini-3.1-Pro in static idea evaluation, and these gains extend to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting more promising ideas, with further gains from continual learning. During policy training, they serve as learned reward models, reserving empirical verification for uncertain cases and enabling substantially more policy updates with the same verification resources. Together, these results show that idea-level critic models help ML agents discover better solutions and learn stronger proposal policies under limited verification budgets.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Training Advisors for LLM Agents from Task Outcomes

    Oct 7, 2026Sergei Polezhaev, Barys Liskavets, Ori Press +1Critic-Free Reinforcement Learning

  2. Steer, Don't Solve: Training Small Critic Models for Large Code Agents

    Jun 20, 2026Shubham Gandhi, Yiqing Xie, Atharva Naik +2Coding AgentsCritic

  3. VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    Oct 1, 2026Caiqi Zhang, Rujun Han, Zifeng Wang +4Long-Horizon AgentsLong-Horizon Task Planning