cs.AIOct 5, 2026

HERA: Harness-Environment Co-Evolution for Reliable Agentic Abstention

Authors: Han Luo, Bingbing Wen, Guang Yang, Zora Zhiruo Wang, Pan Lu, Lucy Lu Wang

Organizations: University of Washington · University of Leeds · Carnegie Mellon University · Stanford University · Allen Institute for AI

Abstract

Large language model (LLM) agents are increasingly capable of acting in complex tool-use environments, yet they often fail to recognize when tasks are infeasible and no valid solution exists. Recent work has formalized this reliability gap as the problem of agentic abstention, and existing approaches typically optimize a model or agent harness against a fixed set of tasks, leading to limited generalization to unseen failure modes. We introduce HERA, a framework for harness-environment co-evolution for agentic abstention. HERA consists of (i) a pipeline to automatically construct verifiable pairs of feasible and infeasible tasks by applying controlled environment mutations that transform solvable tasks into cases requiring abstention, and (ii) a co-evolution procedure in which performance failures on previous tasks are used to drive harness adaptation and generate new execution environments and tasks geared towards previous weaknesses. On held-out evaluation tasks, an evolved harness from HERA improves abstention accuracy from 61.7% to 83.3% while improving feasible-task completion from 68.3% to 76.7%, achieving the highest abstention and feasible-task completion among the compared methods. The resulting best harness transfers across 19 other LLMs, improving abstention accuracy by 15.3 percentage points on average without any model-specific optimization, and enabling smaller models to match the performance of more powerful models at an estimated 85% lower cost.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AgentAbstain: Do LLM Agents Know When Not to Act?

    Jul 11, 2026Xun Liu, Yi Evie Zhang, Vira Kasprova +5Large Language Model AgentsLanguage-Model Agents

  2. Agentic Abstention: Do Agents Know When to Stop Instead of Act?

    Jun 27, 2026Han Luo, Bingbing Wen, Lucy Lu WangLanguage-Model AgentsAbstention

  3. Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

    Sep 22, 2026Laizhen Li, Jiarui Li, Juanjuan Zhao +4Agent HarnessScaffolds