cs.AIOct 4, 2026

EPOCH: Reliable Discovery through Evidence-Governed Search

Authors: Binjie Guo, Aisheng Mo, Ruitong Li, Xinle Deng

Abstract

AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.

Explore similar work

CardsList
  1. Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

    Sep 7, 2026Jingjie Ning, Shanshan Zhong, Xiaochuan Li +1AI Agents for Scientific DiscoveryAlgorithmic Auditing

  2. DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

    Sep 27, 2026Nan Huang, Mario Tapia-Pacheco, Kun Zhou +4AI Agent ReliabilityScientific Hypothesis Generation

  3. Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

    Aug 24, 2026Stephen Chung, Wenyu Du, William J. WesleyMulti-Agent CollaborationMulti-Agent Systems