cs.AIOct 4, 2026

EPOCH: Reliable Discovery through Evidence-Governed Search

Authors: Binjie Guo, Aisheng Mo, Ruitong Li, Xinle Deng

Abstract

AI research agents are increasingly used to search over programs, mathematical constructions, and proofs. However, existing systems typically optimize evaluator feedback without adequately governing how that feedback is interpreted, challenged, and reused. As a result, promising but fragile candidates can be promoted as discoveries, while benchmark improvements, finite certificates, and theorem-level claims are too easily conflated. We introduce EPOCH, an evidence-governed architecture designed to close this gap. EPOCH implements an evidence-governed discovery loop by combining explicit task contracts, typed memory, active falsification, admission checks, and independent replay, so that each candidate is evaluated against the strength and scope of the claim it supports. EPOCH achieves state-of-the-art aggregate performance on AlgoTune, substantially exceeding the strongest baseline in mean normalized score (0.65 vs. 0.53), and attains the highest mean score on the internal Math14 suite (0.57). It further shows favorable held-out behavior under official-test replay and leads the descriptive aggregate on AgentHPO. Across ten discovery problems, EPOCH delivers substantial task-specific advances, including improved executable constructions, optimized algorithms, counterexamples, and proof-supported results. These advances demonstrate its ability to convert search into concrete progress across mathematical and computational domains. Together, the results suggest that evidence governance is a necessary step toward AI research agents that produce not only stronger solutions, but also more trustworthy scientific discoveries.

Explore similar work

Sep 7, 2026cs.MA

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

AI research agents combine public information and experimental feedback to produce measurable results. The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget. Gate 1 validates useful improvement. Gate 2 tests recovery by matched agents given the starting information and observed Web content, with run history and new measurements withheld. Core requires adequate registered controls, zero recoveries, and a finite-sample recovery bound. Optional Gate 3 compares truthful and neutral feedback from a shared checkpoint; Evidence adds a supported effect and a null-policy equivalence check. Controlled SQLite and virtual catalyst audits pass both decision kernels. On real-data response surfaces, Yacht and Ionosphere pass the Core kernel after zero recoveries in 96 attempts, with an upper bound of 0.0468. Each target combines ten observed utilities and six predictions into a 16-entry data product. Yacht scores 0.7677 on reconstruction of all 32 switch effects, with utility-prediction MAE 0.0315 on its six unmeasured configurations. Fresh truthful continuations recover the target level in 9/30 and 16/30 trials, respectively, separating achieved utility from process repeatability. A deterministic verifier reproduces these local decisions from frozen records.
Sep 27, 2026cs.AI

DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hypothesis generation separately rather than testing whether reliable evidence supports valid and novel claims. We introduce DISCERN (Data Integrity and Scientific Capability: Evidence, Reasoning, and Novelty), a controlled benchmark on real, publicly available datasets that evaluates three key levels of an automated research workflow. The first two levels test data integrity and analysis verification under confounds and tool traps, while the third tests hypothesis generation and revision under adversarial review, including counterfactual cases in which evidence consistent with real data and documented scientific phenomena conflicts with established expectations, motivating alternative explanations and testable hypotheses. Across 203 tasks, eight life-science tracks, and eight models, DISCERN shows that strong aggregate performance can mask level-specific weaknesses. Agents earn perfect scores in only 60.8% of Level 1, 34.2% of Level 2, and 0.6% of Level 3 evaluations, with penalties attributed to rejection of sound data, failure to carry recognized limitations into conclusions, and wide variation in hypothesis production. Cross-track rankings by token and code use are substantially more stable than rankings by evidence judgment, suggesting greater consistency in computational effort than in evidence-based reasoning. These profiles identify opportunities for supervised scientific assistance, but current agents do not yet demonstrate reliable autonomous analysis or discovery. Code and data: https://huggingface.co/datasets/discern-bench-anon/discern-benchmark
Aug 24, 2026cs.AI

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate and publish papers. These papers accumulate into a shared body of knowledge that later agents can read, cite and extend. We evaluated the Station on 12 mathematical construction problems from the AlphaEvolve study and two additional case studies. Five of the 12 problems yielded results novel relative to the prior literature: a new infinite family of finite field Kakeya sets, new exact 604-point kissing configurations in eleven dimensions, improved bounds for the discretized Kakeya needle and sign uncertainty problems, and a substantially improved lower bound for Erdős's minimum overlap problem. Agents also discovered novel infinite families for Book Ramsey numbers. Their research extended beyond searching for high-scoring constructions: agents developed explanations of their findings and proved theorems outside the assigned tasks. These explanations guided further discoveries and were preserved in the agents' papers, making the underlying insights easier for external researchers to understand and build upon. All presented discoveries are supported by exact constructions or proofs formally verified in Lean. We release the source code, full agent dialogues, papers and verification code, providing a transparent record of how these discoveries emerged.