cs.SESep 24, 2026

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Authors: Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary

Organizations: Eon

Abstract

In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage. For each generated company, code fills each template and computes an exact answer without a language model. We evaluate 12 agents. Each pairs a model with an agent program, which connects it to the company's systems. The best agent answers 18 of its 24 attempts, three per question, correctly. Four of the six models answer at most 6 of 24 with any program. The hardest questions require picking one of several similar records, such as which of three renewal offers a customer signed. All agents together answered two such questions correctly in only 1 of 84 attempts.

Figures & tables

Explore similar work

CardsList
  1. EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions

    Jun 22, 2026Jincheng Zhong, Weizhi Wang, Che Jiang +5SessionsReproducibility

  2. WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

    Sep 14, 2026Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy +1Enterprise SystemsData Science Agents