cs.LGMar 9, 2026

$OneMillion-Bench: How Far are Language Agents from Human Experts?

Authors: Yang Liu, Jiaqi Li, Jun Bai, Qianyu Yang, Xiaobo Hu, Tao Peng, Zaiyuan Wang, Ran Tian, +15 more

Organizations: BIGAI · Humanlaya · xbench · M-A-P

Abstract

As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce $OneMillion-Bench ($OMB), a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focusing on expert-level problems to ensure meaningful differentiation across agents. Together, $OMB provides a unified testbed for assessing agentic reliability, professional depth, and an indicator of practical readiness in domain-intensive scenarios.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

    Jun 9, 2026Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12Customer Support AutomationLLM Agent Evaluation

  2. OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

    Jul 16, 2026Chengyu Shen, Yujie Fu, Gangtao Xin +13Computer-Use Agent BenchmarksAI Agent Evaluation

  3. WorldBench: Culturally Grounded Benchmark for Multilingual Agents

    Sep 1, 2026Leonardo Ranaldi, Sherrie Shen, Jushi Kai +1Cross-Cultural Language Model EvaluationLLM Agent Evaluation