Agentic Benchmarks

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

13 new papers

A weekly snapshot of new work published in Agentic Benchmarks.

Period ending 2026-09-14

13 new papers

A weekly snapshot of new work published in Agentic Benchmarks.

Period ending 2026-09-07

17 new papers

A weekly snapshot of new work published in Agentic Benchmarks.

Inside this field

Focused directions

414 papers

Latest in Agentic Benchmarks

  1. Can Generalist Agents Automate Data Curation?

    Jun 2, 2026Feiyang Kang, Hanze Li, Adam Nguyen +5Data-CurationAgentic Benchmarks

  2. Monitoring Agentic Systems Before They're Reliable

    Jun 1, 2026Marisa Ferrara Boston, Glen Hanson, Effi Georgala +2Agentic SystemsAgentic Evaluations

  3. BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

    May 29, 2026Srivatsa Kundurthy, Clara Na, Colton Moraine +6SpreadsheetsFinance Domain

  4. EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations

    May 28, 2026Sneheel Sarangi, Maximilian Puelma Touzel, Aurélien Bück-Kaeffer +3User SimulationMulti-Agent Simulations

  5. Realistic honeypot evaluations for scheming propensity

    May 28, 2026Victoria Krakovna, David Lindner, Lewis Ho +2Agentic EvaluationsHoneypots

  6. GTA: Generating Long-Horizon Tasks for Web Agents at Scale

    May 28, 2026Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey +4Web AgentsAgentic Benchmarks