Agentic Evaluations

Recent momentum

+21%

23 papers in the last 28 days · 0.4% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

8 new papers

A weekly snapshot of new work published in Agentic Evaluations.

Period ending 2026-09-14

5 new papers

A weekly snapshot of new work published in Agentic Evaluations.

Period ending 2026-09-07

4 new papers

A weekly snapshot of new work published in Agentic Evaluations.

148 papers

Latest in Agentic Evaluations

  1. Second MOASEI Competition at AAMAS'2026: A Technical Report

    Jul 3, 2026Ceferino Patino, Tyler J. Billings, Alireza Saleh Abadi +4Multi-Agent SystemsAgentic Evaluations

  2. PACE: A Proxy for Agentic Capability Evaluation

    Jul 2, 2026Yueqi Song, Lintang Sutawika, Jiarui Liu +8Agentic BenchmarksAgentic Evaluations

  3. A Framework for Evaluating Agentic Skills at Scale

    Jun 16, 2026Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich +3Agentic EvaluationsLarge Language Model Agents

  4. Monitoring Agentic Systems Before They're Reliable

    Jun 1, 2026Marisa Ferrara Boston, Glen Hanson, Effi Georgala +2Agentic SystemsAgentic Evaluations

  5. BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

    May 29, 2026Srivatsa Kundurthy, Clara Na, Colton Moraine +6SpreadsheetsFinance Domain

  6. EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations

    May 28, 2026Sneheel Sarangi, Maximilian Puelma Touzel, Aurélien Bück-Kaeffer +3User SimulationMulti-Agent Simulations

  7. Realistic honeypot evaluations for scheming propensity

    May 28, 2026Victoria Krakovna, David Lindner, Lewis Ho +2Agentic EvaluationsHoneypots

  8. Holistic Evaluation and Failure Diagnosis of AI Agents

    May 14, 2026Netta Madvil, Gilad Dym, Alon Mecilati +12Agentic Evaluations