Evaluation Benchmarks

Recent momentum

-64%

5 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

3 new papers

A weekly snapshot of new work published in Evaluation Benchmarks.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in Evaluation Benchmarks.

Period ending 2026-09-07

1 new paper

A weekly snapshot of new work published in Evaluation Benchmarks.

96 papers

Latest in Evaluation Benchmarks

  1. The Tool Illusion: Rethinking Tool Use in Web Agents

    Apr 3, 2026Renze Lou, Baolin Peng, Wenlin Yao +5Web AgentsBrowser Agent

  2. Multimodal Language Models as Text-to-Image Model Evaluators

    May 1, 2025Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat +5Text-To-ImageEvaluation Benchmarks

  3. EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

    Date pendingTara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10Voice AgentsEvaluation Benchmarks