Evaluation Benchmarks

Recent momentum

-81%

4 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in Evaluation Benchmarks.

Period ending 2026-09-07

1 new paper

A weekly snapshot of new work published in Evaluation Benchmarks.

145 papers

Latest in Evaluation Benchmarks

Open your feed →
CardsList
  1. FailBench: How Reliable are VLMs at Judging Robot Task Success?

    Sep 3, 2026Zaruhi Navasardyan, Tatul Danielyan, Hrant DavtyanRobotic ManipulationRobots

  2. ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

    Aug 31, 2026Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti +11Multimodal EvaluationEvaluation Benchmarks