Benchmark Suite

Recent momentum

-64%

9 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

4 new papers

A weekly snapshot of new work published in Benchmark Suite.

Period ending 2026-09-07

6 new papers

A weekly snapshot of new work published in Benchmark Suite.

163 papers

Latest in Benchmark Suite

Open your feed →
CardsList
  1. Peg-in-Bench: A Modular Benchmark for High-Precision Robotic Insertion

    Sep 1, 2026Yosel Delgado, José G. Buenaventura-Carreón, Floris Erich +4Robotic ManipulationAssembly

  2. APEX-Accounting

    Jul 29, 2026Julien Benchek, Austin Bennett, Jasmin Kern +8Frontier ModelsBenchmark Suite

  3. LEMUR 2: Unlocking Neural Network Diversity for AI

    Jul 7, 2026Tolgay Atinc Uzun, Waleed Khalid, Saif U Din +17Benchmark SuiteBird

  4. A Fair Benchmarking of Deep Relational Database Learning Models

    Jul 4, 2026Kazi F. Akhter, Bharath Ajendla, Manar D. SamadTabular LearningRelational

  5. PACE: A Proxy for Agentic Capability Evaluation

    Jul 2, 2026Yueqi Song, Lintang Sutawika, Jiarui Liu +8Agentic BenchmarksAgentic Evaluations

  6. Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

    Jun 25, 2026Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi +1Indirect Prompt InjectionLlm-Based Agent