Large-Scale Benchmark

Recent momentum

-44%

5 papers in the last 28 days · 0.1% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

3 new papers

A weekly snapshot of new work published in Large-Scale Benchmark.

64 papers

Latest in Large-Scale Benchmark

Open your feed →
CardsList
  1. Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

    Sep 8, 2026Maximilian Schall, Sedigheh Eslami, Markus Krimmel +4RetrieversRelevance

  2. Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

    May 21, 2026Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11Large Language Model SafetySafety Evaluation

  3. FactoryBench: Evaluating Industrial Machine Understanding

    May 8, 2026Yanis Merzouki, Coral Izquierdo, Matei Ignuta-Ciuncanu +8Circular FactoryLarge-Scale Benchmark