Llm-As-A-Judge

Recent momentum

-7%

14 papers in the last 28 days · 0.4% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

6 new papers

A weekly snapshot of new work published in Llm-As-A-Judge.

Period ending 2026-09-07

5 new papers

A weekly snapshot of new work published in Llm-As-A-Judge.

167 papers

Latest in Llm-As-A-Judge

Open your feed →
CardsList
  1. Does task decomposition improve automatic NLG evaluation?

    Sep 1, 2026Sebastian Steindl, Nikos Voskarides, Alberto Gasparin +1Llm-As-A-JudgeHuman Annotators

  2. Post-hoc Alignment of LLM-judges to Human Judgment Distribution

    Sep 1, 2026Sebastian Steindl, Nikos Voskarides, Alberto Gasparin +1Soft-LabelLlm-As-A-Judge

  3. Automated Textbook Auditing with Multi-Agent LLM Systems

    Jul 13, 2026Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran +1Algorithm AuditingLlm-As-A-Judge

  4. Andha-Dhun: A First Look at Audio Descriptions in Hindi

    Jul 7, 2026Ritabrata Chakraborty, Divy Kala, Nisheeth Bhooshan Gupta +3Indian LanguagesHuman Annotations

  5. RoPoLL: Robust Panel of LLM Judges

    Jun 29, 2026Anish Acharya, Kris W Pan, Brian VerkhovskyLlm-As-A-JudgeLarge Language Model Judges

  6. Counsel: A Meta-Evaluation Dataset for Agentic Tasks

    Jun 19, 2026Sashank Pisupati, Henry Broomfield, Eujeong Choi +5Agentic BenchmarksLlm-As-A-Judge