Automated Evaluation

Momentum

48 papers in the last four weeks, up 12% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 321

All topics
CardsList
  1. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1LLM-as-a-JudgeLanguage Model Generation Evaluation

  2. Auditing Proxy-Based Validation Across Text Spans

    Sep 22, 2026Daein Weon, Dong Ho KangBenchmark ValidityAutomated Evaluation

  3. LLJ Cards: Best practices for the Use of LLMs as Judges

    Sep 21, 2026Khaoula Chehbouni, Melina Medjdoub, Florian Carichon +2LLM EvaluationLLM-as-a-Judge

  4. Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants

    Sep 20, 2026Vaishnav Negi, Lev Sorokin, Soroosh Tayebi Arasteh +1Multi-Turn Dialogue EvaluationVoice Agent Evaluation

  5. Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

    Sep 18, 2026Nikhil Reddy Pottanigari, Ramin Fahimi, Noah Bolger +2LLM EvaluationText Summarization

  6. Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction

    Sep 18, 2026Ruotian Wu, Bill E. Johnson, Gene Saunders +3Grammatical Error CorrectionAutomated Evaluation

  7. greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI

    Sep 17, 2026Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh +1Authorship AttributionAutomated Evaluation

  8. E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

    Sep 17, 2026Haoshen Wang, Dongbo Che, Zeyi Xie +3Multimodal GroundingAudio-Visual Understanding

  9. A Benchmark Framework for Screening Automation in Systematic Reviews

    Sep 16, 2026Gauransh Kumar, Luciano Marchezan, Guillaume Genois +2LLM EvaluationEvidence Selection

  10. I code or AI code: A comparative evaluation of AI-rated scores in classroom observations

    Sep 16, 2026Y. Fong, J. Xiang, T. Y. D. Chan +2LLM-as-a-JudgeEducational Assessment

  11. The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting

    Sep 14, 2026Hisham Ihshaish, Peter Mayhew, Tasnim M. A. Zayet +1Text ClassificationAutomated Evaluation

  12. Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction

    Sep 14, 2026Hayeong Ryu, Sunhee Jo, Seunguk Yu +1Reward ModelingReference-Free Evaluation

  13. WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

    Sep 14, 2026Chenxu Liu, Zilu Zou, Peizhong Gao +9Automated Software TestingWeb Application Generation

  14. ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment

    Sep 14, 2026Junkai Tong, Mingjia Li, Haoran Chen +4Educational AssessmentAutomated Evaluation

  15. AppliedScientist: Automated Scientific Revision Through Iterative AI Reviewing

    Sep 13, 2026Vidushee Vats, Karun Sharma, Shengzhi Li +1AI-Assisted Scientific ResearchAutomated Peer Review

  16. PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews

    Sep 10, 2026Miguel Zabaleta, Baihan LinLLM AuditingAI-Assisted Scientific Research

  17. Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

    Sep 10, 2026Ziyu Zhang, Satoshi NakamuraRubric-Based EvaluationAutomated Evaluation

  18. SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL Evaluation

    Sep 8, 2026Mohammadhossein Malekpour, Mohamed Riahi, Maxime Lamothe +1Mutation TestingText-to-SQL

  19. Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

    Sep 8, 2026Antonin Poché, Fanny Jourdan, Nils Feldhus +6LLM EvaluationAutomated Evaluation

  20. EviSI: An Evidence-Based Evaluation Agent for Simultaneous Interpreting

    Sep 8, 2026Ben Yan, Zongyao Li, Xiaoyu Chen +6LLM-as-a-JudgeSimultaneous Speech Translation

  21. Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection

    Sep 6, 2026Ashanvi Yadav, Shubham BhardwajHate Speech DetectionContent Moderation

  22. Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

    Sep 4, 2026Maria Mahbub, Ashley Rice, Michael R. Munroe +2Automated EvaluationLanguage Model Generation Evaluation

  23. Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

    Sep 3, 2026Xiangyu Wang, Jin Wu, Xiaoyu Li +2Creativity AssessmentLLM-as-a-Judge