Automated Evaluation

Momentum

48 papers in the last four weeks, up 12% on the four weeks before. 0.3% of all new papers.

Jul 13Week of Sep 28

Latest papers 321

All topics
CardsList
  1. WANDR: A Benchmark for Wide and Deep Research

    Aug 14, 2026Vitaliy Polshkov, Marcin Pitera, Jeremy Yang +7AI Agent BenchmarksDeep Research Agents

  2. RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level

    Aug 13, 2026Juan Irving Vasquez, Juan Terven, Laura-Ivoone Garay-JimenezAutomated Evaluation

  3. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

    Aug 13, 2026Chenrun Wang, Mingxuan Zhu, Tiancheng Huang +6LLM EvaluationBenchmark Design

  4. CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

    Aug 13, 2026Nhan Phan, Ilona Lähteenmäki, Anna von Zansen +4Automated EvaluationSpeech Language Models

  5. ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

    Aug 13, 2026Jiale Cui, Yueyao Yuan, Kaixi Zhong +3Human Preference EvaluationAI Agent Benchmarks

  6. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment

    Aug 11, 2026Xiaokang Qu, Yiting LinAutomated Evaluation

  7. FormStruct-Bench:A Hierarchical and Diagnostic Benchmark for Table-Form Document Structure Recognition

    Aug 11, 2026Lujie Ban, Jiangtao Zhu, Yuanheng Yu +2Document UnderstandingTable Structure Recognition

  8. RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement

    Aug 10, 2026Ziheng Jia, Jiaying Qian, Zicheng Zhang +3LLM-as-a-JudgeRubric-Based Evaluation

  9. Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets

    Aug 8, 2026Ximeng Liu, Qianlong Wang, Yingming Mao +6Surrogate-Assisted OptimizationLarge Language Model-Guided Optimization

  10. SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

    Aug 7, 2026Yuheng Zhang, Yuanchun Wang, Fanjin Zhang +4LLM-as-a-JudgeAutomated Evaluation

  11. Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering

    Aug 7, 2026Xiuwei Shang, Li Hu, Xiao Jiang +7LLM-as-a-JudgeSoftware Reverse Engineering

  12. SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

    Aug 7, 2026Jiahui Han, Qinuo Li, Ziheng Peng +6LLM Agent EvaluationAutomated Evaluation

  13. Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

    Aug 6, 2026Hotaka Maeda, Yikai LuEducational AssessmentText Classification

  14. Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

    Aug 6, 2026Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer +3Document Information ExtractionAutomated Evaluation

  15. Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

    Aug 4, 2026S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh +1LLM EvaluationTool-Use Evaluation

  16. LiveEvalBench: Toward Open-World Evaluation for Web Generation

    Aug 4, 2026Yiyao Wang, Zhen Wen, Yinghao Tang +5Web Application GenerationAutomated Evaluation

  17. Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation

    Aug 4, 2026Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright +2Automated EvaluationMT Evaluation

  18. Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search

    Aug 3, 2026Han Wang, Alex Whitworth, Pak Ming Cheung +5VLM EvaluationAutomated Evaluation

  19. ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision

    Aug 3, 2026Wei-Jung Huang, Bonan ShenPairwise ComparisonLLM Agent Evaluation

  20. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

    Aug 3, 2026Can Wang, Haoran Chen, Haowen Gao +3LLM EvaluationBenchmark Design

  21. Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

    Aug 2, 2026Shengwei Xu, Yuxuan Lu, Yifan Wu +2Language Model Generation EvaluationAutomated Evaluation

  22. Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

    Aug 2, 2026Tezan Sahu, Aritra Das, Pankaj Mittal +1Benchmark ConstructionLLM Agent Evaluation

  23. Slides2MindMap: Reconstructing Cognitively Efficient Knowledge Hierarchies from Lecture Slides

    Aug 1, 2026Yuzhi Wang, Rongjun Ye, Shengyuan Chen +5Document UnderstandingEducational Technology

  24. CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

    Jul 31, 2026Mengting Chen, Yanshu Sun, Wanting Liang +5Open-Ended GenerationLLM Evaluation

  25. Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

    Jul 31, 2026Jeonghwan Choi, Taewon Yun, Minjeong Ban +3Reference-Free EvaluationRAG Evaluation