Automated Evaluation

Momentum

36 papers in the last four weeks, up 57% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 289

All topics
CardsList
  1. Can LLMs Fix It Without Code? Toward Automated Verification of No-Code Bug Fixes

    Oct 8, 2026Utku Boran Torun, Veli Karakaya, Eray TüzünAutomated EvaluationLLM Agent Reliability

  2. Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

    Oct 8, 2026Xing Zhang, Guanghui Wang, Yanwei Cui +4Automated EvaluationLLM Agent Self-Improvement

  3. LLM Persuasion Is in the Eye of the Evaluation

    Oct 7, 2026Kamile Dementaviciute, Julija Vaitonyte, Tijl De BiePersuasion in Language ModelsLLM Evaluation

  4. Perceptually Aligned Evaluation of Style Transfer

    Oct 7, 2026Yang Deng, Eleftherios Ioannou, David Mould +3Image Style TransferAutomated Evaluation

  5. CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

    Oct 5, 2026Berke Arda, Ahmetcan Yavuz, Paul Gerry +7Document Information ExtractionLLM Information Extraction

  6. Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation

    Oct 1, 2026Noy Sternlicht, Simra Shahid, Peter Jansen +3Prompt SensitivityLLM-as-a-Judge

  7. Judgement in the Age of Jev: From Evaluation Scarcity to Evaluation Abundance

    Oct 1, 2026Richard HillAutomated EvaluationHuman-AI Decision Making

  8. When Does Exercise-Specific Joint Selection Help? An Audit of Evaluation and Control Design

    Oct 1, 2026Haotian Chen, Jingkun Yu, Yuning Zhang +1Skeleton-Based Action RecognitionClassification

  9. Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study

    Oct 1, 2026Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael ManiasDocument UnderstandingDomain Generalization

  10. Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology

    Sep 30, 2026Julie Krugler Hollek, Michael Zargham, Mala KumarHuman-in-the-Loop EvaluationAutomated Evaluation

  11. LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian

    Sep 30, 2026Claudiu Creanga, Liviu P. DinuLLM EvaluationLLM-as-a-Judge

  12. Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts

    Sep 29, 2026Anton Repushko, Elena ChepelAutomated EvaluationOptical Character Recognition

  13. Do Music Generative Models Understand Musical Qualities? Automatic Music Evaluation with Model-Intrinsic Signals

    Sep 29, 2026Xiaosha Li, Chun Liu, Ziyu WangMusic Generation EvaluationAutomated Evaluation

  14. Mubric: Mutation Testing-Guided Rubric Generation for LLM Evaluation

    Sep 29, 2026Jiayuxuan Yang, Jie M. Zhang, Yiling Lou +1Language Model Generation EvaluationRubric-Based Evaluation

  15. SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression

    Sep 28, 2026Ziwen Zhang, Xiju Wu, Yuheng Jing +9Benchmark DesignAutomated Evaluation

  16. TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation

    Sep 28, 2026Nicolas Dahan, Fran{\cc}ois Yvon, Rachel BawdenAutomated EvaluationDocument-Level MT

  17. SIVIA-RSI: Source-Grounded Adaptation of Diagramming Skills

    Sep 27, 2026Feng Yuan, Yifan Gao, Haoyue Li +1Cross-Domain Transfer LearningAgent Skill Learning

  18. CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation

    Sep 27, 2026Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad VittalaHallucination in Language ModelsFactual Consistency Evaluation

  19. SEEK: Skill-Routed Evaluation with Evolvable Knowledge for Industrial Search

    Sep 24, 2026Zhongxin Huang, Songyang Li, Renzhe Zhou +5LLM EvaluationAutomated Evaluation

  20. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

    Sep 22, 2026Yubo Li, Yidi Miao, Ramayya Krishnan +1LLM-as-a-JudgeLanguage Model Generation Evaluation

  21. Auditing Proxy-Based Validation Across Text Spans

    Sep 22, 2026Daein Weon, Dong Ho KangBenchmark ValidityAutomated Evaluation

  22. LLJ Cards: Best practices for the Use of LLMs as Judges

    Sep 21, 2026Khaoula Chehbouni, Melina Medjdoub, Florian Carichon +2LLM EvaluationLLM-as-a-Judge

  23. Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation

    Sep 18, 2026Nikhil Reddy Pottanigari, Ramin Fahimi, Noah Bolger +2LLM EvaluationText Summarization

  24. greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI

    Sep 17, 2026Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh +1Authorship AttributionAutomated Evaluation

  25. E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews

    Sep 17, 2026Haoshen Wang, Dongbo Che, Zeyi Xie +3Multimodal GroundingAudio-Visual Understanding