Human-in-the-Loop Evaluation

Momentum

11 papers in the last four weeks, up 38% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 102

All topics
CardsList
  1. End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians

    Apr 30, 2026Aaryan Shah, Andrew Hines, Alexia Downs +6HealthcareAI Agent Evaluation

  2. Human-in-the-Loop Benchmarking of Heterogeneous LLMs for Automated Competency Assessment in Secondary Level Mathematics

    Apr 29, 2026Jatin Bhusal, Nancy Mahatha, Aayush Acharya +1LLM EvaluationEducational Assessment

  3. Understanding the Limits of Automated Evaluation for Code Review Bots in Practice

    Apr 27, 2026Veli Karakaya, Utku Boran Torun, Baykal Mehmet Uçar +1LLM-as-a-JudgeHuman-in-the-Loop Evaluation

  4. Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop

    Apr 27, 2026Ashmi Banerjee, Adithi Satish, Wolfgang Wörndl +1LLM EvaluationLLM-as-a-Judge

  5. Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings

    Apr 24, 2026Inês Oliveira e Silva, Sérgio Jesus, Iker Perez +4Human-in-the-Loop EvaluationShapley Value Attribution

  6. CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

    Apr 20, 2026Yijia Shao, Zora Zhiruo Wang, Neel Ahuja +3AI Agent EvaluationHuman-in-the-Loop Evaluation

  7. AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgment

    Apr 16, 2026Oz Levy, Ilya Dikman, Natan Levy +1Requirements EngineeringHuman-in-the-Loop Evaluation

  8. Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback

    Mar 13, 2026Yuki Hirakawa, Takashi Wada, Ryotaro Shimizu +6Virtual Try-OnHuman-in-the-Loop Evaluation

  9. Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VLM Raters

    Feb 1, 2026Lana Do, Gio Jung, Juvenal Francisco Barajas +5VLM EvaluationHuman-in-the-Loop Evaluation

  10. Althea: The Fact-Checking--Metalearning Tradeoff in AI-Assisted Verification

    Dec 29, 2025Svetlana Churina, Kokil Jaidka, Anab Maulana Barik +5Human-in-the-Loop EvaluationAutomated Fact-Checking

  11. No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

    Mar 7, 2025Michael Krumdick, Charles Lovering, Varshini Reddy +2LLM EvaluationLLM-as-a-Judge