Human-in-the-Loop Evaluation

Momentum

11 papers in the last four weeks, up 38% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 102

All topics
CardsList
  1. HARP: The Human--AI Research Platform

    Jul 22, 2026Zeshu Zhu, Natalie Friedman, Kevin Weatherwax +1Human-in-the-Loop EvaluationHuman-AI Interaction

  2. PhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary Monitoring

    Jul 22, 2026Yankai Zheng, Yuhe Liu, Yuxin Ma +8Interpretable MLHuman-in-the-Loop Evaluation

  3. EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

    Jul 20, 2026Jia-Kai Dong, Yi-Cheng Lin, Hung-yi LeeLLM-as-a-JudgeEducational Assessment

  4. Human Grounded Evaluation of Large Language Models for Optical Network Automation

    Jul 20, 2026Kiarash Rezaei, Omran Ayoub, Paolo Monti +1LLM Inference EfficiencyLLM Evaluation

  5. Human-in-the-Loop User Feedback Affects Perceived Accuracy and Trust, but Task Subjectivity Matters

    Jul 20, 2026Donald R. Honeycutt, Mahsan Nourani, Eric D. RaganHuman Preference EvaluationHuman-in-the-Loop Evaluation

  6. Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

    Jul 16, 2026Leanne Tan, Rohan Jaggi, Shaun Khoo +1Human-in-the-Loop EvaluationRubric-Based Evaluation

  7. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

    Jul 8, 2026Mingguang Chen, Licheng Wang, Bo QuLLM Self-RefinementHuman-in-the-Loop Evaluation

  8. Manual, Joystick, or Haptic Control? An In Vitro Comparison of Navigation Strategies for Robotic Interventional Neuroradiology Procedures

    Jul 8, 2026Benjamin Jackson, Nikola Fischer, Harry Robershaw +16Robot TeleoperationHuman-in-the-Loop Evaluation

  9. HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

    Jul 5, 2026Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11LLM Agent EvaluationHuman-in-the-Loop Evaluation

  10. LOTUSim: Multi-Domain Simulator for Marine Robotics

    Jul 3, 2026Cédric Buche, Juliette Grosset, Hélène Lechêne +4Multi-Robot SystemsRobotics Simulation

  11. CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

    Jun 30, 2026Ajmal M., Abin Roy, Afthab Salam Kanniyan +4LLM-as-a-JudgeHuman-in-the-Loop Evaluation

  12. HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

    Jun 29, 2026Xinrui Ruan, Zhenyu Zhao, Waverly Wei +4Human-in-the-Loop EvaluationGenerative AI Evaluation

  13. Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study

    Jun 24, 2026Giulian Biolo, Michael Tezza, Yuanjun Gong +1Automated Program RepairHuman-in-the-Loop Evaluation

  14. Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

    Jun 23, 2026Tian Zheng, Kai-Tai HsuLLM-as-a-JudgeHuman-in-the-Loop Evaluation

  15. When CQs Go Wrong: Challenges in CQ Verification with OE-Assist

    Jun 23, 2026Anna Sofia Lippolis, Mohammad Javad Saeedizade, Robin Keskisärkkä +3Human-in-the-Loop Evaluation

  16. Judgment-Grounded Expansion for Peer Review Generation

    Jun 22, 2026Sheng Lu, Lizhen Qu, Iryna GurevychHuman-in-the-Loop EvaluationAutomated Peer Review

  17. Counsel: A Meta-Evaluation Dataset for Agentic Tasks

    Jun 19, 2026Sashank Pisupati, Henry Broomfield, Eujeong Choi +5LLM-as-a-JudgeLLM Agent Evaluation

  18. AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

    Jun 18, 2026Zilong Zhang, Yi-Ting Hung, Weiyi He +3LLM-as-a-JudgeLLM Auditing

  19. A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies

    Jun 17, 2026Fangyijie Wang, Jianjun Yu, Wentao Shi +4Human-in-the-Loop AnnotationHuman-in-the-Loop Evaluation

  20. RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

    Jun 16, 2026Weizhi Zhang, Zechen Li, Hamid Palangi +16AI Agent EvaluationHuman-in-the-Loop Evaluation

  21. Is Your Trajectory Displacement Safe in Long-tail?

    Jun 15, 2026Qiao Sun, Weicheng Zheng, Yixin Huang +1Autonomous Driving Safety EvaluationAutonomous Driving

  22. Ride, Track, and Recover: Pilot Randomized Trial of a Wearable Digital Self-Management Intervention During a Veteran Endurance-Cycling Program

    Jun 11, 2026Alan Ta, Nilsu Salgin, Caleb Armstrong +2Mental HealthHuman-in-the-Loop Evaluation

  23. RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue

    Jun 11, 2026Sara Candussio, Emanuele Ballarin, Lorenzo Bonin +2AI Agent ReliabilityLanguage Model Safety Evaluation

  24. IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization

    Jun 10, 2026Mingjia Li, Jin Wu, Hong Qian +7Creativity AssessmentPolicy Optimization