Human-in-the-Loop Evaluation

Momentum

11 papers in the last four weeks, up 38% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 102

All topics
CardsList
  1. A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

    Oct 7, 2026Alexandra Coroiu, Andrea Vogt, Viktor Werbilo +3RL from Human FeedbackHuman-Robot Collaboration

  2. What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

    Oct 5, 2026Zhongxiang Sun, Jiahao Yan, Hongkang Zhao +5Software Engineering AgentsHuman-in-the-Loop Evaluation

  3. Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

    Oct 1, 2026Chengguang Gan, Zimeng He, Yoshihiro Tsujii +3AI Agent AuditingWeb Agent Benchmarks

  4. Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology

    Sep 30, 2026Julie Krugler Hollek, Michael Zargham, Mala KumarHuman-in-the-Loop EvaluationAutomated Evaluation

  5. Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots

    Sep 24, 2026Maithili Patel, Sonia ChernovaHuman-in-the-Loop EvaluationRobotic Policy Evaluation

  6. Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

    Sep 24, 2026Dae Woong, Ham, Xuejun Zhao +2Cost-Aware InferenceHuman-in-the-Loop Evaluation

  7. Spatial Action Review: A Visual Analytics Dashboard for Auditing Language-to-Action Hand-offs in Electron Microscopy

    Sep 21, 2026Samia Mohinta, Albert CardonaHuman-in-the-Loop EvaluationMultimodal QA

  8. Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration

    Sep 21, 2026Promise Ekpo, Teju Vijay, Dhruv Mandalik +5Human-in-the-Loop EvaluationRobot Failure Recovery

  9. EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

    Sep 18, 2026Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona ZahidHuman-in-the-Loop EvaluationLLM Reliability

  10. Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

    Sep 15, 2026Igor Cherepanov, David Sessler, Alex Ulmer +2Human-in-the-Loop EvaluationExplainability Evaluation

  11. A Personalized Dynamic Balance Evaluation Paradigm for Hip Exoskeleton-Assisted Walking under Unexpected Ground Perturbations

    Sep 13, 2026Yun Chen, Oluwasegun T. Akinniyi, Qiang ZhangHuman-in-the-Loop EvaluationAssistive Robotics

  12. Do Reasoning Representations Help Humans Evaluate LLM Outputs?

    Sep 8, 2026Jaewoo Lim, Sungbok Shin, Sanghyun HongLLM EvaluationLLM Interpretability

  13. PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

    Aug 31, 2026Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li +7Scientific VisualizationMulti-Agent LLM Systems

  14. Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment

    Aug 30, 2026Zhiyu Chen, Keyu Zhao, Jigao Fu +8AI Agents for Scientific DiscoveryLLM-as-a-Judge

  15. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

    Aug 25, 2026Anupam Purwar, Shashank Singh, Kritika SrivastavaVoice Agent EvaluationLLM-as-a-Judge

  16. Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

    Aug 21, 2026Emma Granqvist, Rocío Mercado, Samuel GenhedenAI Agents for Scientific DiscoveryLLM-as-a-Judge

  17. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

    Aug 18, 2026Emma Yanyang Kong, JJ Tan, Ishan Gupta +8LLM EvaluationLLM-as-a-Judge

  18. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

    Aug 13, 2026Dairu Liu, Zekun Qi, Jiayu Zeng +11Human-in-the-Loop EvaluationHumanoid Motion Tracking

  19. Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

    Aug 11, 2026Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo +3Explainable Artificial IntelligenceHuman-in-the-Loop Evaluation

  20. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol

    Aug 9, 2026Christoph TrattnerTool-Use EvaluationHuman-in-the-Loop Evaluation

  21. EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

    Aug 9, 2026Xudong Wu, Zeqing Wu, Jiarui Zhang +8Human-in-the-Loop EvaluationBuilding Energy Management

  22. Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

    Aug 8, 2026Hannah Cha, Neha Shukla, Solon Barocas +3Human-in-the-Loop EvaluationAI Safety Evaluation

  23. Dynamically Allocating Evaluation Effort for Model Ranking

    Aug 4, 2026Vilém Zouhar, Julia Kreutzer, Alon Lavie +4Model SelectionMulti-Armed Bandits

  24. Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

    Aug 3, 2026Zejun Xie, Xintong Li, Guang Wang +1Post-Hoc CalibrationLearning to Rank

  25. MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

    Aug 3, 2026Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian +1LLM EvaluationHuman-in-the-Loop Evaluation

  26. Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

    Jul 27, 2026Barbara Kitchenham, Sebastián Pizard, Lech Madeyski +3LLM EvaluationHuman-in-the-Loop Evaluation

  27. HARP: The Human--AI Research Platform

    Jul 22, 2026Zeshu Zhu, Natalie Friedman, Kevin Weatherwax +1Human-in-the-Loop EvaluationHuman-AI Interaction

  28. PhaseAware: Interpretable Human-in-the-Loop Rehabilitation Scoring with Boundary Monitoring

    Jul 22, 2026Yankai Zheng, Yuhe Liu, Yuxin Ma +8Interpretable MLHuman-in-the-Loop Evaluation

  29. EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration

    Jul 20, 2026Jia-Kai Dong, Yi-Cheng Lin, Hung-yi LeeLLM-as-a-JudgeEducational Assessment

  30. Human Grounded Evaluation of Large Language Models for Optical Network Automation

    Jul 20, 2026Kiarash Rezaei, Omran Ayoub, Paolo Monti +1LLM Inference EfficiencyLLM Evaluation

  31. Human-in-the-Loop User Feedback Affects Perceived Accuracy and Trust, but Task Subjectivity Matters

    Jul 20, 2026Donald R. Honeycutt, Mahsan Nourani, Eric D. RaganHuman Preference EvaluationHuman-in-the-Loop Evaluation

  32. Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

    Jul 16, 2026Leanne Tan, Rohan Jaggi, Shaun Khoo +1Human-in-the-Loop EvaluationRubric-Based Evaluation

  33. Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

    Jul 8, 2026Mingguang Chen, Licheng Wang, Bo QuLLM Self-RefinementHuman-in-the-Loop Evaluation

  34. Manual, Joystick, or Haptic Control? An In Vitro Comparison of Navigation Strategies for Robotic Interventional Neuroradiology Procedures

    Jul 8, 2026Benjamin Jackson, Nikola Fischer, Harry Robershaw +16Robot TeleoperationHuman-in-the-Loop Evaluation

  35. HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation

    Jul 5, 2026Yaozu Wu, Wei-Chieh Huang, Jizhou Guo +11LLM Agent EvaluationHuman-in-the-Loop Evaluation

  36. LOTUSim: Multi-Domain Simulator for Marine Robotics

    Jul 3, 2026Cédric Buche, Juliette Grosset, Hélène Lechêne +4Multi-Robot SystemsRobotics Simulation

  37. CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

    Jun 30, 2026Ajmal M., Abin Roy, Afthab Salam Kanniyan +4LLM-as-a-JudgeHuman-in-the-Loop Evaluation

  38. HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

    Jun 29, 2026Xinrui Ruan, Zhenyu Zhao, Waverly Wei +4Human-in-the-Loop EvaluationGenerative AI Evaluation

  39. Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study

    Jun 24, 2026Giulian Biolo, Michael Tezza, Yuanjun Gong +1Automated Program RepairHuman-in-the-Loop Evaluation

  40. Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

    Jun 23, 2026Tian Zheng, Kai-Tai HsuLLM-as-a-JudgeHuman-in-the-Loop Evaluation

  41. When CQs Go Wrong: Challenges in CQ Verification with OE-Assist

    Jun 23, 2026Anna Sofia Lippolis, Mohammad Javad Saeedizade, Robin Keskisärkkä +3Human-in-the-Loop Evaluation

  42. Judgment-Grounded Expansion for Peer Review Generation

    Jun 22, 2026Sheng Lu, Lizhen Qu, Iryna GurevychHuman-in-the-Loop EvaluationAutomated Peer Review

  43. Counsel: A Meta-Evaluation Dataset for Agentic Tasks

    Jun 19, 2026Sashank Pisupati, Henry Broomfield, Eujeong Choi +5LLM-as-a-JudgeLLM Agent Evaluation

  44. AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

    Jun 18, 2026Zilong Zhang, Yi-Ting Hung, Weiyi He +3LLM-as-a-JudgeLLM Auditing

  45. A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies

    Jun 17, 2026Fangyijie Wang, Jianjun Yu, Wentao Shi +4Human-in-the-Loop AnnotationHuman-in-the-Loop Evaluation

  46. RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills

    Jun 16, 2026Weizhi Zhang, Zechen Li, Hamid Palangi +16AI Agent EvaluationHuman-in-the-Loop Evaluation

  47. Is Your Trajectory Displacement Safe in Long-tail?

    Jun 15, 2026Qiao Sun, Weicheng Zheng, Yixin Huang +1Autonomous Driving Safety EvaluationAutonomous Driving

  48. Ride, Track, and Recover: Pilot Randomized Trial of a Wearable Digital Self-Management Intervention During a Veteran Endurance-Cycling Program

    Jun 11, 2026Alan Ta, Nilsu Salgin, Caleb Armstrong +2Mental HealthHuman-in-the-Loop Evaluation

  49. RogueAI: A Reverse Turing Test for Detecting Licensed AI Deception in Dialogue

    Jun 11, 2026Sara Candussio, Emanuele Ballarin, Lorenzo Bonin +2AI Agent ReliabilityLanguage Model Safety Evaluation

  50. IntElicit: Eliciting and Assessing Contextualized Creativity via Dialogue Policy Optimization

    Jun 10, 2026Mingjia Li, Jin Wu, Hong Qian +7Creativity AssessmentPolicy Optimization