Human-in-the-Loop Evaluation

Momentum

11 papers in the last four weeks, up 38% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 102

All topics
CardsList
  1. A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

    Oct 7, 2026Alexandra Coroiu, Andrea Vogt, Viktor Werbilo +3RL from Human FeedbackHuman-Robot Collaboration

  2. What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents

    Oct 5, 2026Zhongxiang Sun, Jiahao Yan, Hongkang Zhao +5Software Engineering AgentsHuman-in-the-Loop Evaluation

  3. Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

    Oct 1, 2026Chengguang Gan, Zimeng He, Yoshihiro Tsujii +3AI Agent AuditingWeb Agent Benchmarks

  4. Ontology-Based Contextual AI Evaluations (OB-CAIE) Methodology

    Sep 30, 2026Julie Krugler Hollek, Michael Zargham, Mala KumarHuman-in-the-Loop EvaluationAutomated Evaluation

  5. Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots

    Sep 24, 2026Maithili Patel, Sonia ChernovaHuman-in-the-Loop EvaluationRobotic Policy Evaluation

  6. Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

    Sep 24, 2026Dae Woong, Ham, Xuejun Zhao +2Cost-Aware InferenceHuman-in-the-Loop Evaluation

  7. Spatial Action Review: A Visual Analytics Dashboard for Auditing Language-to-Action Hand-offs in Electron Microscopy

    Sep 21, 2026Samia Mohinta, Albert CardonaHuman-in-the-Loop EvaluationMultimodal QA

  8. Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration

    Sep 21, 2026Promise Ekpo, Teju Vijay, Dhruv Mandalik +5Human-in-the-Loop EvaluationRobot Failure Recovery

  9. EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

    Sep 18, 2026Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona ZahidHuman-in-the-Loop EvaluationLLM Reliability

  10. Beyond Measurement Metrics: A Human-Centered Framework for Semantic Validation of Network Traffic Classification

    Sep 15, 2026Igor Cherepanov, David Sessler, Alex Ulmer +2Human-in-the-Loop EvaluationExplainability Evaluation

  11. A Personalized Dynamic Balance Evaluation Paradigm for Hip Exoskeleton-Assisted Walking under Unexpected Ground Perturbations

    Sep 13, 2026Yun Chen, Oluwasegun T. Akinniyi, Qiang ZhangHuman-in-the-Loop EvaluationAssistive Robotics

  12. Do Reasoning Representations Help Humans Evaluate LLM Outputs?

    Sep 8, 2026Jaewoo Lim, Sungbok Shin, Sanghyun HongLLM EvaluationLLM Interpretability

  13. PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

    Aug 31, 2026Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li +7Scientific VisualizationMulti-Agent LLM Systems

  14. Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment

    Aug 30, 2026Zhiyu Chen, Keyu Zhao, Jigao Fu +8AI Agents for Scientific DiscoveryLLM-as-a-Judge

  15. Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

    Aug 25, 2026Anupam Purwar, Shashank Singh, Kritika SrivastavaVoice Agent EvaluationLLM-as-a-Judge

  16. Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

    Aug 21, 2026Emma Granqvist, Rocío Mercado, Samuel GenhedenAI Agents for Scientific DiscoveryLLM-as-a-Judge

  17. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

    Aug 18, 2026Emma Yanyang Kong, JJ Tan, Ishan Gupta +8LLM EvaluationLLM-as-a-Judge

  18. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

    Aug 13, 2026Dairu Liu, Zekun Qi, Jiayu Zeng +11Human-in-the-Loop EvaluationHumanoid Motion Tracking

  19. Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

    Aug 11, 2026Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo +3Explainable Artificial IntelligenceHuman-in-the-Loop Evaluation

  20. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol

    Aug 9, 2026Christoph TrattnerTool-Use EvaluationHuman-in-the-Loop Evaluation

  21. EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility

    Aug 9, 2026Xudong Wu, Zeqing Wu, Jiarui Zhang +8Human-in-the-Loop EvaluationBuilding Energy Management

  22. Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety

    Aug 8, 2026Hannah Cha, Neha Shukla, Solon Barocas +3Human-in-the-Loop EvaluationAI Safety Evaluation

  23. Dynamically Allocating Evaluation Effort for Model Ranking

    Aug 4, 2026Vilém Zouhar, Julia Kreutzer, Alon Lavie +4Model SelectionMulti-Armed Bandits

  24. Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

    Aug 3, 2026Zejun Xie, Xintong Li, Guang Wang +1Post-Hoc CalibrationLearning to Rank

  25. MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models

    Aug 3, 2026Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian +1LLM EvaluationHuman-in-the-Loop Evaluation

  26. Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews

    Jul 27, 2026Barbara Kitchenham, Sebastián Pizard, Lech Madeyski +3LLM EvaluationHuman-in-the-Loop Evaluation