Human-in-the-Loop Evaluation

Momentum

11 papers in the last four weeks, up 38% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 102

All topics
CardsList
  1. Contemporary AI lacks the imagination to diverge or negate in science

    Jun 6, 2026Honglin Bao, Siyang Wu, Xiao Liu +3LLM-as-a-JudgeScientific Reasoning

  2. Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation

    Jun 6, 2026Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5Human-in-the-Loop EvaluationScientific Reproducibility

  3. CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

    May 30, 2026Fangzhou Lin, Peiran Li, Lingyu Xu +12Image Editing EvaluationText-Guided Image Editing

  4. Bridging Chemists and AI: An Expert-Augmented Framework for Interpretable Route Evaluation

    May 27, 2026Yujia Guo, Mikhail Kabeshov, Tat Hong Duong Le +6Human-in-the-Loop EvaluationDeep Sets

  5. Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

    May 27, 2026Petras Swissler, Mohammadali Rashidioun, Nicholas Sahu +3Multi-Robot SystemsSwarm Robotics

  6. GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

    May 26, 2026Yihang Lin, Yunze Gao, Zeyang Lin +3Benchmark DesignHuman-in-the-Loop Evaluation

  7. Grounding Text Embeddings in Stakeholder Associations

    May 26, 2026Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu +3Text Embedding EvaluationHuman-in-the-Loop Evaluation

  8. Workflow Closure Is Not Scientific Closure in Auto-Research Systems

    May 25, 2026Shuai Wang, Xinyuan Tian, Pangpang Liu +1Human-in-the-Loop EvaluationAutonomous Scientific Discovery

  9. From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

    May 24, 2026Most. Sharmin Sultana Samu, MD. Tanvir Ahmed Seum, Md. Rakibul IslamLLM AuditingHuman-in-the-Loop Annotation

  10. Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

    May 21, 2026Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2HealthcareBenchmark Design

  11. Addressing the Synergy Gap: The Six Elements of the Design Space

    May 20, 2026Tommaso Turchi, Ben Wilson, Matt Roach +2Human-in-the-Loop EvaluationHuman-Centered AI

  12. Can Conversational XAI Improve User Performance? An Experimental Study

    May 19, 2026Sven Kruschel, Julian Rosenberger, Lasse Bohlen +2Explainable Artificial IntelligenceHuman-in-the-Loop Evaluation

  13. QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AI

    May 17, 2026Marjan Veysi, Pirooz Shamsinejadbabaki, Mohammad Zare +1LLM-as-a-JudgeHuman-in-the-Loop Evaluation

  14. Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

    May 13, 2026Deepak Pandita, Flip Korn, Chris Welty +1Inter-Rater ReliabilityLLM Evaluation

  15. Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation

    May 13, 2026Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait +4Human-in-the-Loop EvaluationAutomated Evaluation

  16. Elicitation-Augmented Bayesian Optimization

    May 12, 2026Alvar Haltia, Ville Hyvönen, Samuel KaskiBayesian OptimizationHuman-in-the-Loop Evaluation

  17. LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation

    May 11, 2026Philipp Steigerwald, Mara Stieler, Jennifer Burghardt +2Language Model Generation EvaluationLLM Prompting

  18. Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment

    May 8, 2026I. F. Atasoy, B. Mutlu, E. A. Sezer +1LLM EvaluationLLM-Assisted Annotation

  19. Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?

    May 8, 2026Jane Paik KimLLM EvaluationLLM-as-a-Judge

  20. PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialogue

    May 7, 2026Hyunbae Jeon, Jinho D. ChoiVoice Agent EvaluationSpoken Dialogue Systems

  21. Why Expert Alignment Is Hard: Evidence from Subjective Evaluation

    May 6, 2026Tzu-Mi Lin, Wataru Hirota, Tatsuya Ishigaki +2Human Preference EvaluationLLM Alignment

  22. STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

    May 4, 2026Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6Annotator DisagreementAI Agent Evaluation

  23. Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

    May 3, 2026Christopher Kelly, Angelica Chowdhury, Alexandra Campili +5Human-in-the-Loop EvaluationRubric-Based Evaluation

  24. Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving

    May 1, 2026Ashwin George, Lucas Elbert Suryana, Lorenzo Flipse +5Autonomous DrivingHuman-in-the-Loop Evaluation

  25. HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

    Apr 30, 2026Thibault Bañeras Roux, Jane Wottawa, Mickael Rouvier +2ASR EvaluationHuman-in-the-Loop Evaluation