Human Preference Evaluation

Latest papers 41

All topics
CardsList
  1. How Reliable Are Predicted MOS for Reproducing Human System-Level Preferences in Speech Enhancement?

    Sep 30, 2026Nahomi Kusunoki, Tsubasa Ochiai, Naohiro Tawara +4Human Preference EvaluationSpeech Quality Assessment

  2. JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

    Sep 29, 2026Qi Cao, Kangning Liu, Xuan Kan +10Human Preference EvaluationLLM Evaluation

  3. Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications

    Sep 27, 2026Xianglong Shi, Shifeng Liu, Sirui Zhao +2Human Preference EvaluationLearning to Rank

  4. Is Semantics Enough for Speech Mean Opinion Score Prediction?

    Sep 3, 2026Tianyu Lan, Yufei Shi, Yang Ai +3Human Preference EvaluationNeural Audio Codecs

  5. ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

    Aug 13, 2026Jiale Cui, Yueyao Yuan, Kaixi Zhong +3Human Preference EvaluationAI Agent Benchmarks

  6. How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans

    Aug 10, 2026Hasan Mahmud, Khawaja Abaid Ullah, Mohammad Javad Khojasteh +2Human Preference EvaluationGender Bias in Language Models

  7. Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

    Jul 30, 2026Zheng Wu, Yibo Luo, Pu Zhang +2Human Preference EvaluationLLM-as-a-Judge

  8. Human Preference aligned Tabular Similarity

    Jul 27, 2026Frederik Hoppe, Astrid Franz, Marianne Michaelis +2Human Preference EvaluationTabular Representation Learning

  9. Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

    Jul 20, 2026Jiabing Yang, Yixiang Chen, Yuan Xu +6Human Preference EvaluationPairwise Preference Evaluation

  10. Human-in-the-Loop User Feedback Affects Perceived Accuracy and Trust, but Task Subjectivity Matters

    Jul 20, 2026Donald R. Honeycutt, Mahsan Nourani, Eric D. RaganHuman Preference EvaluationHuman-in-the-Loop Evaluation

  11. Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation

    Jul 15, 2026Yizhou Zhang, Wangjin Zhou, Yi Zhao +3Human Preference EvaluationAlgorithmic Bias

  12. Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

    Jul 6, 2026Robert Morabito, Tyler McDonald, Charitra Viswanath +4Human Preference EvaluationLLM Evaluation

  13. LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

    Jul 1, 2026Ruotong Zhao, Zhiyu Chen, Xurui Liu +7Human Preference EvaluationLLM-as-a-Judge

  14. The Human Creativity Benchmark

    Jun 29, 2026Aspen Hopkins, Allison Nulty, Alexandria Minetti +2Human Preference EvaluationCreativity Assessment

  15. Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

    Jun 18, 2026Masato Takagi, Masaya Kawamura, Reo Shimizu +1Human Preference EvaluationSpeech Generation Evaluation

  16. Humor Style Drives Laughter, Topic Shapes Acceptability: Evaluating Bilingual Personal and Political Robot-Delivered AI Jokes

    Jun 11, 2026Anna-Maria Velentza, Anne-Gwenn BosserHuman Preference EvaluationHuman-Robot Interaction

  17. LLMs Can Better Capture Human Judgments--With the Right Prompts

    Jun 10, 2026Danica Dillion, Chen Cecilia Liu, Baihui Wang +5Human Preference EvaluationLLM Alignment

  18. Re-Centering Humans in LLM Personalization

    Jun 4, 2026Lechen Zhang, Jiarui Liu, Tal AugustHuman Preference EvaluationLLM Evaluation

  19. A Dataset for Dynamic Human Preferences for Vision Language Models

    Jun 2, 2026Hannah Gao, Dylan Hadfield-Menell, Rachel MaVLM EvaluationHuman Preference Evaluation

  20. Personalized Turn-Level User Conversation Satisfaction Benchmark

    May 28, 2026Zhefan Wang, Zhiqiang Guo, Weizhi Ma +3Human Preference EvaluationLLM Personalization

  21. Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

    May 25, 2026Fay Elhassan, David Sasu, Alexandra Kulinkina +2LLM Safety BenchmarksHuman Preference Evaluation

  22. JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

    May 24, 2026Russell Yang, Ruishi Chen, Pierce Kelaita +6Human Preference EvaluationLLM Evaluation