Pairwise Preference Evaluation

Momentum

2 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 69

All topics
CardsList
  1. Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons

    Aug 10, 2026Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal +1Pairwise Preference LearningReward Modeling

  2. TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

    Aug 9, 2026Yidong Wang, Yan Zhan, Ziteng Feng +16Reward ModelingPairwise Preference Evaluation

  3. Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

    Aug 7, 2026Hongyu Luo, He Wang, Huihao Jing +6Creativity AssessmentLLM Evaluation

  4. Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process

    Aug 6, 2026Vitaliy Tsyganok, Sergii Kadenko, Oleh AndriichukPairwise ComparisonPairwise Preference Evaluation

  5. Isotonic Bradley-Terry Model for Paired Comparison Data

    Aug 3, 2026Ryoya YamasakiPairwise Preference LearningPairwise Comparison

  6. AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization

    Jul 30, 2026Shengda Gu, Kai Li, Xinyi Ke +3Automated Algorithm DiscoveryPairwise Preference Evaluation

  7. What do Reward Models Memorize?

    Jul 27, 2026Ivo Verhoeven, Pushkar Mishra, Ekaterina ShutovaReward ModelingPairwise Preference Evaluation

  8. CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

    Jul 23, 2026Hehao Zhang, Danli Wang, Xinyuan Wang +1Pairwise Preference LearningReward Modeling

  9. Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

    Jul 20, 2026Jiabing Yang, Yixiang Chen, Yuan Xu +6Human Preference EvaluationPairwise Preference Evaluation

  10. Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

    Jul 13, 2026Simone Drago, Marco Mussi, Leonardo Bianconi +1Pairwise Preference LearningPairwise Comparison

  11. Geometric mean-based pairwise comparison method with the reference values -- statistical approach

    Jul 10, 2026Konrad Kułakowski, Jacek SzybowskiPairwise ComparisonPairwise Preference Evaluation

  12. Optimal Top-kk Identification from Pairwise Comparisons

    Jul 9, 2026Motti Goldberger, Nils RudiPairwise Preference LearningMulti-Armed Bandits

  13. Attention Limited Reward Learning

    Jul 6, 2026Wenqian XingPairwise Preference LearningReward Modeling

  14. LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

    Jul 1, 2026Ruotong Zhao, Zhiyu Chen, Xurui Liu +7Human Preference EvaluationLLM-as-a-Judge

  15. Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

    Jun 30, 2026Zewen LiuLLM-as-a-JudgePairwise Preference Evaluation

  16. Can LLMs Rank? A Tale of Triads and Triage

    Jun 29, 2026Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler +1Pairwise ComparisonSocial Choice Theory

  17. ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

    Jun 23, 2026Jisu Jeon, Seungyeon Jwa, Joosung Lee +6Audio-Language Model EvaluationPairwise Comparison

  18. A Markov Chain Approach to Preference Alignment

    Jun 21, 2026Takuya Koriyama, Tengyuan LiangPairwise Preference LearningPairwise Comparison

  19. Which Pairs to Compare for LLM Post-Training?

    Jun 17, 2026Jiangze Han, Vineet Goyal, Will MaPairwise Preference LearningPairwise Comparison

  20. PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets

    Jun 17, 2026Junyi Fan, Donald S. WilliamsonPairwise Preference LearningPairwise Preference Evaluation

  21. UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning

    Jun 17, 2026Mohamed Nabail, Leo Kaixuan Cheng, Jingmin Wang +1Pairwise Preference EvaluationRL Exploration

  22. RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing

    Jun 17, 2026Guannan Lai, Haoran Hu, Han-Jia YeLLM EvaluationPairwise Preference Evaluation

  23. Prompt Perturbation for Reliable LLM Evaluation over Comparison Graphs

    Jun 16, 2026Dong Huang, Jianbo Sun, Pengkun YangLLM EvaluationPairwise Comparison

  24. TuneJury: An Open Metric for Improving Music Generation Preference Alignment

    Jun 15, 2026Yonghyun Kim, Junwon Lee, Haiwen Xia +5Text-to-Music GenerationMusic Generation Evaluation

  25. Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling

    Jun 14, 2026Yujin Park, Haejun Chung, Ikbeom JangPairwise Preference LearningPairwise Comparison

  26. Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

    Jun 8, 2026Mina Remeli, Moritz HardtPairwise ComparisonLLM-as-a-Judge

  27. DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity

    Jun 8, 2026Fengyuan Liu, Yongliang Miao, Zirui He +3Pairwise Preference LearningReward Modeling

  28. Local Preferential Bayesian Optimization

    Jun 1, 2026Johanna Menn, Miriam Kober, Paul Brunzema +2Pairwise Preference LearningPairwise Preference Evaluation

  29. From Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement Learning

    May 31, 2026Jun-Jie Yang, Chia-Heng Hsu, Kui-Yuan Chen +1Pairwise Preference LearningRepresentation Learning

  30. Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

    May 29, 2026Jonathan Colaço Carr, Prakash Panangaden, Doina Precup +1Pairwise Preference LearningMarkov Decision Processes

  31. The Representation-Rationalizability Tradeoff in Reward Learning

    May 29, 2026Jing Dong, Yaoliang Yu, Pascal PourpartPairwise Preference LearningReward Modeling

  32. Reward Learning from Best-of-NN Preference Data: Targets, Tradeoffs, and Design Principles

    May 28, 2026Rattana Pukdee, Maria-Florina Balcan, Pradeep RavikumarPairwise Preference LearningReward Modeling

  33. Resolution Diagnostics for Paired LLM Evaluation

    May 28, 2026Anany KotawalaLLM EvaluationPairwise Comparison

  34. From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals

    May 28, 2026Yeyong Yu, Wenya Hu, Xing Wu +1LLM EvaluationPairwise Preference Evaluation

  35. Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

    May 25, 2026Fay Elhassan, David Sasu, Alexandra Kulinkina +2LLM Safety BenchmarksHuman Preference Evaluation

  36. SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

    May 25, 2026Yanhang Li, Zhichao Fan, Zexin ZhuangPairwise ComparisonPairwise Preference Evaluation

  37. JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

    May 24, 2026Russell Yang, Ruishi Chen, Pierce Kelaita +6Human Preference EvaluationLLM Evaluation

  38. RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

    May 20, 2026Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3LLM EvaluationPairwise Comparison

  39. Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation

    May 19, 2026Yuanpei Zhao, Jie Lin, Chao Zhang +5Pairwise Preference LearningPersonalized Image Aesthetics Assessment

  40. Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation

    May 18, 2026Guining Cao, Jiaxin Peng, Chu Zeng +3Pairwise Preference LearningOpen-Ended Generation

  41. CAPS: Cascaded Adaptive Pairwise Selection for Efficient Parallel Reasoning

    May 15, 2026Fangzhou Lin, Shuo Xing, Peiran Li +6Pairwise ComparisonPairwise Preference Evaluation

  42. Active Learners as Efficient PRP Rerankers

    May 14, 2026Jeremías Figueiredo Paschmann, Juan Kaplan, Francisco Nattero +3Pairwise Preference LearningPairwise Comparison

  43. TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

    May 12, 2026Truong Nguyen, Tien-Phat Nguyen, Linh Ngo Van +3Pairwise Preference LearningLLM Alignment

  44. Variance-aware Reward Modeling with Anchor Guidance

    May 12, 2026Shuxing Fang, Ruijian Han, Liangyu Zhang +1Pairwise Preference LearningReward Modeling

  45. Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

    May 10, 2026Yanran LiLLM EvaluationLLM-as-a-Judge

  46. Efficient Ensemble Selection from Binary and Pairwise Feedback

    May 10, 2026Tzeh Yuan Neoh, Nicholas Teh, Je Qin Chooi +2Pairwise ComparisonPairwise Preference Evaluation

  47. Sufficient conditions for a Heuristic Rating Estimation Method application

    May 9, 2026Jacek Szybowski, Konrad Kułakowski, Jiri MazurekPairwise ComparisonPairwise Preference Evaluation

  48. Response Time Enhances Alignment with Heterogeneous Preferences

    May 7, 2026Federico Echenique, Alireza Fallah, Baihe Huang +1Pairwise Preference LearningLLM Alignment

  49. Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring

    May 3, 2026İbrahim Rıza Hallaç, Hasan OğulPairwise Preference LearningPairwise Comparison

  50. Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare

    May 3, 2026Maheed H. Ahmed, Mahsa GhasemiPairwise Preference LearningMulti-Armed Bandits