Annotator Disagreement

Latest papers 56

All topics
CardsList
  1. Language-model ratings of depression reflect the rater more than the patient

    Oct 6, 2026Baihan LinInter-Rater ReliabilityAnnotator Disagreement

  2. QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations

    Sep 28, 2026Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda +4Annotator DisagreementLLM-Assisted Annotation

  3. How Many Humans Are 32 LLM Judges Worth?

    Sep 18, 2026Chao Li, Yingying Yu, Yunfeng LiAnnotator DisagreementLLM-as-a-Judge

  4. Predicting Human Disagreement for Calibrated Dynamic Facial Expression Recognition

    Sep 15, 2026Yiming Wang, Frederick W. B. Li, Jingyun WangAnnotator DisagreementUncertainty Quantification

  5. Uncertainty-Aware Sea-Ice Type Mapping with Multiple Ice Charts

    Sep 8, 2026Samira Alkaee Taleghan, Younghyun Koo, Andrew P. Barrett +1Remote Sensing Image UnderstandingAnnotator Disagreement

  6. Aggregate Disambiguation Systems

    Aug 31, 2026José María Lago, Albert Castellana, Edgars NemšeAnnotator DisagreementConfidence Region Estimation

  7. Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark

    Aug 24, 2026Martin Wessel, Timo Spinde, Jürgen Pfeffer +1Annotator Disagreement

  8. Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

    Aug 4, 2026Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri +2Annotator DisagreementPreference Optimization

  9. Ensemble Diversity Optimization for Subjective Supervision

    Jul 9, 2026Xia Cui, Ziyi Huang, N. R. AbeynayakeAnnotator DisagreementEnsemble Learning

  10. A Hybrid Framework for Song Lyric Annotation Based on Human-LLM Alignment

    Jun 28, 2026Rashini Liyanarachchi, Frank Tran, Md Mahmudul Hasan +2Annotator DisagreementLLM-Assisted Annotation

  11. Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

    Jun 25, 2026Zahra Omidi, John H. L. HansenAnnotator DisagreementLabel Distribution Learning

  12. Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations

    Jun 24, 2026Anna Nedoluzhko, Šárka Zikánová, Jiří Mírovský +2Inter-Rater ReliabilityCoreference Resolution

  13. Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts

    Jun 22, 2026Yi Ren, Lewis Mitchell, Matthew RoughanAnnotator DisagreementText Classification

  14. Quality and Agreement in Multilabel Emotion Annotation: A Case Study and Evaluation Framework

    Jun 19, 2026Emily Öhman, Anna KoufakouSoft-Label LearningAnnotator Disagreement

  15. Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology

    Jun 11, 2026Saba A. Farahani, Elahe Khatibi, Thomas D. Hughes +3Inter-Rater ReliabilityAnnotator Disagreement

  16. A Resource for Enthymeme Detection in Controversial Political Discourse

    Jun 10, 2026Martial Pastor, Nelleke OostdijkAnnotator DisagreementHuman-in-the-Loop Annotation

  17. The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction

    Jun 1, 2026Mirko Lai, Alessandra Urbinati, Simona Frenda +2Annotator DisagreementConformal Prediction

  18. Bayesian Spectral Emotion Transition Discovery from Multi-Annotator Disagreement

    Jun 1, 2026Keito Inoshita, Takato UenoAnnotator DisagreementEmotion Recognition in Conversations

  19. Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection

    May 29, 2026Benedetta Muscato, Beiduo Chen, Gizem Gezici +2Annotator DisagreementHate Speech Detection

  20. When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis

    May 27, 2026Aisha Najera, Alvin Moon, Vedant Srinivasan +1LLM EvaluationAnnotator Disagreement

  21. Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

    May 26, 2026Idris Abdulmumin, Mokgadi Penelope Matloga, Tadesse Destaw Belay +5Annotator DisagreementSentiment Analysis

  22. A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

    May 25, 2026Pawitsapak Akarajaradwong, Wuttikrai Lertprasertphakorn, Chompakorn Chaksangchaichot +1Inter-Rater ReliabilityAnnotator Disagreement

  23. Uncertainty Decomposition via Cyclical SG-MCMC and Soft-label Learning for Subjective NLP

    May 23, 2026Keito Inoshita, Takato UenoSoft-Label LearningBayesian Neural Networks

  24. Calibrating Probabilistic Object Detectors with Annotator Disagreement

    May 23, 2026Zhi Qin Tan, Owen Addison, Yunpeng LiAnnotator DisagreementObject Detection

  25. Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

    May 11, 2026Junyu Lu, Deyi Ji, Xuanyi Liu +5Annotator DisagreementLLM Alignment

  26. Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks

    May 11, 2026Tadesse Destaw Belay, Ibrahim Said Ahmad, Idris Abdulmumin +6Annotator DisagreementClustering

  27. Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation

    May 7, 2026Hakyung Sung, Gyu-Ho ShinAnnotator DisagreementMorphosyntactic Tagging

  28. Who and What? Using Linguistic Features and Annotator Characteristics to Analyze Annotation Variation

    May 7, 2026Maximilian Maurer, Maximilian Linde, Gabriella LapesaAnnotator DisagreementHuman-in-the-Loop Annotation

  29. Understanding Annotator Safety Policy with Interpretability

    May 6, 2026Alex Oesterling, Donghao Ren, Yannick Assogba +4Annotator DisagreementHuman-in-the-Loop Annotation

  30. Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Model

    May 5, 2026Niklas Donhauser, Jakob Fehle, Nils Constantin Hellwig +3Annotator DisagreementLLM-Assisted Annotation

  31. STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems

    May 4, 2026Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6Annotator DisagreementAI Agent Evaluation

  32. Quantifying and Predicting Disagreement in Graded Human Ratings

    May 1, 2026Leixin Zhang, Çağrı ÇöltekinAnnotator Disagreement

  33. Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework

    Apr 30, 2026Sheza Munir, Ahanaf Rodoshi, Sumin Lee +3Annotator DisagreementNeuro-Symbolic Reasoning

  34. LLMs Capture Emotion Labels, Not Emotion Uncertainty: Distributional Analysis and Calibration of Human-LLM Judgment Gaps

    Apr 30, 2026Keito Inoshita, Xiaokang Zhou, Akira Kawai +1LLM EvaluationAnnotator Disagreement

  35. Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement

    Apr 24, 2026Wataru Hirota, Tomoki Taniguchi, Tomoko Ohkuma +6Human Preference EvaluationAnnotator Disagreement

  36. Fine-Grained Perspectives: Modeling Explanations with Annotator-Specific Rationales

    Apr 23, 2026Olufunke O. Sarumi, Charles Welch, Daniel BraunAnnotator DisagreementNatural Language Inference

  37. Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

    Apr 19, 2026Finn Schmidt, Jan Philip Wahle, Terry Ruas +1Annotator DisagreementAutomated Evaluation

  38. Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks

    Apr 18, 2026Nisrine Rair, Alban Goupil, Valeriu Vrabie +1Annotator Disagreement

  39. IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language

    Apr 17, 2026Christina Chance, Rebecca Pattichis, Arjun Subramonian +4Social Media AnalysisAnnotator Disagreement

  40. Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions

    Mar 24, 2026Avrile Floro, Tamara Dhorasoo, Soline Pellez +1Annotator DisagreementLegal Citation Verification

  41. Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks

    Mar 22, 2026Navya Mehrotra, Adam Visokay, Kristina GligorićAnnotator DisagreementLLM-Assisted Annotation

  42. Labels have Human Values: Value Calibration of Subjective Tasks

    Jan 10, 2026Mohammed Fayiz Parappan, Ricardo HenaoAnnotator DisagreementAI Alignment