Inter-Rater Reliability

Latest papers 31

All topics
CardsList
  1. Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

    Oct 7, 2026Himarsha R. Jayanetti, Sivakanesan Dhanushkanda, Shuai Hao +2Sentiment AnalysisSocial Media Analysis

  2. Language-model ratings of depression reflect the rater more than the patient

    Oct 6, 2026Baihan LinInter-Rater ReliabilityAnnotator Disagreement

  3. LLM Judge Validation Under Sparse Overlap: From Inference to Design

    Sep 25, 2026Junxuan Li, Arko Mukherjee, Soumyabrata PalInter-Rater ReliabilityLLM Evaluation

  4. Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons

    Aug 10, 2026Kaustubh Shivshankar Shejole, Tanish Agarwal, Arpit Agarwal +1Pairwise Preference LearningReward Modeling

  5. Consensus Measures for Unstructured Biomedical Text Annotations

    Aug 4, 2026Pascal Wullschleger, Christian Kreis, Martin A. Walter +2Inter-Rater ReliabilitySemantic Textual Similarity

  6. AHEAD: Advancing Multi-Class Label Aggregation with Interpretable Cross-Annotator Modeling

    Jul 20, 2026Ju Chen, Sijia Xu, Jun Feng +2Inter-Rater ReliabilityNoisy-Label Learning

  7. A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

    Jul 8, 2026A. Sayyad, J. Emmons, S. Jones +2Audio-Language Model EvaluationVoice Agent Evaluation

  8. Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations

    Jun 24, 2026Anna Nedoluzhko, Šárka Zikánová, Jiří Mírovský +2Inter-Rater ReliabilityCoreference Resolution

  9. Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias

    Jun 17, 2026Justin D. Norman, Michael U. Rivera, D. Alex HughesInter-Rater ReliabilityLLM Evaluation

  10. Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation

    Jun 15, 2026Truong Vu, Minh Khoi Ho, Yutong XiePrototype-Based LearningInter-Rater Reliability

  11. Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology

    Jun 11, 2026Saba A. Farahani, Elahe Khatibi, Thomas D. Hughes +3Inter-Rater ReliabilityAnnotator Disagreement

  12. Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated

    May 30, 2026Rashid MushkaniVLM EvaluationInter-Rater Reliability

  13. A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

    May 25, 2026Pawitsapak Akarajaradwong, Wuttikrai Lertprasertphakorn, Chompakorn Chaksangchaichot +1Inter-Rater ReliabilityAnnotator Disagreement

  14. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    May 25, 2026Delip Rao, Chris Callison-BurchInter-Rater ReliabilityLLM Evaluation

  15. Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

    May 13, 2026Deepak Pandita, Flip Korn, Chris Welty +1Inter-Rater ReliabilityLLM Evaluation

  16. Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

    May 7, 2026Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3Inter-Rater ReliabilityLLM-as-a-Judge

  17. Multi-Rater Calibrated Segmentation Models

    May 4, 2026Meritxell Riera-Marín, Javier García López, Júlia Rodríguez-Comas +2Inter-Rater ReliabilityImage Segmentation

  18. Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

    Apr 30, 2026M. Riera-Marín, O. K. Sikha, J. Rodríguez-Comas +23Inter-Rater ReliabilityMedical Image Analysis

  19. Honest and Reliable Evaluation and Expert Equivalence Testing of Automated Neonatal Seizure Detection

    Aug 6, 2025Jovana Kljajic, John M. O'Toole, Robert Hogan +1Inter-Rater ReliabilityAutomated Evaluation

  20. Hierarchical Bayesian Crowdsourcing with Item Difficulty

    May 29, 2024Seong Woo Han, Ozan Adıgüzel, Bob CarpenterInter-Rater ReliabilityLatent Variable Models

  21. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

    Date pendingWilliam CabanInter-Rater ReliabilityAI Agent Evaluation