Inter-Annotator Agreement

Momentum

9 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 58

All topics
CardsList
  1. Making COMET Comparable Across Scripts: Diagnosis and Correction of Tokeniser-Induced Script Bias in Indic MT Evaluation

    Oct 6, 2026G. L. John Salvin, Swapnil HingmireMachine Translation QualityUnicode U+2014

  2. COMPASS 2.0: psychometric representational similarity analysis distinguishes symptom structure from personal signal

    Oct 5, 2026Baihan LinPsychometric PropertiesInter-Annotator Agreement

  3. LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction

    Oct 1, 2026Nouha Hayouni, Sheeba Samuel, Alsayed AlgergawyOntologySemantic Relationships

  4. Reproducibility is not construct validity: LLM measurement of institutionally situated communication

    Sep 17, 2026Veronika Batzdorfer, Carlo Romano Marcello Alessandro SantagiustinaInter-Annotator AgreementLarge Language Model Evaluation

  5. Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

    Sep 11, 2026Ziyu Zhang, Satoshi NakamuraInter-Annotator AgreementMulti-Dimensional Evaluation

  6. RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching

    Sep 1, 2026Charles Corbière, Léo Machado, Aubin Charley +3Radiology Report GenerationInter-Annotator Agreement

  7. Augmenting Interviewer Judgments of Patient Experience with Automatic Language Analysis

    Aug 31, 2026Aowen Shi, Michal Balazia, Danilo Postin +4InterviewsClinician Trust

  8. DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

    Aug 31, 2026Yuyang Hong, Jinhui Guo, Jiaqi Gu +6Vision-Language AlignmentInstruction

  9. LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling

    Aug 4, 2026Abhishek Moturu, Babak Taati, Anna GoldenbergNoisy LabelsMedical Imaging Datasets

  10. Consensus Measures for Unstructured Biomedical Text Annotations

    Aug 4, 2026Pascal Wullschleger, Christian Kreis, Martin A. Walter +2Inter-Annotator AgreementMedical Ontology

  11. Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

    Jul 30, 2026Alex Liu, Lief Esbenshade, Michael Xiao +4Llm-As-A-JudgeQualitative Research

  12. SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

    Jul 29, 2026Chuanzhi Xu, Zihan Deng, Huiqi Liang +4Scientific FigurePeer Review

  13. Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs

    Jul 21, 2026Harry Rogers, Sally Shiels, Ashley Tomlinson +5Reference-Guided Multimodal In-Context VerificationClinician Trust

  14. Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations

    Jul 13, 2026Samer Saab, Chaouki AbdallahMulti-Agent Large Language Model SystemsInter-Annotator Agreement

  15. JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes

    Jul 13, 2026Iman Johary, Guillaume Bied, Alexandru C. Mara +1JobsContinuation

  16. Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

    Jun 30, 2026Johannes Hentschel, Emmanouil Karystinaios, Gerhard Widmer +1HarmonizationInter-Annotator Agreement

  17. Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support

    Jun 29, 2026Mizanur Rahman, Abeer Badawi, Elahe Rahimi +4Mental HealthJudges

  18. Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

    Jun 26, 2026Manuel PitaCorrectnessInter-Annotator Agreement

  19. Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication

    Jun 25, 2026Jerome Marston, Tino Kreutzer, Salomé Garnier +3Humanitarian ReportsInter-Annotator Agreement

  20. Interpretable Probabilistic Medical Image Segmentation via Gaussian Process with Explicit Modelling of Annotation Bias and Variability

    Jun 22, 2026Qi Li, Yuliang Huang, Shaheer U. Saeed +7Semi-Supervised Medical Image SegmentationInter-Annotator Agreement

  21. CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System

    Jun 17, 2026Marco Becattini, Niccolò Caselli, Matteo Minin +2Code QualityFeedback

  22. A Two-Stage Statistical Framework for Evaluating Associative Interference in Large Language Models

    Jun 12, 2026Achraf Cohen, Andrew KincaidLarge Language Model BiasInter-Annotator Agreement

  23. Creative Integration: A Decidable Criterion of Creativity

    Jun 11, 2026Yoshinori NomuraCreativityIntelligence

  24. YTClickbait21K: Human-Annotated Multimodal Dataset for YouTube Clickbait Detection Across Diverse Channels and Content Categories

    Jun 10, 2026Md. Minhazul Islam, Md. Tanbeer Jubaer, Amith Khandakar +5Multimodal DatasetContent Moderation

  25. Contemporary AI lacks the imagination to diverge or negate in science

    Jun 6, 2026Honglin Bao, Siyang Wu, Xiao Liu +3Scientific DiscoveryArtificial Intelligence Evaluation

  26. Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

    Jun 5, 2026Kaouther Mouheb, Amos Pomp, Antoine Manenti +9Radiology Report GenerationMedical Vision-Language Models

  27. RadSEM: A Finding-by-Finding Metric for Clinical Consistency in Radiology Reports

    Jun 3, 2026Zhenhong Yang, Zhuoyun Liu, Jintao Fei +4Radiology Report GenerationReport

  28. ACAT: A Collaborative Platform for Efficient Aspect-Based Sentiment Dataset Annotation

    Jun 2, 2026Ana-Maria Luisa Mocanu, Ciprian-Octavian Truica, Elena-Simona ApostolSentimentInter-Annotator Agreement

  29. PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges

    May 29, 2026Swastik Roy, Rajkumar Pujari, Tharindu Kumarage +5Llm-As-A-JudgeJudges

  30. Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora

    May 26, 2026Idris Abdulmumin, Mokgadi Penelope Matloga, Tadesse Destaw Belay +5SentimentInter-Annotator Agreement

  31. A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

    May 25, 2026Pawitsapak Akarajaradwong, Wuttikrai Lertprasertphakorn, Chompakorn Chaksangchaichot +1Llm-As-A-JudgeLarge Language Model Judges

  32. Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse

    May 21, 2026Aisha Ali Al-Athba, Wajdi ZaghouaniArabicComputational Social Science

  33. CompactQE: Interpretable Translation Quality Estimation via Small Open-Weight LLMs

    May 15, 2026Kamil Guttmann, Zofia Fraś, Artur Nowakowski +1Machine Translation QualityMachine Translation

  34. Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation

    May 13, 2026Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait +4Inter-Annotator Agreement

  35. Nürnberg NLP at PsyDefDetect: Multi-Axis Voter Ensembles for Psychological Defence Mechanism Classification

    May 8, 2026Philipp Steigerwald, Eric Rudolph, Jens AlbrechtPsychologyNatural Language Processing

  36. Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement

    May 7, 2026Jessica Huynh, Alfredo Gomez, Athiya Deviyani +3RubricsInter-Annotator Agreement

  37. The First Token Knows: Single-Decode Confidence for Hallucination Detection

    May 6, 2026Mina GabrielSemantic EntropySelf-Consistency

  38. Raising the Ceiling: Better Empirical Fixation Densities for Saliency Benchmarking

    May 5, 2026Susmit Agrawal, Jannis Hollman, Matthias KümmererSaliencyDensity Ratio Estimation

  39. Multi-Rater Calibrated Segmentation Models

    May 4, 2026Meritxell Riera-Marín, Javier García López, Júlia Rodríguez-Comas +2Semi-Supervised Medical Image SegmentationInter-Annotator Agreement

  40. Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

    Apr 27, 2026Aaryan Shah, Andrew Hines, Alexia Downs +6Clinician TrustArtificial Intelligence Evaluation

  41. A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection

    Apr 25, 2026Khalid Hasan, Jamil SaquerMental HealthDepression

  42. MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

    Apr 22, 2026Yingyong Hou, Xinyuan Lao, Huimei Wang +10Inter-Annotator AgreementClinician Trust

  43. Sentiment Analysis of German Sign Language Fairy Tales

    Apr 17, 2026Fabrizio Nunnari, Siddhant Jain, Patrick GebhardSentimentSign Language Translation

  44. The Hitchhikers Guide to Rubric Quality Understanding and Enrichment

    Apr 1, 2026Ankit Aich, Zhengyang Qi, Charles Dickens +7Rubric-Based ScoringRubrics