cs.CLSep 12, 2026

Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus

Authors: Levent Bulut

Organizations: Independent Researcher

Abstract

Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_rules field from a rule-based detector over six craft features -- two prohibitions (emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Three studies are reported. Study 1 (n=120n = 120) scores the detector against blind labels from the scheme's own author. Study 2 (n=100n = 100, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5 (High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor -- closest to the methodology's theoretical core -- the five machine labellers returned positive rates of 00, 11, 4040, 7272 and 7878 out of 100100 scenes, against a human count of 99. Cohen's κκ was at or indistinguishable from chance for five of six labellers, across both human references and both scene sets: 0.0040.004, 0.0150.015, 0.0000.000, 0.0190.019, 0.0270.027. Raw agreement ranged from 74.7%74.7\% to 84.5%84.5\%, an artefact of class imbalance rather than a sign of competence. We deliberately do not resolve this into a single story. Two readings survive: the feature is genuinely inferential and beyond current automatic detection, or the rule's definition is not yet operational enough for any rater to apply consistently -- including the human. Distinguishing them needs a second independent human rater, which this report does not have and therefore does not claim.

Explore similar work

CardsList
  1. Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text

    Oct 2, 2026Claudiu Creanga, Liviu DinuAI-Generated Text DetectionRobustness of Machine-Generated Text Detection

  2. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    May 25, 2026Delip Rao, Chris Callison-BurchInter-Rater ReliabilityLLM Evaluation

  3. Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

    Jun 1, 2026Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli +10Human-in-the-Loop Annotation