cs.CLMar 22, 2026

Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks

Authors: Navya Mehrotra, Adam Visokay, Kristina Gligorić

Organizations: University of Washington

Abstract

Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others. Existing methods for correcting LLM annotation error assume a single ground truth. However, this assumption fails in subjective tasks where disagreement across demographic groups is meaningful. Here we introduce Perspective-Driven Inference, a method that treats the distribution of annotations across groups as the quantity of interest, and estimates it using a small human annotation budget. We contribute an adaptive sampling strategy that concentrates human annotation effort on groups where LLM proxies are least accurate. We evaluate on politeness and offensiveness rating tasks, showing targeted improvements for harder-to-model demographic groups relative to uniform sampling baselines, while maintaining coverage.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?

    Apr 20, 2026Hasan Amin, Harry Yizhou Tian, Xiaoni Duan +3Human AnnotatorsBiases

  2. LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts

    Aug 31, 2026Daniela Occhipinti, Andrea Piergentili, Marco GueriniDemographicsLarge Language Model Bias

  3. Which Demographics do LLMs Default to During Annotation?

    Oct 11, 2024Johannes Schäfer, Aidan Combs, Christopher Bagdon +9Large Language Model AnnotationsLarge Language Model Bias