cs.CLSep 24, 2026

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

Authors: Rong Wang, Kun Sun, Yadong Guo

Organizations: University of Tübingen · Tongji University

Abstract

Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that model--human alignment is associated with explicit linguistic cues, while some rapport-building strategies occur more frequently in misaligned cases. In the categorical task, model predictions exhibit systematic neutral compression, characterized by the overproduction of Neutral labels and the underprediction of Impolite labels. This pattern persists when expert consensus is used as the reference on a diagnostic subset. Our findings highlight the need for pragmatic evaluations that go beyond aggregate agreement metrics by examining directional patterns of model--human disagreement across different human references.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus

    Apr 17, 2026Hitesh Mehta, Arjit Saxena, Garima Chhikara +1Large Language Model BiasPrompt Engineering

  2. How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models

    Apr 17, 2026Judith Sieker, Sina ZarrießLarge Language Model JudgesLinguistics

  3. LLMs Can Better Capture Human Judgments--With the Right Prompts

    Jun 10, 2026Danica Dillion, Chen Cecilia Liu, Baihui Wang +5Human JudgmentMoral Reasoning