Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
Organizations: Old Dominion University, Norfolk, Virginia, USA
Abstract
Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.
Figures & tables
| Pos/Neg/Neutral | Neg/Non-Neg | Pos/Non-Pos | |
| All Humans (R1-R6) | 0.39 | 0.45 | 0.48 |
| All Tools (TB, VD, RB) | 0.43 | 0.49 | 0.43 |
| Humans (R1-R6) with TB | 0.35 | 0.41 | 0.39 |
| Humans (R1-R6) with VD | 0.35 | 0.42 | 0.42 |
| Humans (R1-R6) with RB | 0.43 | 0.51 | 0.50 |
| Humans (R1-R6) with TB, VD, RB | 0.38 | 0.45 | 0.40 |
| Pos/Neg/Neutral | Neg/Non-Neg | Pos/Non-Pos | |
| Majority vs. TB | 0.37 | 0.41 | 0.29 |
| Majority vs. VD | 0.40 | 0.50 | 0.37 |
| Majority vs. RB | 0.73 | 0.88 | 0.61 |
| Majority vs. LLaMA (one-shot) | 0.63 | 0.71 | 0.71 |
| Majority vs. Qwen (one-shot) | 0.48 | 0.60 | 0.61 |
| Majority vs. GPT (one-shot) | 0.57 | 0.67 | 0.67 |