Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks
Organizations: University of Washington
Abstract
Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others. Existing methods for correcting LLM annotation error assume a single ground truth. However, this assumption fails in subjective tasks where disagreement across demographic groups is meaningful. Here we introduce Perspective-Driven Inference, a method that treats the distribution of annotations across groups as the quantity of interest, and estimates it using a small human annotation budget. We contribute an adaptive sampling strategy that concentrates human annotation effort on groups where LLM proxies are least accurate. We evaluate on politeness and offensiveness rating tasks, showing targeted improvements for harder-to-model demographic groups relative to uniform sampling baselines, while maintaining coverage.
Figures & tables
| Avg (Age) | Age 18–34 | Age 35–49 | Age 50+ | |||||
|---|---|---|---|---|---|---|---|---|
| Method | Cov. | Cov. | Cov. | Cov. | ||||
| GPT-5.2 (zero shot) | 88.3 | 7.48 | 95.0 | 4.10 | 80.0 | 6.88 | 90.0 | 11.46 |
| GPT-5.2 (few shot) | 91.7 | 7.52 | 100.0 | 2.45 | 100.0 | 3.81 | 75.0 | 16.31 |
| GPT-5.2 (persona) | 93.3 | 7.02 | 100.0 | 3.58 | 95.0 | 4.60 | 85.0 | 12.87 |
| PPI | 95.0 | 6.35 | 100.0 | 2.93 | 100.0 | 2.49 | 85.0 | 13.63 |
| PDI | 93.3 | 7.10 | 90.0 | 5.51 | 95.0 | 4.58 | 95.0 | 11.23 |
| Group | PDI samples | Uniform samples |
|---|---|---|
| Gender | ||
| Woman | 78.7 | 85.1 |
| Man | 121.3 | 114.9 |
| Race | ||
| White | 135.8 | 140.2 |
| Black/Afr. Am. | 34.3 | 28.8 |
| Avg (Age) | Age 18–34 | Age 35–49 | Age 50+ | |||||
|---|---|---|---|---|---|---|---|---|
| Method | Cov. | Cov. | Cov. | Cov. | ||||
| GPT-5.2 (zero shot) | 11.7 | 20.11 | 15.0 | 14.05 | 5.0 | 18.38 | 15.0 | 27.90 |
| GPT-5.2 (few shot) | 53.3 | 14.21 | 65.0 | 8.01 | 60.0 | 10.06 | 35.0 | 24.54 |
| GPT-5.2 (persona) | 13.3 | 19.68 | 25.0 | 13.12 | 0.0 | 17.50 | 15.0 | 28.40 |
| PPI | 95.0 | 3.86 | 100.0 | 2.38 | 95.0 | 3.57 | 90.0 | 5.64 |
| PDI | 95.0 | 6.04 | 95.0 | 7.31 | 95.0 | 5.57 | 95.0 | 5.24 |
| Group | PDI samples | PPI samples |
|---|---|---|
| Gender | ||
| Woman | 152.3 | 143.7 |
| Man | 47.6 | 56.4 |
| Race | ||
| White | 164.7 | 162.9 |
| Black/Afr. Am. | 6.8 | 9.0 |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Prompting | Politeness | Offensiveness | ||
|---|---|---|---|---|---|
| Cov. | Cov. | ||||
| gpt-5.2 | zero-shot | 0.36 | 6.5 | 0.09 | 18.7 |
| few-shot | 0.64 | 6.6 | 0.09 | 13.5 | |
| persona | 0.55 | 6.4 | 0.09 | 18.4 | |
| claude-sonnet-4.6 | zero-shot | 0.09 | 12.0 | 0.00 | 37.7 |
| few-shot | 0.27 | 10.5 | 0.64 | 6.0 | |
| Prompt Type | Prompt Text |
|---|---|
| Politeness Zero Shot | Consider you read this email from a colleague. How polite is it? Answer using a 1–5 Likert scale where 1 means ‘not polite at all’ and 5 means ‘very polite’. Return ONLY one integer. |
| Offensiveness Zero Shot | Consider you read this comment on Reddit. How offensive is it? Answer using a 1–5 Likert scale where 1 means ‘not offensive at all’ and 5 means ‘very offensive’. Return ONLY one integer. |
| Politeness Few Shot | Here are example politeness ratings (1=very impolite, 5=very polite): {Followed by Text: and Rating: for 10 examples. Rest of the prompt is the same as zero shot.} |
| Offensiveness Few Shot | Here are example offensiveness ratings (1=not offensive, 5=extremely offensive): {Followed by Text: and Rating: for 10 examples. Rest of the prompt is the same as zero shot.} |
| Politeness Persona Prompt | You are {persona}. Answer all questions from the perspective of {persona}. Be consistent with this perspective, but still follow the instructions. {followed by same prompt as Politeness Zero Shot}. persona = a gender aged age with education |
| Offensiveness Persona Prompt | You are {persona}. Answer all questions from the perspective of {persona}. Be consistent with this perspective, but still follow the instructions. {followed by same prompt as Offensiveness Zero Shot}. |
| Offensiveness (1 = not offensive, 5 = very offensive) | ||
| Text | Age | Rating |
| “Women’s studies is dumb though, stop defending your degree?” | 18–24 | 5 |
| 25–29 | 5 | |
| 30–34 | 3 | |
| 40–44 | 5 | |
| 45–49 | 1 | |
| Politeness (1 = very impolite, 5 = very polite) | ||
| Text | Age | Rating |
| “No, tell her that it is off of Buffalo Speedway.” | 18-24 | 1 |
| 25-29 | 3 | |
| 30-34 | 1 | |
| 30-34 | 5 | |
| 50-54 | 4 | |
| Concept | Prevalence | Match Count | Meaning | Example |
|---|---|---|---|---|
| Request for Information | 0.26 | 64 | Does the text include a request for information or clarification? | “I HAVE NOT RECEIVED YOUR DATA FORMS… IF YOU HAVE ANY QUESTIONS, PLEASE FEEL FREE TO CALL ME. Thank you” |
| Hopeful Sentiment | 0.18 | 45 | Does the text express hope or a positive outlook? | “Look forward to seeing you tomorrow.” |
| Polite Closing | 0.09 | 23 | Does the text include a polite closing or sign-off, such as ’best regards’? | “I will leave the legal review up to your team. Please let me know if you have any questions about these CPs. Regards” |
| Attachment Mention | 0.07 | 17 | Does the text mention an attachment or include the word ’attached’? | “I am send you the following attached documents per request.” |
| Expression of Gratitude | 0.04 | 9 | Does the text express gratitude or thanks? | “I would appreciate if it you didn’t blame me for the spills in your cube” |
| Concept | Prevalence | Match Count | Meaning | Example |
|---|---|---|---|---|
| Request for Action | 0.28 | 8 | Does the text contain a request for action or follow-up, indicated by words like ’please’, ’let’, or ’thanks’? | “Please find attached the Summary and Hot List as of 11/16. Please contact me if you have any questions/comments. Thanks” |
| Emotional Sharing | 0.21 | 6 | Does the text involve sharing emotions or personal connections? | “Are you in love with him?” |
| Evaluation and Quality | 0.14 | 4 | Does the text involve evaluating the quality or value of something? | “Let me tell you who won’t win - RU. They are arguably the worst division 1 football team I’ve ever seen. I saw them for the 1st time this weekend on TV. They are horrible. Hope everything’s going well. Life at Enron sucks. I’m sure you’ve been following the situation. I’m sure there’ll be a #1 bestseller written about this circus. Talk to you soon. KR” |
| Motivation and Support | 0.10 | 3 | Does the text discuss the importance of encouragement or support? | “Congrats on keeping your turf. I hope this question doesn’t offend you, but would you trade it for an $80 ENE price? Thanks.” |
| Travel Plans | 0.07 | 2 | Does the text mention any travel plans or locations? | “Hey give me a call! You should come to NYC this weekend to party with your friends!” |
| Human label | LLM error (zero shot) | LLM error (few shot) | LLM error (persona) | |||||||
| Dataset | Dimension | Range | Range | Range | Range | |||||
| MultiPICo | Gender | 2 | 0.30–0.34 | 0.272 | 0.29–0.32 | 0.299 | 0.30–0.33 | 0.391 | 0.29–0.32 | 0.318 |
| Age | 3 | 0.29–0.36 | 0.249 | 0.30–0.35 | 0.398 | 0.28–0.39 | 0.044 | 0.27–0.35 | 0.180 | |
| Ethnicity | 4 | 0.25–0.39 | 0.135 | 0.26–0.35 | 0.546 | 0.31–0.36 | 0.752 | 0.25–0.36 | 0.223 | |
| Nationality | 5 | 0.28–0.37 | 0.406 | 0.27–0.36 | 0.352 | 0.28–0.36 | 0.517 | 0.25–0.38 | 0.031 | |
| CSC | Gender | 2 | 0.31–0.35 | 0.140 | 0.28–0.31 | 0.323 | 0.29–0.30 | 0.839 | 0.29–0.35 | 0.083 |
| Avg (Ethnicity) | White | Black | Asian | Mixed | ||||||
| Method | Cov. | Cov. | Cov. | Cov. | Cov. | |||||
| Group size / LLM error (few shot) | ||||||||||
| 999 / 0.29 | 746 / 0.28 | 117 / 0.38 | 70 / 0.30 | 43 / 0.23 | ||||||
| gpt-oss-120b (zero shot) | 76.3 | 11.63 | 35.0 | 8.86 | 95.0 | 8.64 | 80.0 | 14.73 | 95.0 | 14.28 |
| gpt-oss-120b (few shot) | 41.3 | 17.16 | 5.0 | 12.82 | 35.0 | 20.22 | 45.0 | 19.90 | 80.0 | 15.69 |
| gpt-oss-120b (persona) | 66.3 | 13.43 | 20.0 | 11.52 | 90.0 | 8.84 | 75.0 | 13.90 | 80.0 | 19.46 |
| Dimension | Group | LLM err. | PDI | PPI |
|---|---|---|---|---|
| Gender | Woman | 0.30 | 110.8 | 105.2 |
| Man | 0.29 | 89.2 | 94.8 | |
| Age | Age 18–34 | 0.31 | 95.6 | 96.8 |
| Age 35–49 | 0.30 | 75.2 | 69.8 | |
| Age 50+ | 0.24 | 29.2 | 33.5 | |
| Ethnicity | White | 0.28 | 141.5 | 149.9 |