Groupwise Distortion Guarantees for Preference-Based Alignment
Organizations: University of Pennsylvania
Abstract
Preference-based alignment methods such as reinforcement learning from human feedback (RLHF) and Nash learning from human feedback (NLHF) aggregate pairwise preferences to learn an LLM policy, but a natural goal is maximizing social welfare (average cardinal utility), which comparisons alone need not identify. Gölz, Haghtalab, and Yang (GHY) measure the gap by distortion: the worst-case ratio between the welfare of the best fixed lottery (distribution over responses) and of the learned lottery. They show NLHF is optimal when every user receives the same lottery. Account-based LLMs, however, have information about their users and can serve different lotteries to different people. We give an efficient algorithm, GLHF, that learns a single group-conditioned policy from one comparison per user. Under individual Bradley--Terry comparisons, GLHF asymptotically matches GHY's optimal population distortion bound simultaneously on every group in a prespecified, possibly overlapping collection, with sample complexity growing logarithmically in the number of groups and inversely with the smallest group mass. A sharper guarantee for groups with similar preferences approaches distortion of one when members share a feasible favorite response. In experiments using human coffee ratings and synthetic LLM-generated ratings, GLHF lowers distortion in every evaluated group and substantially reduces worst-group distortion relative to NLHF and other group-agnostic baselines.
Figures & tables
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Definition | |
|---|---|---|
| Universal | 3,369 | Every analysis respondent. |
| Age 45+ | 483 | Reported age 45–54, 55–64, or over 65. |
| Female | 759 | Reported gender is female. |
| Expertise 1–4 | 815 | Self-rated coffee expertise is 1, 2, 3, or 4. |
| Children 1+ | 767 | Reported one, two, three, or more than three children. |
| Works from home | 1,424 | Selected “I primarily work from home.” |
| Group | Definition | |
|---|---|---|
| Universal | 1,497 | Every analysis pseudo-respondent. |
| Not technical | 750 | Persona specifies a nontechnical background. |
| Expressive wording | 747 | Persona prefers expressive writing. |
| Direct wording | 750 | Persona prefers direct writing. |
| Very technical | 747 | Persona specifies a very technical background. |
| University education | 748 | Persona specifies university education. |
| Persona attribute | Discovery | Rating-difference correlation | Selected |
|---|---|---|---|
| Not technical | 82 | yes | |
| Expressive wording | 85 | yes | |
| Direct wording | 82 | yes | |
| Very technical | 85 | yes | |
| High school or less | 83 | no | |
| University education | 84 | yes |
| Method | Deployed output | Comparisons at main setting |
|---|---|---|
| Multigroup GLHF | Lottery depending on group memberships | Active |
| Batch NLHF | One learned lottery for everyone | Passive |
| Pooled GLHF | One learned lottery for everyone | Active |
| Uniform reference | Equal probability for every alternative | None |
| Study | MG | Batch NLHF | Pooled | Ref. | Batch NLHF MG | Pooled MG | Ref. MG |
|---|---|---|---|---|---|---|---|
| Coffee-rating study | 1.0795 | 1.2816 | 1.2528 | 1.1217 | 0.2021 | 0.1732 | 0.0421 |
| Completion-rating study | 1.0695 | 1.3646 | 1.3501 | 1.7357 | 0.2951 | 0.2807 | 0.6662 |
| Feature | First level | Second level |
|---|---|---|
| Age | 22-32 | 62-72 |
| Education | Highest completed education is high school or less | Highest completed education is a university degree |
| Technical background | Not technical; no technical/STEM work, training, or study | Very technical; extensive technical/STEM work, training, or study |
| Preferred writing style | Prefers direct and literal wording | Enjoys expressive and literary wording |