The Missing Coefficients: Bayesian Pairwise Merging for Model Personalization
Organizations: Monash University
Abstract
How can we personalize a shared expert library from a user's pairwise choices? Prior work can realize different reward trade-offs by merging reward-specialized experts, given a vector of trade-off weights. In practice, users can more naturally choose between outputs than specify numerical weights. The challenge is therefore to turn these choices into the coefficients required for merging, while accounting for ambiguity when feedback is limited. Our key idea is to treat the unknown reward weights as latent variables: infer a posterior over them from pairwise choices and reward-score differences, and use its mean directly as the merge coefficients. We instantiate this idea as Bayesian Pairwise Merging (BPM), whose posterior also characterizes which reward trade-offs remain plausible given the feedback. We evaluate BPM on radiology summarization, image captioning, and story generation, spanning text-to-text and image-to-text generation. With 100 feedback per simulated persona, BPM achieves macro decided win rates of 91.7%, 77.1%, and 64.3% against uniform merge. For six pairs of simulated personas, each prefers the model fitted to its own feedback, a pattern also observed in a human proof-of-concept. In simulations under BPM's model and prior, its nominal 90% intervals for temperature-scaled reward weights achieve task-averaged marginal coverage of 88.9% and 89.2% with only 10 and 25 comparisons, respectively. BPM thus enables personalization from pairwise feedback without per-user policy training, while characterizing the coefficient ambiguity left by limited feedback.
Figures & tables
| Summarization | Image Captioning | Story Generation | ||||
| RS | 91.7 [90.3, 93.1] | 88.4 [86.2, 90.4] | 77.1 [74.7, 79.3] | 78.7 [76.6, 80.7] | 64.3 [62.5, 66.1] | 76.1 [74.4, 77.8] |
| Base | 76.8 [74.6, 79.0] | 76.7 [74.0, 79.3] | 74.6 [72.2, 76.7] | 76.4 [74.4, 78.2] | 71.3 [69.4, 73.1] | 81.4 [79.9, 82.8] |
| RM | 74.9 [73.2, 76.7] | 84.1 [81.8, 86.2] | 72.1 [70.3, 74.0] | 75.3 [73.5, 77.0] | 57.3 [56.1, 58.6] | 71.5 [70.2, 72.7] |
| MLE | 28.4 [26.6, 30.3] | 43.9 [38.8, 49.0] | 44.7 [43.0, 46.4] | 45.4 [43.3, 47.6] | 43.4 [42.2, 44.7] | 46.3 [45.0, 47.6] |
| MAP | 28.9 [26.6, 31.2] | 51.1 [46.1, 56.0] | 57.9 [55.5, 60.2] | 58.8 [56.4, 61.3] | 49.5 [48.2, 50.7] | 51.9 [50.8, 53.1] |
| (a) Expert Training Algorithm | (b) Rubric-based Persona | ||||||
|---|---|---|---|---|---|---|---|
| GRPO | RLOO | DPO | Uniform Merge | Base Policy | Ridge-Merge | BT-BoN | |
| 91.7 [90.3,93.1] | 93.9 [92.6,95.1] | 77.7 [74.2,81.1] | 65.7 [63.7,67.7] | 70.8 [69.0,72.8] | 60.4 [59.1,61.7] | 63.4 [61.7,65.1] | |
| 88.4 [86.2,90.4] | 97.7 [96.9,98.4] | 90.3 [88.2,92.1] | 74.7 [72.8,76.5] | 77.8 [76.2,79.3] | 71.4 [70.2,72.7] | 71.1 [69.3,72.7] | |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Definition |
|---|---|
| Input and generated output. | |
| Number of rewards and reward/expert index. | |
| Model and shared base-model parameters. | |
| Shared library of reward-specialized experts. | |
| Probability simplex in . | |
| Merge coefficients, set to in BPM. |
| Data role | Summarization | Image Captioning | Story Generation |
|---|---|---|---|
| Expert-training pool | 85.6k reports | 600 images | 1,000 prompts |
| Prior calibration | 1,000 patients | 1,000 images | 1,000 prompts |
| Pool simulation | 250 reports | 250 images | 250 prompts |
| Reward standardization | 250 reports | 250 images | 250 prompts |
| Test | 451 reports | 500 images | 500 prompts |
| Setting | Summarization | Image Captioning | Story Generation |
|---|---|---|---|
| Training inputs per expert | 1,600 | 600 | 1,000 |
| Optimizer steps | 100 | 100 | 100 |
| Learning rate | |||
| Completions per prompt | 8 | 8 | 8 |
| Completions per optimizer step | 128 | 128 | 128 |
| Sampling temperature | 1.0 | 1.0 | 0.7 |
| Persona | Nonzero criterion weights |
|---|---|
| Literary | Prose elegance ; showing rather than telling ; imagery ; avoidance of purple prose . |
| Genre | Engagement ; surprise ; dialogue quality ; coherence . |
| Librarian | Readability ; coherence ; empathy ; avoidance of purple prose . |
| Flash | Brevity ; surprise ; imagery . |
| Screenwriter | Dialogue quality ; showing rather than telling ; imagery . |
| Description | Comparator | Wins | Losses | Ties | Win rate (%) |
|---|---|---|---|---|---|
| Picture Book | Uniform | 44 | 8 | 48 | 84.6 [72.5, 92.0] |
| Magazine Editor | Uniform | 84 | 7 | 9 | 92.3 [85.0, 96.2] |
| Picture Book | Magazine Editor model | 89 | 5 | 6 | 94.7 [88.1, 97.7] |
| Magazine Editor | Picture Book model | 93 | 3 | 4 | 96.9 [91.2, 98.9] |
| Persona | Uniform | Base Policy | Ridge-Merge | BT-BoN | |
|---|---|---|---|---|---|
| Literary | 100 | 69.4 [65.9, 72.9] | 80.6 [77.7, 83.5] | 61.6 [58.8, 64.4] | 67.4 [64.5, 70.7] |
| 500 | 82.6 [79.7, 85.2] | 90.2 [88.2, 92.2] | 77.6 [74.9, 80.1] | 81.9 [79.0, 84.7] | |
| Genre | 100 | 65.7 [62.3, 69.0] | 68.0 [64.5, 71.2] | 60.8 [58.1, 63.5] | 64.0 [60.8, 67.1] |
| 500 | 68.8 [65.4, 72.1] | 70.8 [67.6, 74.2] | 66.1 [63.6, 68.5] | 65.9 [62.7, 69.0] | |
| Librarian | 100 | 59.6 [55.8, 63.2] | 49.8 [45.7, 53.5] | 52.3 [49.4, 55.1] | 64.4 [61.2, 67.5] |
| 500 | 67.2 [63.7, 70.6] | 57.7 [53.9, 61.5] | 61.1 [58.4, 63.9] | 68.9 [65.6, 72.0] |
| Judge | vs. uniform-weight retrain | vs. bpm merge | vs. other persona |
|---|---|---|---|
| Ward | 98.37 [97.20,99.53] | 59.09 [36.36,77.27] | 100.00 [99.06,100.00] |
| ED | 85.40 [81.37,89.13] | 48.76 [41.79,55.72] | 97.95 [96.59,99.09] |
| RLOO experts | DPO experts | |||
|---|---|---|---|---|
| Persona | ||||
| Ward | 92.4 [89.7, 94.9] | 97.4 [96.0, 98.5] | 74.9 [68.7, 80.6] | 87.2 [83.8, 90.2] |
| ED | 94.8 [93.0, 96.3] | 96.4 [94.8, 97.8] | 86.8 [83.6, 89.9] | 93.4 [90.9, 95.5] |
| Complete | 93.3 [90.8, 95.6] | 99.2 [98.5, 99.8] | 78.5 [72.7, 83.8] | 92.3 [89.5, 94.8] |
| Concise | 96.5 [95.2, 97.7] | 97.9 [96.4, 99.1] | 83.2 [77.8, 87.9] | 92.6 [90.1, 94.7] |
| Faithful | 94.3 [91.9, 96.5] | 98.2 [97.0, 99.3] | 72.7 [65.9, 79.4] | 89.2 [85.4, 92.6] |
| Interval | Coverage | Width | Coverage | Width |
|---|---|---|---|---|
| bpm posterior | 88.9 | 0.176 | 89.2 | 0.191 |
| Prior only | 89.8 | 0.171 | 89.9 | 0.171 |
| MAP bootstrap | 8.2 | 0.024 | 20.4 | 0.261 |
| MLE bootstrap | 53.1 | 11.144 | 61.6 | 8.974 |
| SS sequential Gaussian | 99.9 | 3.036 | 99.7 | 2.797 |
| Setting | coverage range | coverage range |
|---|---|---|
| Generating weight configurations | ||
| One weight , | 93.6–96.6 | 93.8–96.9 |
| One weight , | 93.8–96.6 | 93.8–96.8 |
| One weight , | 93.8–96.6 | 93.8–96.6 |
| Three weights , | 81.2–90.8 | 81.2–91.0 |
| Three weights , | 81.1–89.9 | 82.3–90.8 |
| Task | Shrinkage win rate (%) | Shrinkage variability | BPM variability | ||
|---|---|---|---|---|---|
| Summarization | 25 | 0.70 | 77.6 [75.6, 79.4] | 0.299 | 0.112 |
| Summarization | 100 | 0.20 | 69.8 [67.8, 71.7] | 0.646 | 0.465 |
| Image Captioning | 100 | 0.25 | 41.9 [40.0, 43.8] | 0.733 | 0.313 |
| Story Generation | 100 | 0.60 | 44.5 [43.3, 45.7] | 0.328 | 0.392 |
| Method | Variability [95% CI] | Difference from BPM [95% CI] | |
|---|---|---|---|
| 25 | BPM | 0.112 [0.080, 0.142] | – |
| 25 | MAP | 0.231 [0.098, 0.377] | +0.119 [+0.014, +0.240] |
| 25 | MLE | 0.998 [0.494, 1.482] | +0.886 [+0.412, +1.346] |
| 25 | Shrinkage | 0.299 [0.148, 0.445] | +0.187 [+0.067, +0.307] |
| 25 | RM | 0.331 [0.297, 0.367] | +0.219 [+0.173, +0.263] |
| 25 | SS | 1.498 [1.216, 1.705] | +1.386 [+1.092, +1.587] |
| Method | Variability [95% CI] | Difference from BPM [95% CI] | |
|---|---|---|---|
| 25 | BPM | 0.358 [0.214, 0.495] | – |
| 25 | MAP | 0.366 [0.205, 0.540] | +0.007 [-0.209, +0.245] |
| 25 | MLE | 1.592 [1.433, 1.747] | +1.234 [+0.938, +1.511] |
| 25 | Shrinkage | 0.000 [0.000, 0.000] | -0.358 [-0.495, -0.214] |
| 25 | RM | 0.323 [0.261, 0.389] | -0.036 [-0.212, +0.157] |
| 25 | SS | 0.742 [0.174, 1.375] | +0.383 [-0.277, +1.138] |
| Method | Variability [95% CI] | Difference from BPM [95% CI] | |
|---|---|---|---|
| 25 | BPM | 0.088 [0.059, 0.116] | – |
| 25 | MAP | 0.119 [0.056, 0.190] | +0.032 [-0.010, +0.079] |
| 25 | MLE | 1.398 [1.249, 1.554] | +1.310 [+1.152, +1.482] |
| 25 | Shrinkage | 0.000 [0.000, 0.000] | -0.088 [-0.116, -0.059] |
| 25 | RM | 0.439 [0.389, 0.493] | +0.351 [+0.297, +0.414] |
| 25 | SS | 1.306 [1.113, 1.501] | +1.219 [+1.038, +1.405] |
| Task | |||||
|---|---|---|---|---|---|
| Summarization | |||||
| Image Captioning | |||||
| Story Generation | |||||
| BPM exact vs. SS exact | BPM exact vs. BPM factor | ||
|---|---|---|---|
| Task | |||
| Summarization | |||
| Image Captioning | |||
| Story Generation | |||
| Summarization | Image Captioning | Story Generation | ||||
|---|---|---|---|---|---|---|
| BPM | SS | BPM | SS | BPM | SS | |
| 10 | ||||||
| 25 | ||||||
| 50 | ||||||
| 100 | ||||||
| 250 | ||||||
| BPM Selection | Selection Uniform-weight selection | BPM Uniform-weight selection | |
|---|---|---|---|
| 10 | |||
| 25 | |||
| 50 | |||
| 100 | |||
| 250 | |||
| 500 |
| task | compared method | Dec. | PP [95% CI] | PP bounds | OI | Dec. | PP [95% CI] | PP bounds | OI |
|---|---|---|---|---|---|---|---|---|---|
| Summarization | RS (uniform merge) | 91.7 | 76.5 [75.4, 77.7] | 71.3–81.8 | 21.1 | 88.4 | 74.0 [72.7, 75.3] | 65.3–82.7 | 34.9 |
| Base policy | 76.8 | 69.3 [67.8, 70.7] | 62.7–75.9 | 26.4 | 76.7 | 68.5 [67.2, 69.9] | 59.3–77.7 | 36.8 | |
| L2W | 59.6 | 63.1 [62.2, 63.9] | 59.5–66.6 | 14.0 | 81.9 | 72.5 [71.5, 73.6] | 68.4–76.7 | 16.6 | |
| Ridge-Merge | 74.9 | 62.8 [61.8, 63.8] | 57.1–68.4 | 22.6 | 84.1 | 67.9 [66.6, 69.2] | 58.9–76.9 | 35.9 | |
| BT + best-of-8 | 82.5 | 70.0 [68.8, 71.2] | 64.2–75.8 | 23.0 | 85.1 | 69.6 [68.2, 70.9] | 60.4–78.7 | 36.5 | |
| Task | Cohen’s | Decided agreement | Max. win-rate change |
|---|---|---|---|
| Summarization | – | – | |
| Image Captioning | – |
| Task | Compared method | Original judge | gpt-5-mini |
|---|---|---|---|
| Summarization | Uniform merge | ||
| Ridge-Merge | |||
| DPO | |||
| Image Captioning | Uniform merge | ||
| Ridge-Merge | |||
| DPO |