Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-level information and overlook diverse user preferences. We identify group preference collapse, where multi-user personalized MLLMs become insensitive to individual preferences and drift toward dominant population-level choices due to suppressed preference signals and unreliable preference use during generation. We propose PrefMoE, a preference-centric framework that separates stable profile information from preference-related representations. PrefMoE decomposes preferences into shared prototypes and personalized residuals, preserves individualized residuals with imbalance-aware learning, counterfactual pseudo-user augmentation, and residual decorrelation, and routes profile and preference factors through separate LoRA adaptation paths. Experiments across multiple MLLM backbones show that PrefMoE improves preference-sensitive personalization while substantially reducing preference collapse. Project page: https://prefmoe.github.io/.
Figures & tables
Figure 1: Illustration of group preference collapse . Although users have diverse preferences, existing personalized MLLMs tend to predict the dominant preference across users (“Yoga”) rather than reflecting the target user’s specific preference. Bars indicate the percentage of users from each ground-truth preference group whose Top-1 preference is the dominant preference (“Yoga”).
Figure 2: Overview of PrefMoE . PrefMoE separates profile and preference representations, models preferences with shared prototypes and personalized residuals, and regularizes the residuals with contrastive learning and decorrelation. A hierarchical MoE router then activates profile- and preference-aware LoRA experts for query-dependent personalized reasoning.
Method
Type
0-turn
10-turn
Overall ↑
Preference ↑
Profile ↑
Collapse ↓
Overall ↑
Preference ↑
Profile ↑
Collapse ↓
LLaVA-1.5-7B
NT
0.3601
0.3387
0.3716
0.6228
0.3243
0.3067
0.3337
0.6164
LLaVA-1.5-13B
NT
0.4023
0.3653
0.4222
0.6027
0.3720
0.3333
0.3928
0.6119
LLaVA-OV-72B
NT
0.5100
0.4800
0.5262
0.5297
0.4620
0.4333
0.4774
0.5525
DeepSeek-VL2-Tiny
NT
0.3683
0.3613
0.3720
0.7169
0.3497
0.3440
0.3527
0.6895
DeepSeek-VL2-Small
NT
0.3866
0.3853
0.3871
0.6895
0.3668
0.3667
0.3669
0.6712
Table 1: Major comparisons with SOTAs under 0-turn and 10-turn settings.
Figure 3: Cross-domain results on food and fashion using LLaVA-1.5-7B.
0-turn
10-turn
E
P
I
C
D
M
Overall ↑
Preference ↑
Profile ↑
Collapse ↓
Overall ↑
Preference ↑
Profile ↑
Collapse ↓
✓
0.5562
0.4120
0.6337
0.6027
0.4928
0.3584
0.5651
0.6347
✓
✓
0.6000
0.5700
0.6159
0.3333
0.5266
0.4902
0.5461
0.3653
✓
✓
✓
0.6279
0.6133
0.6358
0.3288
0.5506
0.5275
0.5631
0.3607
✓
✓
✓
✓
0.6303
0.6253
0.6330
0.2283
0.5571
0.5393
0.5667
0.2603
✓
✓
✓
✓
✓
0.6382
0.6467
0.6337
0.1457
0.5688
0.5622
0.5723
0.1781
Table 2: Component-wise ablation study of the proposed method. E, P, I, C, D, and M denote the basic user embedding module, profile factor learning, imbalance-aware residual preservation, counterfactual user augmentation, preference decorrelation, and hierarchical MoE router, respectively.
Figure 6Figure 7
Figure 9: Qualitative examples. <SKS> denotes the target personalized identity or concept. Green boxes indicate correct predictions, while red marks indicate incorrect predictions.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Users
Samples
Judgments
Agreement ↑
Confidence (1–5)
50
200
2,000
99.40%
4.97
Appendix
Table 3: Human audit of reconstructed MMPB-Clean labels. Judgments denotes individual annotation decisions, not the number of annotators.
Statistic
Train
Evaluation
Total
Dataset scale
QA pairs
10371
2145
12516
Images
8272
2131
10016
Human users
50
50
50
Non-human concepts
61
61
61
Preference QA by facet
Appendix
Table 4: Detailed statistics of MMPB-Clean. The same evaluation split is used for both 0-turn and 10-turn evaluation settings.
Group
Accuracy ↑
Collapse share among errors ↓
Fine-Tuning
PrefMoE
Fine-Tuning
PrefMoE
High
50
73
–
20
Mid
–
65
–
27
Low
28
53
84
32
Appendix
Table 5: Available preference-frequency diagnostics. All entries are percentages. A dash denotes a value not reported in this diagnostic.
Group
Method
Accuracy ↑
Collapse ↓
Male ( n=28 )
Fine-Tuning
0.5108
0.3391
PrefMoE
0.6724
0.1258
Female ( n=22 )
Fine-Tuning
0.5026
0.3468
PrefMoE
0.6785
0.1201
Appendix
Table 6: Gender-stratified diagnostic using available user annotations (LLaVA-1.5-7B).
Method
Setting
≤4
5–8
≥9
Overall
LLaVA-1.5-7B
NT
0.6020
0.6103
0.6760
0.6228
LLaVA-1.5-13B
NT
0.5750
0.5958
0.6490
0.6027
LLaVA-OV-72B
NT
0.5140
0.5333
0.5370
0.5297
LLaVA-1.5-7B
FFT
0.3180
0.3381
0.3790
0.3425
LLaVA-1.5-13B
FFT
0.2920
0.3056
0.3420
0.3105
LLaVA-OV-72B
FFT
0.2970
0.3056
0.3580
0.3151
Appendix
Table 7: Collapse breakdown for LLaVA-family models across user frequency groups.
Method
Setting
Entertainment
Fashion
Lifestyle
Shopping
Travel
Overall
LLaVA-1.5-7B
NT
0.6250
0.6458
0.6222
0.6000
0.6364
0.6228
LLaVA-1.5-13B
NT
0.5833
0.6250
0.6000
0.5778
0.6061
0.6027
LLaVA-OV-72B
NT
0.5208
0.5625
0.4889
0.5333
0.5455
0.5297
LLaVA-1.5-7B
FFT
0.3542
0.3750
0.3333
0.2889
0.3636
0.3425
LLaVA-1.5-13B
FFT
0.3125
0.2917
0.2889
0.2667
0.3939
0.3105
LLaVA-OV-72B
FFT
0.3333
0.3542
0.3111
0.2444
0.3030
0.3151
Appendix
Table 8: Facet-wise preference collapse on the LLaVA family.
Method
Type
Entertainment
Travel
Lifestyle
Shopping
Fashion
Ovr.Preference
LLaVA-1.5-7B
NT
0.3310
0.3440
0.3370
0.3420
0.3399
0.3387
LLaVA-1.5-13B
NT
0.3590
0.3670
0.3610
0.3710
0.3688
0.3653
LLaVA-OV-72B
NT
0.4740
0.4860
0.4770
0.4820
0.4812
0.4800
LLaVA-1.5-7B
FFT
0.4370
0.4560
0.3870
0.4620
0.4626
0.4413
LLaVA-1.5-13B
FFT
0.4520
0.4710
0.4050
0.4860
0.4846
0.4600
LLaVA-OV-72B
FFT
0.5850
0.6130
0.5310
0.6310
0.6121
0.5947
Appendix
Table 9: Facet-wise 0-turn preference accuracy on the LLaVA family.
Method
Train Cost
Eval Cost
Peak Memory
LLaVA-1.5-7B
–
0.5–1 GPU-hours
15–20 GB/GPU
Yo’LLaVA-7B
5–10 GPU-hours
1–2 GPU-hours
28–34 GB/GPU
LOVA3-7B
6–12 GPU-hours
1–2 GPU-hours
28–34 GB/GPU
TG-LLaVA-7B
6–12 GPU-hours
1–2 GPU-hours
28–34 GB/GPU
PrefMoE (LLaVA-1.5-7B)
8–16 GPU-hours
1–2 GPU-hours
28–34 GB/GPU
PrefMoE (LLaVA-1.5-13B)
24–32 GPU-hours
2–3 GPU-hours
34–39 GB/GPU
Appendix
Table 10: Computational cost comparison. GPU-hours are computed as the number of GPUs multiplied by wall-clock time.
Figure 10: Row-normalized transitions from ground-truth to predicted preference-frequency groups. High, Mid, and Low denote at least nine, five to eight, and at most four users, respectively. All panels share the same color scale. Each row sums to one; the diagonal denotes a matching frequency group, not exact-answer accuracy.
Method
0-turn
10-turn
Accuracy ↑
Collapse ↓
Accuracy ↑
Collapse ↓
Fine-Tuning
0.400
0.425
0.325
0.475
TG-LLaVA
0.475
0.325
0.425
0.375
PrefMoE
0.575
0.175
0.500
0.250
Appendix
Table 11: Multi-label diagnostic with 40 questions and two acceptable preferences per question.
Method
Overall ↑
Preference ↑
Profile ↑
Collapse ↓
Fine-Tuning
0.4921
0.4236
0.5289
0.3714
Yo’LLaVA
0.4794
0.4618
0.4889
0.2348
TG-LLaVA
0.5155
0.5524
0.4957
0.2586
LOVA3
0.5079
0.5417
0.4898
0.4568
PrefMoE
0.6542
0.6487
0.6571
0.1519
Appendix
Table 12: Preference queries conditioned on explicit weather contexts (0-turn).
Figure 11: Scores before and after two sequential rounds of preference changes. Circles show initial scores and squares show final scores; each segment connects the same method before and after updating. Overall, Preference, and Profile denote accuracies (higher-is-better); Collapse is lower-is-better.
Figure 12: Scores before and after two sequential rounds of profile changes, using the same layout and axis ranges as Figure 11 . Circles show initial scores and squares show final scores. Accuracies are higher-is-better and collapse is lower-is-better.
Method
Initial
After update 10
Initial-to-final drop
TG-LLaVA
0.5413
0.4956
0.0457
PrefMoE
0.6751
0.6065
0.0686
Appendix
Table 13: Ten-update accuracy, averaged equally across the ten user groups. These summary values are computed from the group trajectories below.
Figure 13: Overall accuracy across ten sequential updates, with one panel per user group. TG-LLaVA and PrefMoE are evaluated under the same protocol. Open markers at I show initial scores; filled markers show reported update-stage scores. Group t changes at update t , and unreported pre-update entries are omitted without interpolation. All panels share the same axes.