Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user's lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for personalized image generation from user histories. The benchmark comprises two complementary tasks and a multi-axis evaluation protocol that assesses target fidelity, visual quality, user distinguishability, semantic alignment with the user's history, and task-specific utility. Grounded in real-world e-commerce and social media settings, the benchmark includes: (1) Personalized Scene Generation, which places a given object in a scene that reflects a user's preferences and lifestyle, motivated by personalized product presentation; and (2) Personalized Creative Generation, which generates a novel image on a specified topic that is faithful to a user's aesthetic and visual identity, motivated by social media content creation. We further propose PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward. Across both tasks, PEARL outperforms strong baselines, achieving an average improvement of 15% across personalization metrics.
Figures & tables
Figure 1: Personalized Image Generation based on User History.
Figure 2: Training procedure of the Pearl framework. We warm start with the Silver Personalization Trajectory, and optimize both policies with the Render-in-the-Loop optimization.
Retrieval
Image Quality
MLLM Judge ( 1 – 5 )
Method
H@5 ↑
MRR ↑
Aes ↑
Cont. ↑
Overall ↑
Sty. ↑
PMG ( Shen et al., 2024 )
0.2143
0.1535
5.11
3.58
3.48
2.48
Pigeon ( Xu et al., 2025 )
0.0985
0.0834
4.60
4.32
1.79
1.98
LaVIT ( Jin et al., 2024 )
0.2162
0.1441
4.38
3.94
2.92
2.37
LLaVA ( Liu et al., 2023 )
0.1873
0.1367
5.17
4.32
2.89
2.30
Pearl (ours)
0.2297
0.1639
5.33
4.36
3.92
2.88
Table 1: Main results on Personalized Scene Generation for e-commerce setting on Amazon. (Bold marks the best per column; underline marks second-best.)
Target-based Image Metrics
Contrastive R@1
MLLM Judge ( 1 – 5 )
Method
CIS ↑
DIS ↑
LPIPS ↓
MS-SSIM ↑
Aes. ↑
Inter-cat. ↑
Intra-cat. ↑
Sty. ↑
Cont. ↑
Overall ↑
PMG ( Shen et al., 2024 )
0.601
0.206
0.729
0.047
5.84
0.597
0.466
3.167
3.075
3.157
Pigeon ( Xu et al., 2025 )
0.561
0.256
0.741
0.053
4.28
0.748
0.659
3.341
3.258
3.334
LaVIT ( Jin et al., 2024 )
0.616
0.238
0.744
0.068
5.37
0.640
0.395
3.344
3.143
3.331
LLaVA ( Liu et al., 2023 )
0.575
0.171
0.735
0.053
5.48
0.432
0.285
3.256
3.179
3.227
Pearl (ours)
0.645
0.289
0.735
0.055
5.45
0.876
0.701
3.554
3.277
3.541
Table 2: Main results on Personalized Creative Generation for social media setting on Instagram. (Bold marks the best per column; underline marks second-best.)
Figure 3: Qualitative Results.
Figure 4: Ablation: Pearl vs. Pearl -Reflection. The visualization normalizes and aggregates the metrics across both tasks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Personalized Scene Generation (E-commerce: Amazon)
Table 4: Human evaluation of the original SDXL outputs. Entries are counts (percentages), with 25 pairs per task and question. The baseline is PMG for scenes and Pigeon for creative generation. “Other” denotes insufficient evidence for history fit and neither image satisfactory for product preservation or topic adherence.
Retrieval
Image Quality
MLLM Judge ( 1 – 5 )
Method
H@5 ↑
MRR ↑
Aes ↑
Cont. ↑
Overall ↑
Sty. ↑
Pearl -Reflection
0.2336
0.1583
5.32
4.14
3.99
2.74
Pearl (ours)
0.2297
0.1639
5.33
4.36
3.92
2.88
Appendix
Table 5: Ablation Study for Personalized Scene Generation. Bold marks the best per column.
Target-based Image Metrics
Contrastive R@1
MLLM Judge ( 1 – 5 )
Method
CIS ↑
DIS ↑
LPIPS ↓
MS-SSIM ↑
Inter-cat. ↑
Intra-cat. ↑
Sty. ↑
Cont. ↑
Overall ↑
PEARL-Reflection
0.635
0.285
0.732
0.057
0.838
0.701
3.529
3.331
3.516
Pearl (ours)
0.645
0.289
0.735
0.055
0.876
0.701
3.554
3.277
3.541
Appendix
Table 6: Ablation Study for Personalized Creative Generation. Bold marks the better of the two per column.
Figure 5: Additional qualitative results on Personalized Creative Generation. For each of three users, the top row shows five historical posts alongside a held-out target post (rightmost), and the bottom row shows generations from LaVIT, LLaVA, PMG, Pigeon, and Pearl . The examples illustrate recurring choices in color, framing, and scene presentation across interiors, food, and quilting content.
Symbol
Description
Problem Setup
u
A user.
Hu
User u ’s multimodal history; a sequence of prior activities.
hi(u)
The i -th history entry for user u (e.g., a review with product image, or a post with caption).
School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China · Nanjing University, China