Personalized image generation has remained narrowly focused on conditional synthesis from curated visual exemplars, rather than capturing who a user is. In practice, however, a user's personal context is much richer, comprising reviews, posts, images, captions, and metadata accumulated over time. A truly personalized generator should leverage this history to produce images aligned with the user's lifestyle and aesthetic preferences. To this end, we introduce the first unified benchmark for personalized image generation from user histories. The benchmark comprises two complementary tasks and a multi-axis evaluation protocol that assesses target fidelity, visual quality, user distinguishability, semantic alignment with the user's history, and task-specific utility. Grounded in real-world e-commerce and social media settings, the benchmark includes: (1) Personalized Scene Generation, which places a given object in a scene that reflects a user's preferences and lifestyle, motivated by personalized product presentation; and (2) Personalized Creative Generation, which generates a novel image on a specified topic that is faithful to a user's aesthetic and visual identity, motivated by social media content creation. We further propose PEARL, which couples a multimodal reasoner with a frozen image generator in an interleaved reason-reflect loop optimized with differential data reward. Across both tasks, PEARL outperforms strong baselines, achieving an average improvement of 15% across personalization metrics.
Figures & tables
Figure 1: Personalized Image Generation based on User History.
Figure 2: Training procedure of the Pearl framework. We warm start with the Silver Personalization Trajectory, and optimize both policies with the Render-in-the-Loop optimization.
Retrieval
Image Quality
MLLM Judge ( 1 – 5 )
Method
H@5 ↑
MRR ↑
Aes ↑
Cont. ↑
Overall ↑
Sty. ↑
PMG ( Shen et al., 2024 )
0.2143
0.1535
5.11
3.58
3.48
2.48
Pigeon ( Xu et al., 2025 )
0.0985
0.0834
4.60
4.32
1.79
1.98
LaVIT ( Jin et al., 2024 )
0.2162
0.1441
4.38
3.94
2.92
2.37
LLaVA ( Liu et al., 2023 )
0.1873
0.1367
5.17
4.32
2.89
2.30
Pearl (ours)
0.2297
0.1639
5.33
4.36
3.92
2.88
Table 1: Main results on Personalized Scene Generation for e-commerce setting on Amazon. (Bold marks the best per column; underline marks second-best.)
Target-based Image Metrics
Contrastive R@1
MLLM Judge ( 1 – 5 )
Method
CIS ↑
DIS ↑
LPIPS ↓
MS-SSIM ↑
Aes. ↑
Inter-cat. ↑
Intra-cat. ↑
Sty. ↑
Cont. ↑
Overall ↑
PMG ( Shen et al., 2024 )
0.601
0.206
0.729
0.047
5.84
0.597
0.466
3.167
3.075
3.157
Pigeon ( Xu et al., 2025 )
0.561
0.256
0.741
0.053
4.28
0.748
0.659
3.341
3.258
3.334
LaVIT ( Jin et al., 2024 )
0.616
0.238
0.744
0.068
5.37
0.640
0.395
3.344
3.143
3.331
LLaVA ( Liu et al., 2023 )
0.575
0.171
0.735
0.053
5.48
0.432
0.285
3.256
3.179
3.227
Pearl (ours)
0.645
0.289
0.735
0.055
5.45
0.876
0.701
3.554
3.277
3.541
Table 2: Main results on Personalized Creative Generation for social media setting on Instagram. (Bold marks the best per column; underline marks second-best.)
Figure 3: Qualitative Results.
Figure 4: Ablation: Pearl vs. Pearl -Reflection. The visualization normalizes and aggregates the metrics across both tasks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Personalized Scene Generation (E-commerce: Amazon)
Table 4: Human evaluation of the original SDXL outputs. Entries are counts (percentages), with 25 pairs per task and question. The baseline is PMG for scenes and Pigeon for creative generation. “Other” denotes insufficient evidence for history fit and neither image satisfactory for product preservation or topic adherence.
Retrieval
Image Quality
MLLM Judge ( 1 – 5 )
Method
H@5 ↑
MRR ↑
Aes ↑
Cont. ↑
Overall ↑
Sty. ↑
Pearl -Reflection
0.2336
0.1583
5.32
4.14
3.99
2.74
Pearl (ours)
0.2297
0.1639
5.33
4.36
3.92
2.88
Appendix
Table 5: Ablation Study for Personalized Scene Generation. Bold marks the best per column.
Target-based Image Metrics
Contrastive R@1
MLLM Judge ( 1 – 5 )
Method
CIS ↑
DIS ↑
LPIPS ↓
MS-SSIM ↑
Inter-cat. ↑
Intra-cat. ↑
Sty. ↑
Cont. ↑
Overall ↑
PEARL-Reflection
0.635
0.285
0.732
0.057
0.838
0.701
3.529
3.331
3.516
Pearl (ours)
0.645
0.289
0.735
0.055
0.876
0.701
3.554
3.277
3.541
Appendix
Table 6: Ablation Study for Personalized Creative Generation. Bold marks the better of the two per column.
Figure 5: Additional qualitative results on Personalized Creative Generation. For each of three users, the top row shows five historical posts alongside a held-out target post (rightmost), and the bottom row shows generations from LaVIT, LLaVA, PMG, Pigeon, and Pearl . The examples illustrate recurring choices in color, framing, and scene presentation across interiors, food, and quilting content.
Symbol
Description
Problem Setup
u
A user.
Hu
User u ’s multimodal history; a sequence of prior activities.
hi(u)
The i -th history entry for user u (e.g., a review with product image, or a post with caption).
Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt. To this end, we introduce PIPBench, the first profile-inclusive benchmark for evaluating personalized image generation. We further propose a novel data construction pipeline that leverages psychological and demographic profiling dimensions for both real-user data collection and scalable agent-based data generation. Using PIPBench, we conduct a thorough evaluation of representative line of methods. Our experiments reveal key limitations in existing methods, suggesting new challenges and opportunities for personalized text-to-image synthesis. Project page: https://wuyuhang05.github.io/PIPBench/
Yuhang Wu, Shuxiang Zhang, Wee Hian Ching +2
College of AI, Tsinghua University · Sun Yat-sen University · Shanghai Qi Zhi Institute
Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste. Human preferences are pluralistic: one user favoring muted, nostalgic portraits may prefer vibrant street photography, while another gravitates toward dreamy film aesthetics. Existing methods require dense interaction histories or per-user fine-tuning, failing in cold-start settings and collapsing context-dependent preferences into a static representation. We introduce zero-shot image personalization from personas (ZIPP), which conditions image generation on natural-language personas (concise descriptors of a user's identity and aesthetic sensibilities) without any user-specific data or weight updates. ZIPP uses an LLM to rewrite prompts from the perspective of a given persona, steering diffusion models toward personalized outputs. To mine personas at scale, we train an inductive Graph Attention Network over a 22M-user Reddit interaction graph with dual contrastive objectives aligning graph structure with visual behavior, then verbalize learned representations into natural-language personas via an MLLM. We introduce ZIPBench, the first zero-shot personalization benchmark with 1.5K users, graph-mined personas, and 40K generated images. Across four benchmarks and 14 LLMs spanning five model families, persona conditioning yields consistent gains (13-20%), with frontier models benefiting most. In the few-shot setting, ZIPP matches or exceeds fine-tuned baselines trained on 100+ examples per user. ZIPP achieves the lowest preference distributional divergence (CMMD 0.16 vs. 0.55), and IPF-normalized demographic evaluation shows it substantially reduces subpopulation bias present in existing methods. Human evaluation confirms a 79% win rate over generic generation and 58-65% over all fine-tuned baselines.
Harini SI, Somesh Singh, Yaman Kumar Singla +2
Adobe Media and Data Science Research (MDSR) · SUNY at Buffalo · IIIT-Delhi
Users increasingly expect image generation models to quickly adapt to highly diverse and personalized requirements, such as producing images with distinctive styles or characteristics. Traditional approaches rely on fine-tuning, which is costly and difficult to scale. To cope with these limitations, the community has accumulated a growing library of fine-tuned modules and adapters, where each component targets specific generation needs and collectively serves as a foundation for handling new demands. This naturally raises a question: instead of repeatedly training new models, can we systematically exploit this expanding ecosystem to better fulfill user instructions? To this end, we present Polaris, an intelligent retrieval framework that automatically selects and integrates suitable models from the model library based on a user's instructions. The key insight is that harnessing such a massive and heterogeneous pool requires not only finding the most relevant modules among thousands of candidates, but also aligning them effectively for instruction-driven generation and editing. Polaris addresses this challenge by indexing over 6,500 checkpoints and 75,000 adapters, and retrieving the most relevant components given a user's input and instruction. In doing so, it delivers scalable, controllable, and well-aligned generation -- without any additional training.
Zhi-Kai Chen, Jun-Peng Jiang, Jun-Jie Tao +2
School of Artificial Intelligence, Nanjing University, China · National Key Laboratory for Novel Software Technology, Nanjing University, China · Nanjing University, China