The demand for personalized LLMs is shifting from style imitation toward content quality. We investigate whether self-distillation can bridge this gap in existing fine-tuning paradigm. To address this limitation, we introduce MIRROR(Meta- personalization by Internalizing Reference-Revealed On-policy Reflections), a novel self-distillation framework that shifts LLM personalization from imitation toward preference internalization. First, we replace reference-token imitation with reference-revealed on-policy self-distillation, aligning the model's next-token distributions along its own generation trajectories with those of its reference-conditioned self, thereby internalizing user preferences rather than reproducing reference wording.Second, we introduce MIRROR-F, a focal plug-in that augments on-policy distributional alignment with selective supervision over informative reference tokens, thereby strengthening content generation while preserving user-specific expression. Across three personalized generation benchmarks, two model scales, and complementary reference-based and LLM-based evaluations, MIRROR and MIRROR-F achieve leading overall personalization performance and superior text quality, while exhibiting less catastrophic forgetting than SFT-based baselines on three unseen personalized generation tasks. The gains are consistent across model scales and application scenarios, translating to improved performance in LLM personalization tasks.
Figures & tables
Figure 1: Overview of personalization paradigms. (a) A case study of the quality defects of existing methods. (b) Retrieval-based personalization. (c) SFT-based personalization. (d) MIRROR, which aligns a self-student with a self-teacher on the student’s own on-policy trajectories.
Figure 2: Overview of MIRROR-F. A frozen reference-conditioned teacher ( h,x,y ) supervises the trainable student ( h,x ) on the student’s own rollouts via LMIRROR , while LANC anchors only the selected informative tokens of the reference y .
Datasets
Methods ( → )
Base
Retrieval-based Methods
PEFT-based Methods
Ours
Metrics ( ↓ )
Qwen3
RAG
LatestK
LLM-TRSR
SFT
OPPU
PerCE
NextQuill
MIRROR
MIRROR-F
Qwen3-1.7B
Abstract Generation
ROUGE-1
0.3717
0.3811
0.3685
0.3758
0.3755
0.3082
0.3670
0.3698
0.3967
0.4015
METEOR
0.2125
0.2262
0.2145
0.2311
0.2365
0.2073
0.2240
0.2437
0.2506
0.2529
BERTScore
0.8577
0.8675
0.8657
0.8656
0.8668
0.8443
0.8621
0.8694
0.8587
0.8701
News Headline Generation
ROUGE-1
0.1422
0.1495
0.1402
0.1387
0.1468
0.1559
0.1527
0.1594
0.1674
0.1763
Table 1: Main results of MIRROR and MIRROR-F on three personalized generation benchmarks Bold numbers denote the best performance, while underlined numbers denote the second best.
Figure 3: Catastrophic forgetting on held-out OOD personalized generation tasks from the Amazon Review dataset. Performance degradation is measured relative to the corresponding base model under the same evaluation protocol; smaller values indicate better retention. Bold numbers denote the smallest degradation, while underlined numbers denote the second-smallest.
Datasets ( → )
Abstract Generation
News Headline Generation
Paper Topic Generation
Methods ( ↓ )
Content
Style
C&S
Content
Style
C&S
Content
Style
C&S
Qwen3-1.7B
Base
0.669
8.360
7.523
0.077
4.520
2.644
0.353
8.000
5.766
SFT
0.563
7.990
6.808
0.080
4.730
2.765
0.417
8.360
6.267
PerCE
0.492
7.810
6.365
0.068
4.810
2.745
0.413
8.100
6.116
NextQuill
0.542
8.000
6.708
0.078
4.670
2.725
0.392
8.140
6.028
Table 2: ExPert evaluation of personalized generation quality on three benchmarks. The judge model scores each generated response for Content , Style , and their combination C&S (Content & Style); higher is better. Bold numbers denote the best performance among trained methods, while underlined numbers denote the second best among trained methods.
Figure 4: Comparison of model performance scores based on the G-eval framework.
Figure 5: The impact of fluctuations in the value of α on the performance of the model’s stylistic and syntactic personalization assessment.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Item
Content
Question
Generate an abstract for the title “Learning the Language: The Importance of Studying Written Directions in Designing Navigational Technologies for the Blind” using the following items: (1) Independent navigation; (2) Text-to-Speech directions; (3) Language and cues; (4) Qualitative analysis; and (5) Landmarks and alternate routes.
Gold answer
Independent navigation is important to individuals who are blind and visually impaired (VI). Researchers have long explored how blind and VI people navigate to inform the design of more useful, accessible wayfinding devices. However, there has been little research on the role language plays in providing effective text-to-speech directions for this population. Through qualitative analysis, this work examines the language and cues expressed in written navigational directions, including how distances are represented, how direction is indicated, and what landmarks are referenced. The analysis further identifies warnings about overshooting a destination, alternative routes that are easier to navigate, and the accessibility of destinations for people with disabilities.
Appendix
Table 3: Question–answer case for Abstract Generation.
Method
Answer
Base
This paper explores the importance of studying written directions in the design of navigational technologies for the blind, focusing on how independent navigation can be enhanced through effective use of text-to-speech directions. We conducted a qualitative analysis of how language and cues are used in written directions to guide users, emphasizing the need for clear, concise, and contextually relevant information. The study also examines the role of landmarks and alternate routes in providing users with multiple options for navigation, ensuring that they can make informed decisions in complex environments. Our findings highlight the significance of incorporating user-centered design principles that consider the unique needs and experiences of blind individuals when developing navigational technologies.
NextQuill
Independent navigation is a challenge for many people who are blind. One promising solution is to use text-to-speech directions to guide them. However, there has been little inquiry into how people who are blind use and interpret directions, and how technologies can be designed to better support them. We conducted a qualitative analysis of 11 blind participants’ experiences with and suggestions for text-to-speech directions. We found that users rely on language and cues to interpret directions, and that they had strong preferences for how these directions should be presented. We present two findings that can be used to guide the design of future text-to-speech directions: (1) the use of landmarks and alternate routes can be beneficial to users, and (2) users prefer directions that are presented in a conversational tone.
MIRROR
Independent navigation is a critical skill for individuals who are blind, enabling them to move confidently and safely in their environment. However, the effectiveness of navigational technologies for the blind often hinges on the clarity and usability of the directions provided. This paper explores the importance of studying written directions in the design of navigational technologies for the blind, focusing on how language and cues influence user experience. Through qualitative analysis of user interactions, we examine how text-to-speech directions are perceived and understood by users, highlighting the role of language in conveying information about landmarks and alternate routes. Our findings reveal that the way directions are structured and the cues used can significantly impact the user’s ability to navigate independently. We argue that a deeper understanding of language and cues in written directions is essential for designing more effective navigational technologies that support the unique needs of blind users.
MIRROR-F
Independent navigation for the blind is significantly influenced by the clarity and effectiveness of navigational technologies. This study emphasizes the importance of studying written directions in the design of navigational technologies. Through qualitative analysis, we explore how text-to-speech directions, language, and cues impact the user’s ability to navigate independently. The study highlights the role of language in conveying information and the necessity of clear, concise cues that aid in understanding the environment. We examine how landmarks and alternate routes are described and how these elements can be integrated into navigational systems to enhance usability. The findings underscore the need for a deep understanding of language in the design process, ensuring that navigational technologies are not only functional but also intuitive and accessible for users with visual impairments.
Appendix
Table 4: Different model answers for the Abstract Generation case.
Item
Content
Question
Generate a headline for the following article: “Um, OK. For those of us whose leggy limbs haven’t been immortalized in print, we can turn to H&M’s updated classics, like …”
Gold answer
H&M Model Collection Inspired By Joan Smalls, Daphne Groeneveld & More (PHOTOS)
Appendix
Table 5: Question–answer case for News Headline Generation.
Method
Answer
PerCE
7 Ways To Make Your Pantyhose Look Like A Model (PHOTOS)
NextQuill
Jenna Holmes’ ’90s Style Is Too Retro For A 2013 Cover (PHOTOS)
MIRROR
H&M’s Updated Classics: Leggy Limbs For Every Body Type (PHOTOS)
MIRROR-F
H&M’s Updated Classics: Leggy Limbs For Every Style (PHOTOS)
Appendix
Table 6: Different model answers for the News Headline Generation case.
Item
Content
Question
Generate a title for the following abstract: “This paper addresses the conversion of discrete-time single input-single output PieceWise Affine (PWA) models from state space to input-output form. Necessary and sufficient conditions are given for a PWA state space model to admit equivalent input-output representations. When an equivalent input-output model exists, a constructive procedure is presented to derive both its parameters and the partition of the regressors domain. It is shown that the number of modes and the number of parameters may grow considerably when converting a PWA state space model into an equivalent input-output representation. Numerical examples highlight the role of the derived necessary and sufficient conditions for input-output realization of PWA state space models.”
Gold answer
Input-Output Realization Of Piecewise Affine State Space Models
Appendix
Table 7: Question–answer case for Paper Topic Generation.
Method
Answer
PerCE
On the Input-Output Representation of Piecewise Affine State Space Models
NextQuill
Input-output realization of piecewise affine state space models
MIRROR
Conversion of Discrete-Time Piecewise Affine Models to Input-Output Form
MIRROR-F
Conversion of Piecewise Affine State Space Models to Input-Output Form
Appendix
Table 8: Different model answers for the Paper Topic Generation case.
Category
Hyperparameter
Value
Model and data
Backbone
Qwen3-1.7B/4B
Dataset
LongLaMP,LaMP-4,LaMP-5
Optimization
Optimizer
AdamW
Learning rate
1×10−4
Per-device batch size
1
Gradient accumulation
1
Appendix
Table 9: Training hyperparameters used for the MIRROR and MIRROR-F comparison.
Backbone
Task
Base
ContextSFT
PerCE
NextQuill
MIRROR(Ours)
MIRROR-F(Ours)
Qwen3-1.7B
Abstract
0.933
0.895
0.785
0.839
0.949
0.952
News
0.774
0.447
0.404
0.358
0.387
0.631
Paper
0.544
0.566
0.527
0.521
0.588
0.596
Qwen3-4B
Abstract
0.951
0.919
0.854
0.819
0.957
0.963
News
0.686
0.458
0.416
0.461
0.676
0.640
Paper
0.645
0.564
0.557
0.543
0.615
0.624
Appendix
Table 10: Rule-based information completeness evaluation. Higher values indicate better coverage of information-bearing content. Bold numbers denote the best performance among trained methods, while underlined numbers denote the second-best. The Base model is reported as a non-personalized reference and is excluded from the ranking.
α
Coherence
Consistency
Fluency
Relevance
Overall
0.1
4.480
4.810
4.980
4.580
4.713
0.3
4.440
4.830
5.000
4.550
4.705
0.7
4.500
4.880
4.990
4.570
4.735
0.9
4.440
4.690
5.000
4.470
4.650
Appendix
Table 11: Representative G-Eval results under different values of α on the Scholarly Title task. Moderate anchoring achieves the highest overall professional-quality score, whereas an excessively large anchoring coefficient leads to performance degradation. Bold numbers denote the best result in each column.
Qwen3-1.7B
Qwen3-4B
Mode
ROUGE-1
METEOR
BERTScore
ROUGE-1
METEOR
BERTScore
Thinking off
0.2661
0.1437
0.8378
0.3334
0.1985
0.8467
Thinking on
0.2908
0.1616
0.8351
0.3220
0.1847
0.8340
Δ
+0.0247
+0.0179
−0.0027
−0.0114
−0.0138
−0.0127
Appendix
Table 12: Thinking-mode ablation on LongLaMP. Values are macro-averaged across Abstract Generation, Product Review, and Topic Writing. Δ denotes the change from thinking off to thinking on; positive values indicate improvement.
Large language models (LLMs) are typically aligned with population-level preferences, despite substantial variation across individual users. We introduce POPI, a user-level personalization framework that separates the problem into two components connected by a natural-language interface: a shared inference model that distills heterogeneous user signals into a concise preference summary, and a shared generator that conditions on this summary to produce personalized responses. Both components are trained under a unified preference-optimization objective, with reinforcement learning handling the non-differentiable inference step. This objective decomposes into generator approximation error and summary informativeness, revealing how a single loss simultaneously drives accurate generation and informative summarization. Because the interface is natural language, learned summaries can be inferred once per user and reused across different generators -- including frozen, black-box commercial APIs. Across four personalization benchmarks, POPI generally improves personalization quality while reducing context overhead by up to an order of magnitude.
Yizhuo Chen, Xin Liu, Ruijie Wang +7
University of Illinois Urbana-Champaign · Amazon · University of Notre Dame
Large Language Models (LLMs) exhibit strong implicit personalization ability, yet most existing approaches treat this behavior as a black box, relying on prompt engineering or fine tuning on user data. In this work, we adopt a mechanistic interpretability perspective and hypothesize the existence of a sparse set of Preference Heads, attention heads that encode user specific stylistic and topical preferences and exert a causal influence on generation. We introduce Differential Preference Steering (DPS), a training free framework that (1) identifies Preference Heads through causal masking analysis and (2) leverages them for controllable and interpretable personalization at inference time. DPS computes a Preference Contribution Score (PCS) for each attention head, directly measuring its causal impact on user aligned outputs. During decoding, we contrast model predictions with and without Preference Heads, amplifying the difference between personalized and generic logits to selectively strengthen preference aligned continuations. Experiments on widely used personalization benchmarks across multiple LLMs demonstrate consistent gains in personalization fidelity while preserving content coherence and low computational overhead. Beyond empirical improvements, DPS provides a mechanistic explanation of where and how personalization emerges within transformer architectures. Our implementation is publicly available.
Weixu Zhang, Ye Yuan, Changjiang Han +7
McGill University · Mila - Quebec AI Institute · MBZUAI +2
Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. However, most existing personalization approaches rely on implicit representations within model parameters, making it difficult to interpret user-specific preferences or effectively handle long-context dependencies. To address these challenges, we propose PrefReward, a novel preference-aware generative framework that explicitly models user styles through a structured preference matrix and integrates it into the decoding process as a reward signal. PrefReward consists of two stages: (1) extracting a user-specific preference matrix that summarizes individual stylistic tendencies, and (2) using the matrix to guide generation via a KL-divergence-based reward function. Experiments on the LongLaMP dataset show that PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability.
Yue Wu, Chengbing Wang, Yimeng Bai +3
University of Science and Technology of China · The Chinese University of Hong Kong · National University of Singapore