MIRROR: From Imitation to Internalization in LLM Personalization
Organizations: Baidu Inc.
Abstract
The demand for personalized LLMs is shifting from style imitation toward content quality. We investigate whether self-distillation can bridge this gap in existing fine-tuning paradigm. To address this limitation, we introduce MIRROR(Meta- personalization by Internalizing Reference-Revealed On-policy Reflections), a novel self-distillation framework that shifts LLM personalization from imitation toward preference internalization. First, we replace reference-token imitation with reference-revealed on-policy self-distillation, aligning the model's next-token distributions along its own generation trajectories with those of its reference-conditioned self, thereby internalizing user preferences rather than reproducing reference wording.Second, we introduce MIRROR-F, a focal plug-in that augments on-policy distributional alignment with selective supervision over informative reference tokens, thereby strengthening content generation while preserving user-specific expression. Across three personalized generation benchmarks, two model scales, and complementary reference-based and LLM-based evaluations, MIRROR and MIRROR-F achieve leading overall personalization performance and superior text quality, while exhibiting less catastrophic forgetting than SFT-based baselines on three unseen personalized generation tasks. The gains are consistent across model scales and application scenarios, translating to improved performance in LLM personalization tasks.
Figures & tables
| Datasets | Methods ( ) | Base | Retrieval-based Methods | PEFT-based Methods | Ours | ||||||
| Metrics ( ) | Qwen3 | RAG | LatestK | LLM-TRSR | SFT | OPPU | PerCE | NextQuill | MIRROR | MIRROR-F | |
| Qwen3-1.7B | |||||||||||
| Abstract Generation | ROUGE-1 | 0.3717 | 0.3811 | 0.3685 | 0.3758 | 0.3755 | 0.3082 | 0.3670 | 0.3698 | 0.3967 | 0.4015 |
| METEOR | 0.2125 | 0.2262 | 0.2145 | 0.2311 | 0.2365 | 0.2073 | 0.2240 | 0.2437 | 0.2506 | 0.2529 | |
| BERTScore | 0.8577 | 0.8675 | 0.8657 | 0.8656 | 0.8668 | 0.8443 | 0.8621 | 0.8694 | 0.8587 | 0.8701 | |
| News Headline Generation | ROUGE-1 | 0.1422 | 0.1495 | 0.1402 | 0.1387 | 0.1468 | 0.1559 | 0.1527 | 0.1594 | 0.1674 | 0.1763 |
| Datasets ( ) | Abstract Generation | News Headline Generation | Paper Topic Generation | ||||||
| Methods ( ) | Content | Style | C&S | Content | Style | C&S | Content | Style | C&S |
| Qwen3-1.7B | |||||||||
| Base | 0.669 | 8.360 | 7.523 | 0.077 | 4.520 | 2.644 | 0.353 | 8.000 | 5.766 |
| SFT | 0.563 | 7.990 | 6.808 | 0.080 | 4.730 | 2.765 | 0.417 | 8.360 | 6.267 |
| PerCE | 0.492 | 7.810 | 6.365 | 0.068 | 4.810 | 2.745 | 0.413 | 8.100 | 6.116 |
| NextQuill | 0.542 | 8.000 | 6.708 | 0.078 | 4.670 | 2.725 | 0.392 | 8.140 | 6.028 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Content |
| Question | Generate an abstract for the title “Learning the Language: The Importance of Studying Written Directions in Designing Navigational Technologies for the Blind” using the following items: (1) Independent navigation; (2) Text-to-Speech directions; (3) Language and cues; (4) Qualitative analysis; and (5) Landmarks and alternate routes. |
| Gold answer | Independent navigation is important to individuals who are blind and visually impaired (VI). Researchers have long explored how blind and VI people navigate to inform the design of more useful, accessible wayfinding devices. However, there has been little research on the role language plays in providing effective text-to-speech directions for this population. Through qualitative analysis, this work examines the language and cues expressed in written navigational directions, including how distances are represented, how direction is indicated, and what landmarks are referenced. The analysis further identifies warnings about overshooting a destination, alternative routes that are easier to navigate, and the accessibility of destinations for people with disabilities. |
| Method | Answer |
| Base | This paper explores the importance of studying written directions in the design of navigational technologies for the blind, focusing on how independent navigation can be enhanced through effective use of text-to-speech directions. We conducted a qualitative analysis of how language and cues are used in written directions to guide users, emphasizing the need for clear, concise, and contextually relevant information. The study also examines the role of landmarks and alternate routes in providing users with multiple options for navigation, ensuring that they can make informed decisions in complex environments. Our findings highlight the significance of incorporating user-centered design principles that consider the unique needs and experiences of blind individuals when developing navigational technologies. |
| NextQuill | Independent navigation is a challenge for many people who are blind. One promising solution is to use text-to-speech directions to guide them. However, there has been little inquiry into how people who are blind use and interpret directions, and how technologies can be designed to better support them. We conducted a qualitative analysis of 11 blind participants’ experiences with and suggestions for text-to-speech directions. We found that users rely on language and cues to interpret directions, and that they had strong preferences for how these directions should be presented. We present two findings that can be used to guide the design of future text-to-speech directions: (1) the use of landmarks and alternate routes can be beneficial to users, and (2) users prefer directions that are presented in a conversational tone. |
| MIRROR | Independent navigation is a critical skill for individuals who are blind, enabling them to move confidently and safely in their environment. However, the effectiveness of navigational technologies for the blind often hinges on the clarity and usability of the directions provided. This paper explores the importance of studying written directions in the design of navigational technologies for the blind, focusing on how language and cues influence user experience. Through qualitative analysis of user interactions, we examine how text-to-speech directions are perceived and understood by users, highlighting the role of language in conveying information about landmarks and alternate routes. Our findings reveal that the way directions are structured and the cues used can significantly impact the user’s ability to navigate independently. We argue that a deeper understanding of language and cues in written directions is essential for designing more effective navigational technologies that support the unique needs of blind users. |
| MIRROR-F | Independent navigation for the blind is significantly influenced by the clarity and effectiveness of navigational technologies. This study emphasizes the importance of studying written directions in the design of navigational technologies. Through qualitative analysis, we explore how text-to-speech directions, language, and cues impact the user’s ability to navigate independently. The study highlights the role of language in conveying information and the necessity of clear, concise cues that aid in understanding the environment. We examine how landmarks and alternate routes are described and how these elements can be integrated into navigational systems to enhance usability. The findings underscore the need for a deep understanding of language in the design process, ensuring that navigational technologies are not only functional but also intuitive and accessible for users with visual impairments. |
| Item | Content |
| Question | Generate a headline for the following article: “Um, OK. For those of us whose leggy limbs haven’t been immortalized in print, we can turn to H&M’s updated classics, like …” |
| Gold answer | H&M Model Collection Inspired By Joan Smalls, Daphne Groeneveld & More (PHOTOS) |
| Method | Answer |
| PerCE | 7 Ways To Make Your Pantyhose Look Like A Model (PHOTOS) |
| NextQuill | Jenna Holmes’ ’90s Style Is Too Retro For A 2013 Cover (PHOTOS) |
| MIRROR | H&M’s Updated Classics: Leggy Limbs For Every Body Type (PHOTOS) |
| MIRROR-F | H&M’s Updated Classics: Leggy Limbs For Every Style (PHOTOS) |
| Item | Content |
| Question | Generate a title for the following abstract: “This paper addresses the conversion of discrete-time single input-single output PieceWise Affine (PWA) models from state space to input-output form. Necessary and sufficient conditions are given for a PWA state space model to admit equivalent input-output representations. When an equivalent input-output model exists, a constructive procedure is presented to derive both its parameters and the partition of the regressors domain. It is shown that the number of modes and the number of parameters may grow considerably when converting a PWA state space model into an equivalent input-output representation. Numerical examples highlight the role of the derived necessary and sufficient conditions for input-output realization of PWA state space models.” |
| Gold answer | Input-Output Realization Of Piecewise Affine State Space Models |
| Method | Answer |
| PerCE | On the Input-Output Representation of Piecewise Affine State Space Models |
| NextQuill | Input-output realization of piecewise affine state space models |
| MIRROR | Conversion of Discrete-Time Piecewise Affine Models to Input-Output Form |
| MIRROR-F | Conversion of Piecewise Affine State Space Models to Input-Output Form |
| Category | Hyperparameter | Value |
| Model and data | Backbone | Qwen3-1.7B/4B |
| Dataset | LongLaMP,LaMP-4,LaMP-5 | |
| Optimization | Optimizer | AdamW |
| Learning rate | ||
| Per-device batch size | 1 | |
| Gradient accumulation | 1 |
| Backbone | Task | Base | ContextSFT | PerCE | NextQuill | MIRROR(Ours) | MIRROR-F(Ours) |
| Qwen3-1.7B | Abstract | 0.933 | 0.895 | 0.785 | 0.839 | 0.949 | 0.952 |
| News | 0.774 | 0.447 | 0.404 | 0.358 | 0.387 | 0.631 | |
| Paper | 0.544 | 0.566 | 0.527 | 0.521 | 0.588 | 0.596 | |
| Qwen3-4B | Abstract | 0.951 | 0.919 | 0.854 | 0.819 | 0.957 | 0.963 |
| News | 0.686 | 0.458 | 0.416 | 0.461 | 0.676 | 0.640 | |
| Paper | 0.645 | 0.564 | 0.557 | 0.543 | 0.615 | 0.624 |
| Coherence | Consistency | Fluency | Relevance | Overall | |
| 0.1 | 4.480 | 4.810 | 4.980 | 4.580 | 4.713 |
| 0.3 | 4.440 | 4.830 | 5.000 | 4.550 | 4.705 |
| 0.7 | 4.500 | 4.880 | 4.990 | 4.570 | 4.735 |
| 0.9 | 4.440 | 4.690 | 5.000 | 4.470 | 4.650 |
| Qwen3-1.7B | Qwen3-4B | |||||
| Mode | ROUGE-1 | METEOR | BERTScore | ROUGE-1 | METEOR | BERTScore |
| Thinking off | 0.2661 | 0.1437 | 0.8378 | 0.3334 | 0.1985 | 0.8467 |
| Thinking on | 0.2908 | 0.1616 | 0.8351 | 0.3220 | 0.1847 | 0.8340 |
| +0.0247 | +0.0179 | |||||