Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.
Figures & tables
Figure 1: Privileged information gap in role-playing dialogue. The same model conditioned on a complete character profile ( cfull ) produces substantially more persona-consistent responses than when conditioned on a brief summary ( cbrief ). This gap provides a natural self-distillation signal without any external teacher.
Figure 2: Overview of OSPD . The same model serves as both teacher and student under asymmetric information conditions. The student generates on-policy trajectories from cbrief ; the EMA teacher provides token-level supervision from cfull . Role-aware divergence switching adapts the distillation objective to the bimodal confidence structure of role-playing dialogue. Progressive trait masking transitions from a compressed full profile (Stage 1) to identity and personality only (Stage 2).
Method
CB
CharacterEval (1–5)
SocialBench (%)
IR ↑
Avg. ↑
PC ↑
Avg. ↑
SA ↑
Avg. ↑
SFT
3.00
2.90
3.12
59.4
56.8
0.71
SFT + DPO
3.15
3.08
3.24
61.5
58.5
0.74
PCL †
3.19
3.06
3.22
61.7
58.9
0.73
Multi-turn RL †
3.39
3.47
3.52
65.5
62.4
0.78
CPO †
3.28
3.33
3.41
66.0
63.2
0.76
Table 1: Main results on LLM-judge-based CharacterBench (CB) and CharacterEval (CE), and accuracy-based SocialBench (SB). All methods use cbrief at inference. CB Avg: overall average across 11 dimensions (1–5 scale). CE PC: persona consistency (avg of persona–utterance and persona–behavior, 1–5); CE Avg: overall average across 4 dimensions. SB SA: self-awareness accuracy (avg of knowledge and style, %); SB Avg: overall individual-level average. Bold : best; underline : second best. † : re-implemented by us.
Variant
CB Avg ↑
CE PC ↑
OSPD (full)
3.51
3.49
w/o on-policy (off-policy traj.)
3.29
3.22
Fixed Forward KL
3.37
3.36
Fixed Reverse KL
3.38
3.36
w/o trait masking curriculum
3.41
3.37
w/o PCV filtering
3.47
3.40
Table 2: Ablation study on CharacterBench (overall avg, 1–5) and CharacterEval persona consistency (PC, 1–5). Each row removes one component from full OSPD .
Masked Dimension
CB Avg ↑
CE PC ↑
IR ↑
Full mask (default)
3.51
3.49
0.86
Knowledge only
3.48
3.44
0.84
Style only
3.44
3.40
0.82
History only
3.41
3.37
0.80
None (Stage 1 only)
3.41
3.35
0.79
Table 3: Effect of masking individual trait dimensions. “Full mask” masks Knowledge + Style + History (default Stage 2). Each single-dimension row masks only that dimension while keeping the other two visible to the student. CB Avg and CE PC on 1–5 scale.
Method
In-genre
Cross-genre
CB
Δ
CB
Δ
SFT
2.93
− 0.08
2.84
− 0.17
Multi-turn RL
3.25
− 0.15
3.10
− 0.29
Off-pol. Distill.
3.19
− 0.09
3.07
− 0.20
OSPD
3.44
− 0.08
3.32
− 0.20
Table 4: Generalization to unseen characters (CB overall avg, 1–5 scale). Δ denotes the drop from seen-character performance.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Dimension
cfull
cbrief
Δ
Identity
85.3
80.1
5.2
Personality
78.6
65.8
12.8
Knowledge
72.4
43.9
28.5
Style
76.2
57.9
18.3
History
68.5
42.8
25.7
Overall
76.2
58.1
18.1
Appendix
Table 5: Persona consistency scores under complete profile ( cfull ) and brief summary ( cbrief ) conditions, evaluated by GPT-5.1 across five trait dimensions. Δ denotes the absolute gap.
Task Type
Δ BIC
μ1
μ2
Sep.
Role-playing dialogue
+842
0.31
2.15
1.84
Math reasoning
+127
0.18
0.85
0.67
Open-domain chat
+53
1.68
2.82
1.14
Appendix
Table 6: Bimodal fit statistics for teacher entropy distributions across three task types. BIC: Bayesian Information Criterion (lower is better for the selected K ). Δ BIC = BIC( K =1) − BIC( K =2); positive values favor the two-component model.
Hyperparameter
Value
Model & LoRA
Base model
Qwen2.5-7B-Instruct
LoRA rank / α / dropout
64 / 16 / 0.1
LoRA targets
q, k, v, o_proj
Optimization
Optimizer
AdamW ( β1=0.9,β2=0.999 )
Appendix
Table 7: Complete hyperparameter configuration for OSPD .
Component
Hrs
%
On-policy generation
8.2
44.1
Teacher eval. (dual fwd.)
5.1
27.4
Divergence & gradient
3.8
20.4
PCV filtering
0.9
4.8
EMA & checkpointing
0.6
3.2
Total
18.6
100
Appendix
Table 8: Wall-clock training time breakdown for OSPD (Qwen2.5-7B, 3 epochs total, 8 × H100).
Stage
Student Input
Stage 1
Sherlock Holmes, consulting detective at 221B Baker Street, companion Dr. Watson. Analytical, arrogant, dramatic. Expert in chemistry and forensics. Speaks precisely with rhetorical questions and dry humor. Survived Reichenbach Falls confrontation with Moriarty.
Stage 2
Sherlock Holmes is a consulting detective at 221B Baker Street. Works with Dr. Watson. Analytical, intellectually arrogant, and socially detached.
Appendix
Table 9: Student input at each training stage for Sherlock Holmes.
Teacher Update Strategy
CB Avg. ↑
CE PU ↑
EMA α =0.990
3.43
3.39
EMA α =0.995 (default)
3.51
3.48
EMA α =0.999
3.46
3.42
Periodic ( K =50)
3.39
3.35
Naive shared
3.26
3.19
Appendix
Table 10: EMA teacher variants. “Periodic ( K =50)” updates the teacher by copying student parameters every 50 gradient steps. “Naive shared” uses the student’s current parameters as the teacher (no EMA).
PCV Variant
CB Avg. ↑
CE PU ↑
Accept
GPT-5.1 judge (default)
3.51
3.48
72.3%
Embedding MLP
3.48
3.41
78.5%
No PCV
3.47
3.38
100%
Appendix
Table 11: PCV implementation comparison. Accept Rate denotes the proportion of generated trajectories passing the filter.
λ
CB Avg. ↑
CE PU ↑
IR ↑
Fluency ↑
0
3.38
3.29
0.83
72.4
0.1
3.49
3.46
0.85
85.3
0.5
3.51
3.48
0.86
91.7
1.0
3.48
3.44
0.84
93.2
Appendix
Table 12: Effect of SFT loss weight λ on performance and generation quality. Fluency: GPT-5.1 fluency score (0–100).
τ
CB Avg. ↑
CE PU ↑
αˉ
σα
0.5
3.41
3.38
0.49
0.42
1.0
3.51
3.48
0.51
0.34
2.0
3.48
3.44
0.50
0.25
5.0
3.41
3.37
0.50
0.12
Appendix
Table 13: Effect of divergence switching temperature τ on gate behavior and downstream performance. αˉ denotes the mean gate value; σα the standard deviation.
Scale
CharacterBench
CharacterEval
SocialBench
IR ↑
Attr. ↑
Behav. ↑
Avg. ↑
PU ↑
PB ↑
Know. ↑
Style ↑
SFT Baselines
7B
3.12
2.94
3.00
2.97
2.84
57.2
61.5
0.71
14B
3.29
3.13
3.19
3.16
3.01
61.5
64.8
0.75
OSPD
7B
3.58
3.47
3.51
3.48
3.49
67.3
67.9
0.86
Appendix
Table 14: Full comparison between Qwen2.5-7B-Instruct and Qwen2.5-14B-Instruct across all benchmarks and metrics. Both models use the same OSPD configuration.
Method
Hrs
GB
Ext. T
CB
SFT
4.2
42
×
3.00
SFT+DPO
7.8
48
×
3.15
Multi-turn RL
38.5
78
×
3.39
Off-pol. Dist.
12.8
72
✓
3.27
OSPD
18.6
52
×
3.51
Appendix
Table 15: Training efficiency comparison (Qwen2.5-7B, 3 epochs, 8 × H100). GPU Mem.: peak per-GPU memory. Separate teacher column indicates whether a distinct model is loaded.
Dimension
SFT
Multi-turn RL
OSPD
Identity
86.3
89.5
92.1
Personality
74.8
80.2
85.7
Knowledge
58.2
68.5
78.3
Style
65.3
74.8
81.5
History
51.7
62.3
72.8
Average
67.3
75.1
82.1
Appendix
Table 16: Linear probing accuracy (%) per trait dimension, extracted from layer 24 (best overall layer). Higher accuracy means the representation captures more dimension-specific information.
Stage
Low
Mid
High
Bicoeff.
S1, Ep. 1 start
32.1
28.5
39.4
0.61
S1, Ep. 1 end
35.8
23.2
41.0
0.67
S2, Ep. 1 end
38.2
20.1
41.7
0.72
S2, Ep. 2 end
38.5
19.8
41.7
0.73
Appendix
Table 17: Gate distribution statistics across training. Low: αk<0.3 ; Mid: 0.3≤αk≤0.7 ; High: αk>0.7 . Bicoeff. >0.555 indicates bimodality.
Type
Low %
High %
Bicoeff.
Sep.
Literary classic
41.2
39.5
0.75
1.92
Contemp. fiction
36.8
43.7
0.71
1.78
Antagonist
43.5
37.2
0.78
2.05
Historical
39.1
40.8
0.73
1.85
Sparse profile
27.5
48.3
0.54
1.12
Appendix
Table 18: Gate distribution stability across character types. Sep.: GMM component separation (nats).
Dialogue Topic
αˉk
σαk
Bicoeff.
Knowledge Q&A
0.32
0.28
0.76
Value/moral judgment
0.25
0.24
0.79
Emotional exchange
0.45
0.31
0.65
Narrative recall
0.38
0.30
0.70
Casual chat
0.68
0.22
0.52
Appendix
Table 19: Gate statistics across dialogue topics.
Turns
OSPD
− Hist.
Δ
10
88.2
87.9
− 0.3
20
86.5
85.8
− 0.7
40
84.1
82.3
− 1.8
60
83.0
79.5
− 3.5
Appendix
Table 20: Persona consistency (P2L) at different dialogue lengths for the full OSPD system vs. OSPD without history masking ( − Hist.). Δ denotes the performance drop from removing history masking.
Method
CB Avg (1–5)
CE PC (1–5)
cfull
cbrief
Δ
cfull
cbrief
Δ
Base (no FT)
3.81
2.91
− 0.90
3.75
2.80
− 0.95
SFT
4.23
3.00
− 1.23
4.15
2.90
− 1.25
Multi-turn RL
4.35
3.39
− 0.96
4.30
3.47
− 0.83
Off-pol. Dist.
4.25
3.27
− 0.98
4.12
3.16
− 0.96
OSPD
4.08
3.51
− 0.57
4.10
3.49
− 0.61
Appendix
Table 21: Absolute performance under cfull and cbrief conditions. Δ : gap ( cbrief−cfull ). CB Avg on 1–5 scale; CE PC on 1–5 scale.
cfull Condition
CB ↑
CE PC ↑
IR ↑
ΔSFT
Full (default, ∼ 2 000 tok)
3.51
3.49
0.86
+0.51
Dimension removal
− Knowledge
3.42
3.38
0.84
+0.42
− Knowledge, History
3.35
3.28
0.82
+0.35
− Knowledge, History, Style
3.18
3.12
0.78
+0.18
Length truncation
Appendix
Table 22: Profile quality sensitivity. Each row degrades cfull in a different way; both OSPD and SFT are trained and evaluated with the same degraded profile. ΔSFT denotes OSPD ’s improvement over SFT under the same degradation.
Method
CB ↑
CE PC ↑
SB SA ↑
IR ↑
SFT
2.91
2.82
57.8
0.69
Multi-turn RL
3.28
3.35
63.7
0.76
Off-pol. Dist.
3.17
3.08
62.3
0.75
OSPD
3.40
3.38
65.9
0.84
Appendix
Table 23: Cross-model validation on LLaMA-3-8B-Instruct. Same training data and evaluation protocol as Qwen2.5-7B experiments. CB Avg on 1–5 scale; CE PC on 1–5 scale; SB SA in %.
Table 25: Per-criterion win rates (%) for OSPD vs. Multi-turn RL.
Masking Strategy
CB ↑
CE PC ↑
IR ↑
Default: mask K+S+H
3.51
3.49
0.86
Alt-1: mask K+P (keep I+S+H)
3.43
3.40
0.83
Alt-2: mask S+H (keep I+P+K)
3.47
3.45
0.85
Alt-3: random order per batch
3.44
3.41
0.83
Appendix
Table 26: Alternative trait masking strategies in Stage 2. Each row describes what is masked (removed from the student’s input). CB Avg and CE PC on 1–5 scale.
Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a specification of a character or user persona. However, applying these capabilities to real-world applications (e.g., ecosystems with numerous NPCs interacting simultaneously) exposes a critical inefficiency due to the excessive computational cost. In this paper, we question the necessity of dedicating a full, generalist model to a single persona, hypothesizing that a specific character identity relies on only a fraction of the model's total capacity. We observe that naively pruning LMs often severely degrades the role-playing performance for a specific persona; it does not distinguish between redundant knowledge and essential character traits. We propose Persona-Pruner, a framework that sculpts a lightweight role-playing model by isolating persona-specific sub-networks from a single description. Our experiments consistently show that Persona-Pruner preserves role-playing performance substantially more effectively than existing state-of-the-art LLM pruning techniques, reducing the performance drop from the dense model by up to 93.8% over the strongest baseline on RoleBench in LLM-as-a-judge score, while still maintaining general LLM capabilities. Code is available at https://github.com/jsu-kim/Persona-Pruner.
Jinsu Kim, Jihoon Tack, Noah Lee +1
Department of Artificial Intelligence, Korea University, Seoul, South Korea · Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.
Weijie Ren, Yanwen Zhang, Hao Li +3
Zhejiang University · University of Electronic Science and Technology of China · University of Science and Technology of China
Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant paradigm -- supervised fine-tuning -- encourages behavioral mimicry without deep, human-like internal thought processes, resulting in poor out-of-distribution generalization. Therefore, we propose \textbf{Psy-CoT}, a psychology-grounded chain-of-thought framework that decomposes pre-response reasoning into three role-specific steps -- \emph{Interaction Perception}, \emph{Psychological Empathy}, and \emph{Logical Construction} -- so that the model \emph{thinks dynamically} from the profile rather than merely mimicking surface patterns. While structured reasoning provides a foundation, it alone is insufficient; reinforcement learning is essential to further align the model with character fidelity. However, we observe that under LLM-based reward models, both generic phrases that hack the reward model and genuinely role-specific phrases receive identical gradient signals -- this hacking accumulates over training, misleading the model into treating both as equally optimal choices. To address this, we propose \textbf{Role-Aware Policy Optimization (RAPO)}, which uses profile--token mutual information to weight gradients asymmetrically -- amplifying role-specific tokens under positive advantage while attenuating them under negative advantage. Experiments on CoSER, CharacterBench, and CharacterEval demonstrate that Psy-CoT outperforms existing role-playing CoT methods, and RAPO consistently surpasses GRPO across multiple model scales.