OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue
Organizations: Fudan University · Shanghai Innovation Institute
Abstract
Maintaining persona consistency across multi-turn dialogues remains a core challenge for role-playing language models. Off-policy distillation from external teachers incurs distribution mismatch that compounds across dialogue turns, while reinforcement learning struggles with reward ambiguity inherent in subjective persona fidelity. We propose OSPD, an on-policy self-distillation framework where the same model serves as both teacher and student under asymmetric information: the teacher receives a complete character profile while the student sees only a brief summary, and the student generates trajectories from its own policy. We find that teacher confidence in role-playing dialogue exhibits a bimodal structure---sharply peaked at character-critical tokens yet diffuse at generic utterances---and introduce role-aware divergence switching to match this structure. A progressive trait masking curriculum further forces staged internalization of character knowledge along semantic dimensions. Experiments on CharacterBench, CharacterEval, and SocialBench show that OSPD substantially improves persona consistency over supervised fine-tuning and multi-turn RL baselines, without requiring any external teacher or reward model.
Figures & tables
| Method | CB | CharacterEval (1–5) | SocialBench (%) | IR | ||
| Avg. | PC | Avg. | SA | Avg. | ||
| SFT | 3.00 | 2.90 | 3.12 | 59.4 | 56.8 | 0.71 |
| SFT + DPO | 3.15 | 3.08 | 3.24 | 61.5 | 58.5 | 0.74 |
| PCL † | 3.19 | 3.06 | 3.22 | 61.7 | 58.9 | 0.73 |
| Multi-turn RL † | 3.39 | 3.47 | 3.52 | 65.5 | 62.4 | 0.78 |
| CPO † | 3.28 | 3.33 | 3.41 | 66.0 | 63.2 | 0.76 |
| Variant | CB Avg | CE PC |
| OSPD (full) | 3.51 | 3.49 |
| w/o on-policy (off-policy traj.) | 3.29 | 3.22 |
| Fixed Forward KL | 3.37 | 3.36 |
| Fixed Reverse KL | 3.38 | 3.36 |
| w/o trait masking curriculum | 3.41 | 3.37 |
| w/o PCV filtering | 3.47 | 3.40 |
| Masked Dimension | CB Avg | CE PC | IR |
| Full mask (default) | 3.51 | 3.49 | 0.86 |
| Knowledge only | 3.48 | 3.44 | 0.84 |
| Style only | 3.44 | 3.40 | 0.82 |
| History only | 3.41 | 3.37 | 0.80 |
| None (Stage 1 only) | 3.41 | 3.35 | 0.79 |
| Method | In-genre | Cross-genre | ||
| CB | CB | |||
| SFT | 2.93 | 0.08 | 2.84 | 0.17 |
| Multi-turn RL | 3.25 | 0.15 | 3.10 | 0.29 |
| Off-pol. Distill. | 3.19 | 0.09 | 3.07 | 0.20 |
| OSPD | 3.44 | 0.08 | 3.32 | 0.20 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Dimension | |||
| Identity | 85.3 | 80.1 | 5.2 |
| Personality | 78.6 | 65.8 | 12.8 |
| Knowledge | 72.4 | 43.9 | 28.5 |
| Style | 76.2 | 57.9 | 18.3 |
| History | 68.5 | 42.8 | 25.7 |
| Overall | 76.2 | 58.1 | 18.1 |
| Task Type | BIC | Sep. | ||
| Role-playing dialogue | +842 | 0.31 | 2.15 | 1.84 |
| Math reasoning | +127 | 0.18 | 0.85 | 0.67 |
| Open-domain chat | +53 | 1.68 | 2.82 | 1.14 |
| Hyperparameter | Value |
| Model & LoRA | |
| Base model | Qwen2.5-7B-Instruct |
| LoRA rank / / dropout | 64 / 16 / 0.1 |
| LoRA targets | q, k, v, o_proj |
| Optimization | |
| Optimizer | AdamW ( ) |
| Component | Hrs | % |
| On-policy generation | 8.2 | 44.1 |
| Teacher eval. (dual fwd.) | 5.1 | 27.4 |
| Divergence & gradient | 3.8 | 20.4 |
| PCV filtering | 0.9 | 4.8 |
| EMA & checkpointing | 0.6 | 3.2 |
| Total | 18.6 | 100 |
| Stage | Student Input |
| Stage 1 | Sherlock Holmes, consulting detective at 221B Baker Street, companion Dr. Watson. Analytical, arrogant, dramatic. Expert in chemistry and forensics. Speaks precisely with rhetorical questions and dry humor. Survived Reichenbach Falls confrontation with Moriarty. |
| Stage 2 | Sherlock Holmes is a consulting detective at 221B Baker Street. Works with Dr. Watson. Analytical, intellectually arrogant, and socially detached. |
| Teacher Update Strategy | CB Avg. | CE PU |
| EMA =0.990 | 3.43 | 3.39 |
| EMA =0.995 (default) | 3.51 | 3.48 |
| EMA =0.999 | 3.46 | 3.42 |
| Periodic ( =50) | 3.39 | 3.35 |
| Naive shared | 3.26 | 3.19 |
| PCV Variant | CB Avg. | CE PU | Accept |
| GPT-5.1 judge (default) | 3.51 | 3.48 | 72.3% |
| Embedding MLP | 3.48 | 3.41 | 78.5% |
| No PCV | 3.47 | 3.38 | 100% |
| CB Avg. | CE PU | IR | Fluency | |
| 0 | 3.38 | 3.29 | 0.83 | 72.4 |
| 0.1 | 3.49 | 3.46 | 0.85 | 85.3 |
| 0.5 | 3.51 | 3.48 | 0.86 | 91.7 |
| 1.0 | 3.48 | 3.44 | 0.84 | 93.2 |
| CB Avg. | CE PU | |||
| 0.5 | 3.41 | 3.38 | 0.49 | 0.42 |
| 1.0 | 3.51 | 3.48 | 0.51 | 0.34 |
| 2.0 | 3.48 | 3.44 | 0.50 | 0.25 |
| 5.0 | 3.41 | 3.37 | 0.50 | 0.12 |
| Scale | CharacterBench | CharacterEval | SocialBench | IR | ||||
| Attr. | Behav. | Avg. | PU | PB | Know. | Style | ||
| SFT Baselines | ||||||||
| 7B | 3.12 | 2.94 | 3.00 | 2.97 | 2.84 | 57.2 | 61.5 | 0.71 |
| 14B | 3.29 | 3.13 | 3.19 | 3.16 | 3.01 | 61.5 | 64.8 | 0.75 |
| OSPD | ||||||||
| 7B | 3.58 | 3.47 | 3.51 | 3.48 | 3.49 | 67.3 | 67.9 | 0.86 |
| Method | Hrs | GB | Ext. T | CB |
| SFT | 4.2 | 42 | × | 3.00 |
| SFT+DPO | 7.8 | 48 | × | 3.15 |
| Multi-turn RL | 38.5 | 78 | × | 3.39 |
| Off-pol. Dist. | 12.8 | 72 | ✓ | 3.27 |
| OSPD | 18.6 | 52 | × | 3.51 |
| Dimension | SFT | Multi-turn RL | OSPD |
| Identity | 86.3 | 89.5 | 92.1 |
| Personality | 74.8 | 80.2 | 85.7 |
| Knowledge | 58.2 | 68.5 | 78.3 |
| Style | 65.3 | 74.8 | 81.5 |
| History | 51.7 | 62.3 | 72.8 |
| Average | 67.3 | 75.1 | 82.1 |
| Stage | Low | Mid | High | Bicoeff. |
| S1, Ep. 1 start | 32.1 | 28.5 | 39.4 | 0.61 |
| S1, Ep. 1 end | 35.8 | 23.2 | 41.0 | 0.67 |
| S2, Ep. 1 end | 38.2 | 20.1 | 41.7 | 0.72 |
| S2, Ep. 2 end | 38.5 | 19.8 | 41.7 | 0.73 |
| Type | Low % | High % | Bicoeff. | Sep. |
| Literary classic | 41.2 | 39.5 | 0.75 | 1.92 |
| Contemp. fiction | 36.8 | 43.7 | 0.71 | 1.78 |
| Antagonist | 43.5 | 37.2 | 0.78 | 2.05 |
| Historical | 39.1 | 40.8 | 0.73 | 1.85 |
| Sparse profile | 27.5 | 48.3 | 0.54 | 1.12 |
| Dialogue Topic | Bicoeff. | ||
| Knowledge Q&A | 0.32 | 0.28 | 0.76 |
| Value/moral judgment | 0.25 | 0.24 | 0.79 |
| Emotional exchange | 0.45 | 0.31 | 0.65 |
| Narrative recall | 0.38 | 0.30 | 0.70 |
| Casual chat | 0.68 | 0.22 | 0.52 |
| Turns | OSPD | Hist. | |
| 10 | 88.2 | 87.9 | 0.3 |
| 20 | 86.5 | 85.8 | 0.7 |
| 40 | 84.1 | 82.3 | 1.8 |
| 60 | 83.0 | 79.5 | 3.5 |
| Method | CB Avg (1–5) | CE PC (1–5) | ||||
| Base (no FT) | 3.81 | 2.91 | 0.90 | 3.75 | 2.80 | 0.95 |
| SFT | 4.23 | 3.00 | 1.23 | 4.15 | 2.90 | 1.25 |
| Multi-turn RL | 4.35 | 3.39 | 0.96 | 4.30 | 3.47 | 0.83 |
| Off-pol. Dist. | 4.25 | 3.27 | 0.98 | 4.12 | 3.16 | 0.96 |
| OSPD | 4.08 | 3.51 | 0.57 | 4.10 | 3.49 | 0.61 |
| Condition | CB | CE PC | IR | |
| Full (default, 2 000 tok) | 3.51 | 3.49 | 0.86 | +0.51 |
| Dimension removal | ||||
| Knowledge | 3.42 | 3.38 | 0.84 | +0.42 |
| Knowledge, History | 3.35 | 3.28 | 0.82 | +0.35 |
| Knowledge, History, Style | 3.18 | 3.12 | 0.78 | +0.18 |
| Length truncation | ||||
| Method | CB | CE PC | SB SA | IR |
| SFT | 2.91 | 2.82 | 57.8 | 0.69 |
| Multi-turn RL | 3.28 | 3.35 | 63.7 | 0.76 |
| Off-pol. Dist. | 3.17 | 3.08 | 62.3 | 0.75 |
| OSPD | 3.40 | 3.38 | 65.9 | 0.84 |
| Comparison | Win | Tie | Lose | |
| OSPD vs. SFT | 62.8 | 18.4 | 18.8 | 0.71 |
| OSPD vs. Multi-turn RL | 43.2 | 28.4 | 28.4 | 0.65 |
| Criterion | Win | Tie | Lose |
| Persona consistency | 42.0 | 30.8 | 27.2 |
| Knowledge accuracy | 48.4 | 26.0 | 25.6 |
| Linguistic style | 44.8 | 24.4 | 30.8 |
| Overall preference | 43.2 | 28.4 | 28.4 |
| Masking Strategy | CB | CE PC | IR |
| Default: mask K+S+H | 3.51 | 3.49 | 0.86 |
| Alt-1: mask K+P (keep I+S+H) | 3.43 | 3.40 | 0.83 |
| Alt-2: mask S+H (keep I+P+K) | 3.47 | 3.45 | 0.85 |
| Alt-3: random order per batch | 3.44 | 3.41 | 0.83 |