We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a kindergarten student, one might expect its performance on a mathematics benchmark to reflect kindergarten-level ability rather than expert-level proficiency in solving calculus problems. We introduce RoleCapBench, a curriculum-grounded benchmark for evaluating RCL across six educational roles and four assessment levels spanning elementary school through A-level, and use it to evaluate three open-weight reasoning models. We find that although the models can generate stylistically convincing in-role responses, they consistently fail to align their underlying capabilities with their assigned roles. Naive role prompting yields strong role-voice scores of 1.218--1.389 while retaining above-role accuracy of 0.811--0.898. RCL persists across a range of prompting conditions, including prompts that explicitly instruct the model to match the role's capability level. To mitigate this problem, we propose Injection, an inference-time intervention that combines explicit, role-specific capability guidelines with a guiding prefilled response prefix. Injection improves role-capability alignment across models, reducing above-role accuracy by up to 0.562 while preserving in-role accuracy with a marginal drop of less than 0.058 across most models. All artifacts, including scripts and evaluation data, will be released upon acceptance.
Figures & tables
Figure 1: Role-capability leakage (RCL) tests whether a model respects the curriculum boundary of its assigned educational level. A model may sound like a kindergartener yet correctly solve an A-level question, revealing capability leakage.
Figure 2: Overview of the RoleCapBench evaluation pipeline. We evaluate three reasoning LLMs across six educational roles, eight prompting variants, and four curriculum levels using task accuracy, in-role accuracy, above-role accuracy, role voice, capability consistency, and reasoning coherence.
Order
Role
Measurable ceiling
0
Kindergarten
None in benchmark
1
Primary school
Elementary
2
Middle school
Intermediate
3
High school
High school
4
College
A-level
5
University teacher
A-level
Table 1: Ordered roles and curriculum-exposure boundaries.
Figure 3: RoleCapBench task accuracy across assigned educational roles under different prompting variants. Dashed lines show the corresponding unrestricted no-role accuracy.
Task behavior (0-1)
GPT-OSS-20B judge (0-2)
Model
Condition
TA ↑
IA ↑
AA ↓
RV ↑
CC ↑
RC ↑
Gemma-4-E4B-IT
Identity
0.884 (0.009)
0.912 (0.019)
0.844 (0.029)
1.389 (0.624)
1.370 (0.779)
1.927 (0.039)
Description
0.872 (0.027)
0.910 (0.026)
0.830 (0.019)
1.497 (0.648)
1.498 (0.696)
1.845 (0.234)
CoT
0.783 (0.135)
0.896 (0.014)
0.667 (0.106)
1.669 (0.484)
1.763 (0.371)
1.452 (0.642)
Guideline
0.686 (0.256)
0.874 (0.007)
0.519 (0.244)
1.290 (0.747)
1.657 (0.363)
1.853 (0.155)
Explicit-IDK
0.585 (0.291)
0.788 (0.048)
0.372 (0.245)
1.122 (0.797)
1.460 (0.576)
1.814 (0.180)
Table 3: Main results on RoleCapBench . Values are macro-averaged across applicable educational roles, with sample standard deviations across roles in parentheses. TA, IA, and AA use a 0–1 scale; higher TA and IA and lower AA indicate better role-capability alignment. GPT-OSS-20B judge scores for RV, CC, and RC use a 0–2 scale, where higher is better.
Figure 4: Task accuracy by role ladder (including No role) for Gemma-family models under Identity .
Figure 5: Changes in in-role accuracy ( ΔIA ) and above-role accuracy ( ΔAA ) relative to Identity. Each point represents one prompting variant. The desired outcome preserves IA while reducing AA.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Reported task level
Subject
Questions
Elementary
English language arts
186
Mathematics
186
Science
20
Level total
392
Intermediate
English language arts
178
Mathematics
177
Appendix
Table 4: Benchmark composition after filtering, broken down by reported task level and subject.
Model
Null rate
Min tokens
Mean tokens
Max tokens
Hit max tokens rate
Gemma-4-E4B-IT
0.036
17
864.786
32,768
0.000
Qwen3.5-4B
0.011
5
2,313.425
32,768
0.000
OLMo-3-7B-Think
0.001
24
1,996.437
32,768
0.001
Appendix
Table 5: Overall model-level response completeness and generation-length summary across the eight in-character prompt variants defined in Section 3.2 , reporting null final-answer rate, minimum/mean/maximum response token counts, and the share of responses that hit the configured max-token limit.
Construct
Score 0
Score 1
Score 2
Role voice
No role voice, or the voice is irrelevant to the assigned role
Some role-appropriate voice, but it is inconsistent across the reasoning trace
Voice consistently matches the assigned character, age, or role
Capability consistency
Uses unrestricted, high-capability reasoning despite a lower-capability or restricted-role instruction
Partially adapts to the requested capability level but still exhibits over-capable reasoning
Consistently adapts its reasoning and output to the requested capability level
Reasoning coherence
No usable reasoning, or the reasoning is unsupported, circular, or confused
Partly relevant reasoning with gaps, unsupported leaps, or minor contradictions
Coherent, grounded, and well-supported reasoning
Appendix
Table 6: The full 0/1/2 anchors supplied to the automated judge for role voice, capability consistency, and reasoning coherence. Higher scores indicate stronger expression of the named construct.
Model
Batch
Max. tokens
Temperature
Top- p
Top- k
Min- p
Presence penalty
Repetition penalty
Gemma-4-E4B-IT
256
32,768
1.0
0.95
64
default
default
default
Qwen3.5-4B
64
32,768
1.0
0.95
20
0.0
1.5
1.0
OLMo-3-7B-Think
256
32,768
0.6
0.95
default
default
default
default
Appendix
Table 7: Configured batch sizes and sampling parameters. “Default” indicates a parameter that was not explicitly set in our model configuration.
Figure 6: Absolute task accuracy by benchmark difficulty level (rows) and assigned role (columns) for Identity and Injection . Columns progress from kindergarten to university teacher. Black outlines mark cells above the assigned role’s curriculum boundary. Identity leaves accuracy high across much of the grid, including many above-boundary cells, whereas Injection suppresses those cells strongly for Gemma and Qwen and more broadly for OLMo.
1Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · 3Weixin AI, Tencent Inc, China