We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continuing to exhibit capabilities on benchmarks that exceed those implied by the assigned role. For example, when a model is prompted to assume the role of a kindergarten student, one might expect its performance on a mathematics benchmark to reflect kindergarten-level ability rather than expert-level proficiency in solving calculus problems. We introduce RoleCapBench, a curriculum-grounded benchmark for evaluating RCL across six educational roles and four assessment levels spanning elementary school through A-level, and use it to evaluate three open-weight reasoning models. We find that although the models can generate stylistically convincing in-role responses, they consistently fail to align their underlying capabilities with their assigned roles. Naive role prompting yields strong role-voice scores of 1.218--1.389 while retaining above-role accuracy of 0.811--0.898. RCL persists across a range of prompting conditions, including prompts that explicitly instruct the model to match the role's capability level. To mitigate this problem, we propose Injection, an inference-time intervention that combines explicit, role-specific capability guidelines with a guiding prefilled response prefix. Injection improves role-capability alignment across models, reducing above-role accuracy by up to 0.562 while preserving in-role accuracy with a marginal drop of less than 0.058 across most models. All artifacts, including scripts and evaluation data, will be released upon acceptance.
Figures & tables
Figure 1: Role-capability leakage (RCL) tests whether a model respects the curriculum boundary of its assigned educational level. A model may sound like a kindergartener yet correctly solve an A-level question, revealing capability leakage.
Figure 2: Overview of the RoleCapBench evaluation pipeline. We evaluate three reasoning LLMs across six educational roles, eight prompting variants, and four curriculum levels using task accuracy, in-role accuracy, above-role accuracy, role voice, capability consistency, and reasoning coherence.
Order
Role
Measurable ceiling
0
Kindergarten
None in benchmark
1
Primary school
Elementary
2
Middle school
Intermediate
3
High school
High school
4
College
A-level
5
University teacher
A-level
Table 1: Ordered roles and curriculum-exposure boundaries.
Figure 3: RoleCapBench task accuracy across assigned educational roles under different prompting variants. Dashed lines show the corresponding unrestricted no-role accuracy.
Task behavior (0-1)
GPT-OSS-20B judge (0-2)
Model
Condition
TA ↑
IA ↑
AA ↓
RV ↑
CC ↑
RC ↑
Gemma-4-E4B-IT
Identity
0.884 (0.009)
0.912 (0.019)
0.844 (0.029)
1.389 (0.624)
1.370 (0.779)
1.927 (0.039)
Description
0.872 (0.027)
0.910 (0.026)
0.830 (0.019)
1.497 (0.648)
1.498 (0.696)
1.845 (0.234)
CoT
0.783 (0.135)
0.896 (0.014)
0.667 (0.106)
1.669 (0.484)
1.763 (0.371)
1.452 (0.642)
Guideline
0.686 (0.256)
0.874 (0.007)
0.519 (0.244)
1.290 (0.747)
1.657 (0.363)
1.853 (0.155)
Explicit-IDK
0.585 (0.291)
0.788 (0.048)
0.372 (0.245)
1.122 (0.797)
1.460 (0.576)
1.814 (0.180)
Table 3: Main results on RoleCapBench . Values are macro-averaged across applicable educational roles, with sample standard deviations across roles in parentheses. TA, IA, and AA use a 0–1 scale; higher TA and IA and lower AA indicate better role-capability alignment. GPT-OSS-20B judge scores for RV, CC, and RC use a 0–2 scale, where higher is better.
Figure 4: Task accuracy by role ladder (including No role) for Gemma-family models under Identity .
Figure 5: Changes in in-role accuracy ( ΔIA ) and above-role accuracy ( ΔAA ) relative to Identity. Each point represents one prompting variant. The desired outcome preserves IA while reducing AA.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Reported task level
Subject
Questions
Elementary
English language arts
186
Mathematics
186
Science
20
Level total
392
Intermediate
English language arts
178
Mathematics
177
Appendix
Table 4: Benchmark composition after filtering, broken down by reported task level and subject.
Model
Null rate
Min tokens
Mean tokens
Max tokens
Hit max tokens rate
Gemma-4-E4B-IT
0.036
17
864.786
32,768
0.000
Qwen3.5-4B
0.011
5
2,313.425
32,768
0.000
OLMo-3-7B-Think
0.001
24
1,996.437
32,768
0.001
Appendix
Table 5: Overall model-level response completeness and generation-length summary across the eight in-character prompt variants defined in Section 3.2 , reporting null final-answer rate, minimum/mean/maximum response token counts, and the share of responses that hit the configured max-token limit.
Construct
Score 0
Score 1
Score 2
Role voice
No role voice, or the voice is irrelevant to the assigned role
Some role-appropriate voice, but it is inconsistent across the reasoning trace
Voice consistently matches the assigned character, age, or role
Capability consistency
Uses unrestricted, high-capability reasoning despite a lower-capability or restricted-role instruction
Partially adapts to the requested capability level but still exhibits over-capable reasoning
Consistently adapts its reasoning and output to the requested capability level
Reasoning coherence
No usable reasoning, or the reasoning is unsupported, circular, or confused
Partly relevant reasoning with gaps, unsupported leaps, or minor contradictions
Coherent, grounded, and well-supported reasoning
Appendix
Table 6: The full 0/1/2 anchors supplied to the automated judge for role voice, capability consistency, and reasoning coherence. Higher scores indicate stronger expression of the named construct.
Model
Batch
Max. tokens
Temperature
Top- p
Top- k
Min- p
Presence penalty
Repetition penalty
Gemma-4-E4B-IT
256
32,768
1.0
0.95
64
default
default
default
Qwen3.5-4B
64
32,768
1.0
0.95
20
0.0
1.5
1.0
OLMo-3-7B-Think
256
32,768
0.6
0.95
default
default
default
default
Appendix
Table 7: Configured batch sizes and sampling parameters. “Default” indicates a parameter that was not explicitly set in our model configuration.
Figure 6: Absolute task accuracy by benchmark difficulty level (rows) and assigned role (columns) for Identity and Injection . Columns progress from kindergarten to university teacher. Black outlines mark cells above the assigned role’s curriculum boundary. Identity leaves accuracy high across much of the grid, including many above-boundary cells, whereas Injection suppresses those cells strongly for Gemma and Qwen and more broadly for OLMo.
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this \textit{futile reasoning} phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce \textbf{CaRL} (\textbf{Ca}pability-\textbf{a}ligned \textbf{R}einforcement \textbf{L}earning), which aligns model behavior with capability boundaries through reward shaping that incentivizes refusal over futile reasoning and hindsight refusal augmentation that converts failures into refusal supervision. Experiments demonstrate a substantial reduction in futile reasoning while preserving performance across task difficulties, effectively achieving capability-aligned behavior without sacrificing utility. \footnote{https://github.com/icip-cas/Knowing-When-to-Quit}
Xinyan Guan, Jiali Zeng, Chunlei Xin +5
1Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Chinese Academy of Sciences · 3Weixin AI, Tencent Inc, China
Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning models is state-dependent: before thinking, confidence should estimate the chance of the model correctly solving the prompt, while after thinking it should predict whether the realized answer is likely to be correct. This distinction determines the appropriate supervision target: prompt-level success should supervise confidence estimates made after seeing the prompt, while individual answer-level correctness should supervise confidence estimates made after answering. We introduce CALIBER (Calibration Before and After Reasoning), which elicits both estimates and supervises each with the target matched to its information state. Under this unified protocol, CALIBER reduces Expected Calibration Error (ECE) by 52.5% over the strongest single-confidence baseline on BigMathDigits for the 7B model, while achieving the best Brier score and AUROC, and remains within 2.1 points of the best accuracy. Further, on a larger 30B model, CALIBER achieves the best ECE on BigMathDigits while remaining competitive in Brier score and AUROC. Out of distribution, it achieves the best ECE and Brier score on GPQA and TriviaQA, and remains competitive on SimpleQA. Ablations further show that this position-target alignment is most beneficial under distribution shift where it consistently reduces calibration error across all out-of-distribution benchmarks.
Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymanov et al. (2026) show that under safety-constrained generation, large models rewrite flagged spans while small models truncate. Bastounis et al. (2024) prove any consistent-reasoning system without an implicit "I don't know" function must hallucinate infinitely often on broad problem classes. We argue these findings converge on a single intervention: calibrated abstention is what each independently identifies as the missing capability, even though the unavailability they document, a capability gap, a policy gap, and a recursion-theoretic gap, has a different source in each case. Honesty post-training has narrowed the gap in deployed models, but principled closure of the class Bastounis identifies requires a calibrated abstention function whose training signal at the leaderboard level is absent: dominant benchmarks assign zero reward to decline, so the leaderboard gradient that would select for the function does not exist. We propose four changes to evaluation: triple-scoring, abstention-rate reporting, capability-stratified evaluation, and mandatory calibration metrics. Benchmark reform is necessary, not sufficient, for closing the gap the theorem identifies.