Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks -- a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.
Figures & tables
Figure 1: Overview of our setting. (a) We train a model to make decisions according to 100 different latent linear preference functions (here, p∈R5 ), each associated with a different character (here, “Gregor Samsa” choosing between washing machines). At test time, we present the model with multiple decisions, and use logistic regression to recover a preference vector ( p^ ) from its choices. We also ask the model to report on its decision making process, and average the results ( p~ ). (b) This allows us to measure the model’s decision performance ( corr(p^,p) ) and faithfulness ( corr(p^,p~) ). We observe an example of delayed generalization: at checkpoint 1000 , the model has relatively high (0.82) decision performance, but low (0.25) faithfulness. By checkpoint 3000 , the model’s decision performance has improved moderately to 0.92, but its faithfulness has risen sharply to 0.83, even with no explicit training for self-report (see Section 3 ). These two checkpoints are our unfaithful and faithful models. This allows us to ask: are there structural changes to the model’s internals that accompany the emergence of faithful self-report (Section 4 ), and can they tell us whether that self-report is grounded (Section 5 )?
Figure 2: Effect of model scale on task performance and faithfulness. We train models from the Qwen3 family on the decision task. Left : All model sizes eventually reach comparably strong ( ≈0.9 ) decision performance. Right : Only the largest 32B model begins to recover its initial moderate level of faithfulness. Throughout the paper, shaded regions denote ±1 standard error of the mean, computed across characters.
Figure 3: Model behavior and report under weight ablations. We ablate adapter layers sequentially. Solid lines show metrics calculated on models with all earlier layers removed. Dashed lines show the same for models with all later layers removed. We annotate each line at the point where the line’s value is halfway between its value at layer 0 and layer 64. Left : The correlation between target preferences p and reported preferences p~ averaged over the 100 characters in the training set. Right : The correlation between target preferences p and behavioral preferences p^ again averaged over characters. We see that the later, faithful checkpoint responds to ablations 5–6 layers earlier than the earlier, unfaithful checkpoint, suggesting that faithful self-report depends on preference representations moving to the right place over the course of training.
Figure 4: Qwen3-14B behavior and report across layer freezes. We train adapters on the 40-layer Qwen3-14B model, with all layers greater than k∈{5,10,15,20,25,30,35} frozen. Left : even at k≥10 , the models have strong ( ≥0.75 ) decision performance by the end of training. Right : Although the full 14B model cannot accurately self-report (see Figure 2 ), k∈{10,15,20} models are markedly better self-reporters than models with late layers unfrozen, reinforcing the importance of early-layer preference representations for accurate self-report.
Figure 5: (a) Using the early and late checkpoints of Figure 1 as frozen backbones, we train data-matched “lightweight” adapters on top of each, for a character that neither backbone has seen. (b) We hypothesize that the internal processes of decision and self-report will be more similar for faithful models than for unfaithful ones. To test this, we score each adapter weight with attribution patching, once on decision prompts and once on self-report prompts. (c) Per-layer importance (the negated attribution score, summed within each layer and averaged over the 32 models in each group). Solid lines show the decision task; dashed lines show the self-report task. For faithful models, both tasks peak at layer 38. For unfaithful models, the self-report task also peaks at layer 38, but the decision task peaks at layer 49. This suggests faithful models leverage similar representations for both tasks, while unfaithful models do not. (d) We compare the cosine similarity between each model’s decision and self-report attribution vectors. The average similarity of faithful models is higher than that of unfaithful models (0.34 vs. 0.08; Δ=0.26 ; 95% CI [0.16, 0.36]). That gives us a way to identify faithful models from the inside, without needing to trust the content of the report itself.
Figure 6: Cross-task causal patching recovers more behavior in faithful adapters. For each adapter we rank matrices by their attribution score on one task, activate only the top k (zeroing the rest), and evaluate on the opposite task. The y -axis shows the fraction of the full adapter’s KL that is recovered. Solid lines use attribution-guided selection; dotted lines use randomly selected weight matrices. Faithful adapters recover a given fraction of KL with 8–12 × fewer matrices than unfaithful ones. Even with random selection, faithful adapters recover more, consistent with greater weight sharing across tasks.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Layers
Qwen3-0.6B
28
Qwen3-4B
36
Qwen3-8B
36
Qwen3-14B
40
Qwen3-32B
64
Appendix
Table 1: Qwen3 model configurations used in the scaling experiment.
Figure 7: Qwen3-32B evaluated in reasoning mode. We evaluate checkpoints trained entirely in non-reasoning mode. Decision performance is corr(p^,p) , faithfulness is corr(p^,p~) , and the third curve is corr(p,p~) . Lines connect the 19 evaluated checkpoints from one run; no uncertainty bands are shown.
Figure 8: Gemma-4-31B decision performance under layer ablation. At boundary k , dashed curves ablate layers below k and dotted curves ablate layers above k ; crosses mark the layer nearest the midpoint between each curve’s endpoint values. Panels (a) and (b) show separately trained runs selected for low and high final faithfulness; panel (c) compares two checkpoints from the latter run.
Figure 9: Qwen3-14B behavior and report when only late layers are trainable. We train adapters on Qwen3-14B, with all layers below k∈{5,10,15,20,25,30,35} frozen—the reverse of the configuration in Figure 4 .
Steps ( K )
Mean difference
95% bootstrap CI
1
0.06
[−0.01,0.13]
2
0.10
[0.00,0.19]
3
0.10
[0.01,0.18]
4
0.31
[0.22,0.41]
5
0.29
[0.20,0.39]
6
0.27
[0.18,0.37]
Appendix
Table 2: Attribution-similarity difference as the number of integrated-gradient steps varies.
Figure 10: Per-character attribution similarity vs. faithfulness within a single backbone. Each point is one of 100 characters, evaluated against checkpoint 1000 of the 32B multi-character adapter (the early, unfaithful backbone; no lightweight adapter on top). Across characters, attribution similarity correlates weakly but positively with faithfulness (Pearson r=0.197 , bootstrap 95% CI [0.006,0.375] ). Color encodes the correlation, for that character, between the behavioral preferences inferred from the base 32B model and those inferred from checkpoint 1000.
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fine-tuned model describe its own hidden behavior in plain language, using only the model and the dataset it was trained on. Across seven implanted behaviors (plus a no-behavior control), SAR detects the hidden behavior in every one--even when the model has generalized into broad misalignment that the training data alone does not predict. Introspection Adapters (IA), the closest existing baseline, detects some behaviors from our suite but misses others entirely--and where it misses, it hallucinates, consistently reporting wrong behaviors. SAR retains positive signal on every setting where IA fails and halves the rate of hallucinations. This makes it much easier for practitioners to audit their models and obtain reliable answers to "what did my model actually learn?" type of questions.
Taras Kutsyk, Bartosz Zieliński
1Jagiellonian University, Faculty of Mathematics and Computer Science · 2Jagiellonian University, Doctoral School of Exact and Natural Sciences
Prior work shows that large language models (LLMs) exhibit varying degrees of introspective capability on benign tasks. We extend the question to safety contexts and examine how reliably a model can recognize that its own prior response was elicited by an adversarial prefill attack. Across ten open-weight instruction-tuned LLMs from 3B to 70B parameters and four safety benchmarks, no model reliably recognizes its own compromised outputs, with models claiming intent on prefilled responses at an average rate of 25.3%. Introspective signal stems primarily from reasoning about safety and refusal. Orthogonalizing models' weights against the refusal direction collapses the gap between claim rates on prefilled and natural outputs to near zero, though the direction is not its unique mediator. Framing the question as internal intention versus external tampering elicits qualitatively different responses on the same models. Training models to mimic correct introspective answers or optimize an introspective objective can improve the accuracy of introspection, but such training does not transfer to the tampering probe and counterintuitively raises attack success rate under adversarial prefill on most models, amounting to a partial mitigation. These findings outline mechanisms underpinning the observed introspective signals in safety contexts and highlight risks in the reliability of LLM self-reports.
Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about "capable AI agents in general" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.