Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Figures & tables
Figure 1: Three judges of explanation quality. Each judge sees only what it requires: The utility judge answers the task query from the explanation; the source-support judge checks the explanation’s claims with the source; and the writing judge assesses only the explanation.
Figure 2: Better reconstruction accompanies more confabulation and writing defects. As the FVE reconstruction rises (solid line), so does the confabulation index (dashed line) across models. The index I(t)=21[R(t)/R(0)+D(t)/D(0)] averages assertion risk R and writing defects per explanation D relative to supervised initialization ( I=1 ). Shading shows 95% bootstrap intervals.
Figure 3: Better reconstruction accompanies more relevant information. Hard-negative separation across the 12 classification datasets for point-reconstruction NLA (solid) and Flow-NLA (dashed); higher is better.
Utility ↑
Confabulation ↓
Writing ↓
Target model
Class.
Suffix
Behavior
Risk
Coverage
Defects
Released Anthropic NLAs
Qwen2.5-7B
79.4
98.0
82.7
92.6
74.8
2.45
Gemma-3-12B
84.7
96.6
83.4
83.1
67.5
1.78
Gemma-3-27B
85.7
98.2
84.1
76.3
55.4
1.92
Llama-3.3-70B
82.7
99.0
65.1
76.4
56.9
1.78
Point-reconstruction reruns
Qwen2.5-7B
79.7
97.6
74.9
83.7
68.3
1.63
Table 1: Utility and confabulation in released NLAs and our reruns. All scores are percentages except writing defects, which are counts per explanation.
Figure 4: Denoising can reward information beyond the conditional mean. H is uniform over four unit directions; we follow the realized activation H=(1,0) (gold outline). (a) The horizontal explanation zhor keeps the horizontal pair (blue) and excludes the vertical pair (open circles). Both pairs have mean zero ( × ), so the optimal expected point loss is the same with or without zhor . (b) Denoising also observes the noisy activation xt=atH+btϵ (black arrow: noise btϵ ), with candidates scaled by at . (c) Without the explanation, xt is about equally close to the true and upward candidates, so the optimal estimate hedges between them. (d) With zhor , the upward candidate is excluded and the estimate moves to the true direction, lowering the denoising error. In (c)–(d), arrows show the displacement atEϕ[H∣xt,z]−xt from each noisy observation to its optimal signal estimate (without z in (c)); shading shows the density of xt .
Figure 5: Learning activation explanations with Flow-NLA. An activation h from a frozen target model is passed to a verbalizer, which generates an explanation z , and corrupted with Gaussian noise to form xt . The flow reconstructor learns to denoise xt conditioned on z , modeling the distribution of activations compatible with the explanation. Denoising errors across noise levels define a reward derived from a conditional diffusion likelihood bound. This reward trains the verbalizer to produce explanations that make the observed activation more predictable under the conditional model.
Figure 6: Changes in utility, confabulation, and writing defects from supervised initialization for point-reconstruction NLA (solid) and Flow-NLA (dashed). Markers show raw measurements, curves show three-checkpoint means, and shading shows pointwise 95% bootstrap intervals.
Figure 7: Changes from supervised initialization in mean utility, assertion risk, token coverage, and writing defects, averaged over the second half of the training range shared by both methods. Filled circles denote NLA; open squares denote Flow-NLA; lines show the gap between methods.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Utility ↑
Confabulation ↓
Writing ↓
Target model
Method
Class.
Suffix
Behavior
Risk
Coverage
Defects
Qwen2.5-7B
Released
79.4
98.0
82.7
92.6
74.8
2.45
Point recon.
79.7
97.6
74.9
83.7
68.3
1.63
Flow-NLA
80.3
96.6
72.9
71.9
54.7
0.99
Gemma-3-12B
Released
84.7
96.6
83.4
83.1
67.5
1.78
Point recon.
84.2
96.2
82.4
77.8
44.2
1.38
Appendix
Table 2: Utility and confabulation across NLA training methods. We compare released NLAs, our point-reconstruction reruns, and Flow-NLA within the same target model. All scores are percentages except writing defects, which are counts per explanation. Best results are highlighted within each target model. Checkpoint selection is described in Appendix B.2 .
Target model
Revision or release
Hidden state
Ours
Qwen2.5-7B-Instruct
a09a354
hidden_states[21]
Gemma-3-12B-it
96b6f1e
hidden_states[33]
Apertus-v1.5-8B
a411d83
hidden_states[22]
Released
Qwen2.5-7B-Instruct
nla-qwen2.5-7b-L20
hidden_states[21]
Gemma-3-12B-it
nla-gemma3-12b-L32
hidden_states[33]
Gemma-3-27B-it
nla-gemma3-27b-L41
hidden_states[42]
Appendix
Table 3: Activation extraction sites. Revisions are Hugging Face commit prefixes; released NLAs are identified by their repository names. A released NLA at layer K reads the output of decoder block K , i.e., hidden_states[ K+1 ] .
Task
Judge’s task
Examples
Options
Metric
Classification
Source properties (12 tasks)
1,344
2–14
Mean accuracy
Suffix selection
True 32-token continuation
500
10
Accuracy
Response behavior
Answer, refuse, or clarify
248
3
Macro-F1
Hard negatives
Contrastive question on source pair
1,316 (658 pairs)
2
Log-prob. ratio
Appendix
Table 4: Overview of our various utility tasks
Figure 8: Changes in utility, confabulation, and writing defects from supervised initialization for the point-reconstruction NLA reruns. Shading shows pointwise 95% bootstrap intervals.
Figure 9: Matched predictive gain, mismatch penalty, and total hard-negative separation across the 12 classification datasets for point-reconstruction NLA (solid) and Flow-NLA (dashed). Curves are three-checkpoint means over raw markers; shading shows 95% bootstrap intervals.
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are reconstruction-dependent, so the score tracks gist, not specific facts. Under exact synthetic ground truth, the standard recipe develops co-adapted private codes (false wording the reconstruction depends on) in 5/5 runs, and fixes that leave the target model unchanged do not help. We contribute two audit protocols, the grounded-vs-true cross and the evaluator swap, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors): linear heads trained alongside the target model to keep designated content decodable. On RECAP-trained sandbox models, fresh verbalizers state the designated content truly and the codes vanish, at a +0.001-nat cost. This replicates on a pretrained Pythia-160M: the content becomes reliably probe-decodable, though a fresh verbalizer conveys it only in part (truth 0.44-0.46 vs a near-zero control). For interpretability, high reconstruction does not certify individual claims. For AI safety, RECAP makes designated internal content independently checkable against probes rather than asserted by prose a model can game: an independent probe scores the verbalizer's true claims above its false ones (AUC 0.96, vs 0.82 without RECAP). Against an adversary that edits an explanation to maximize the reconstruction score while lying (suppressing ~87% of its lie penalty), the RECAP probe still flags the lies (AUC 0.95) while the control probe collapses to chance (0.51).
Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, without requiring a model forward pass for position selection. On three of four datasets, explaining just 5% of positions retains nearly all of the success rate from explaining every position, where success means obtaining an explanation about the threat. The benefit varies with the audit task. We also show that pretrained verbalizers recover words that models have learned to conceal through fine-tuning, without additional verbalizer training. These results identify where auditors can concentrate explanation generation and show that useful explanations can extend beyond the model a verbalizer was trained to describe.
Federico Torrielli, Gianluca Barmina, Andrea Blasi Núñez +4
Department of Computer Science, University of Turin, Italy · Department of Mathematics and Computer Science, University of Southern Denmark, Denmark
Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own activations. We introduce Universal Activation Verbalizer (UAV), a framework that uses a shared decoder to explain activations from heterogeneous donor models. UAV learns a lightweight adapter that converts donor activations into soft tokens in decoder's embedding space, and further supports adapter-only transfer by reusing a frozen decoder-side LoRA while training only a new adapter for another donor. Across classification, fact retrieval, and gist summarization, UAV remains competitive with strong self-explanation baselines while enabling cross-model verbalization across model families and scales. Ablations show that decoder-side tuning mainly improves task behavior, whereas the adapter provides the activation-grounded factual and semantic information needed for faithful explanations. Code and data are available at https://github.com/hy-zhao23/ActExp.
Haiyan Zhao, Zirui He, Guanchu Wang +3
New Jersey Institute of Technology · University of North Carolina at Charlotte · Cisco Research +1