Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for predicting model behavior while also increasingly introducing unsupported details and exhibiting writing defects. To assess these changes separately, we introduce a standardized evaluation framework for unstructured NLA explanations, measuring information recoverable from explanations, contextual support for their claims, and writing quality. To address confabulation and writing defects, we move beyond predicting a single activation: explanations can distinguish distributions of possible activations even when their means and optimal point-reconstruction rewards are identical. We introduce Flow-NLA, which models the distribution of activations compatible with an explanation and trains the verbalizer using a diffusion likelihood bound. Across Qwen, Gemma, and Apertus, this richer signal retains the utility gains of point reconstruction while curbing the growth of confabulation and writing defects, opening up a direction for improving activation-derived training to encourage more informative, supported, and readable explanations. Code and evaluation prompts will be made publicly available upon acceptance.
Figures & tables
Figure 1: Three judges of explanation quality. Each judge sees only what it requires: The utility judge answers the task query from the explanation; the source-support judge checks the explanation’s claims with the source; and the writing judge assesses only the explanation.
Figure 2: Better reconstruction accompanies more confabulation and writing defects. As the FVE reconstruction rises (solid line), so does the confabulation index (dashed line) across models. The index I(t)=21[R(t)/R(0)+D(t)/D(0)] averages assertion risk R and writing defects per explanation D relative to supervised initialization ( I=1 ). Shading shows 95% bootstrap intervals.
Figure 3: Better reconstruction accompanies more relevant information. Hard-negative separation across the 12 classification datasets for point-reconstruction NLA (solid) and Flow-NLA (dashed); higher is better.
Utility ↑
Confabulation ↓
Writing ↓
Target model
Class.
Suffix
Behavior
Risk
Coverage
Defects
Released Anthropic NLAs
Qwen2.5-7B
79.4
98.0
82.7
92.6
74.8
2.45
Gemma-3-12B
84.7
96.6
83.4
83.1
67.5
1.78
Gemma-3-27B
85.7
98.2
84.1
76.3
55.4
1.92
Llama-3.3-70B
82.7
99.0
65.1
76.4
56.9
1.78
Point-reconstruction reruns
Qwen2.5-7B
79.7
97.6
74.9
83.7
68.3
1.63
Table 1: Utility and confabulation in released NLAs and our reruns. All scores are percentages except writing defects, which are counts per explanation.
Figure 4: Denoising can reward information beyond the conditional mean. H is uniform over four unit directions; we follow the realized activation H=(1,0) (gold outline). (a) The horizontal explanation zhor keeps the horizontal pair (blue) and excludes the vertical pair (open circles). Both pairs have mean zero ( × ), so the optimal expected point loss is the same with or without zhor . (b) Denoising also observes the noisy activation xt=atH+btϵ (black arrow: noise btϵ ), with candidates scaled by at . (c) Without the explanation, xt is about equally close to the true and upward candidates, so the optimal estimate hedges between them. (d) With zhor , the upward candidate is excluded and the estimate moves to the true direction, lowering the denoising error. In (c)–(d), arrows show the displacement atEϕ[H∣xt,z]−xt from each noisy observation to its optimal signal estimate (without z in (c)); shading shows the density of xt .
Figure 5: Learning activation explanations with Flow-NLA. An activation h from a frozen target model is passed to a verbalizer, which generates an explanation z , and corrupted with Gaussian noise to form xt . The flow reconstructor learns to denoise xt conditioned on z , modeling the distribution of activations compatible with the explanation. Denoising errors across noise levels define a reward derived from a conditional diffusion likelihood bound. This reward trains the verbalizer to produce explanations that make the observed activation more predictable under the conditional model.
Figure 6: Changes in utility, confabulation, and writing defects from supervised initialization for point-reconstruction NLA (solid) and Flow-NLA (dashed). Markers show raw measurements, curves show three-checkpoint means, and shading shows pointwise 95% bootstrap intervals.
Figure 7: Changes from supervised initialization in mean utility, assertion risk, token coverage, and writing defects, averaged over the second half of the training range shared by both methods. Filled circles denote NLA; open squares denote Flow-NLA; lines show the gap between methods.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Utility ↑
Confabulation ↓
Writing ↓
Target model
Method
Class.
Suffix
Behavior
Risk
Coverage
Defects
Qwen2.5-7B
Released
79.4
98.0
82.7
92.6
74.8
2.45
Point recon.
79.7
97.6
74.9
83.7
68.3
1.63
Flow-NLA
80.3
96.6
72.9
71.9
54.7
0.99
Gemma-3-12B
Released
84.7
96.6
83.4
83.1
67.5
1.78
Point recon.
84.2
96.2
82.4
77.8
44.2
1.38
Appendix
Table 2: Utility and confabulation across NLA training methods. We compare released NLAs, our point-reconstruction reruns, and Flow-NLA within the same target model. All scores are percentages except writing defects, which are counts per explanation. Best results are highlighted within each target model. Checkpoint selection is described in Appendix B.2 .
Target model
Revision or release
Hidden state
Ours
Qwen2.5-7B-Instruct
a09a354
hidden_states[21]
Gemma-3-12B-it
96b6f1e
hidden_states[33]
Apertus-v1.5-8B
a411d83
hidden_states[22]
Released
Qwen2.5-7B-Instruct
nla-qwen2.5-7b-L20
hidden_states[21]
Gemma-3-12B-it
nla-gemma3-12b-L32
hidden_states[33]
Gemma-3-27B-it
nla-gemma3-27b-L41
hidden_states[42]
Appendix
Table 3: Activation extraction sites. Revisions are Hugging Face commit prefixes; released NLAs are identified by their repository names. A released NLA at layer K reads the output of decoder block K , i.e., hidden_states[ K+1 ] .
Task
Judge’s task
Examples
Options
Metric
Classification
Source properties (12 tasks)
1,344
2–14
Mean accuracy
Suffix selection
True 32-token continuation
500
10
Accuracy
Response behavior
Answer, refuse, or clarify
248
3
Macro-F1
Hard negatives
Contrastive question on source pair
1,316 (658 pairs)
2
Log-prob. ratio
Appendix
Table 4: Overview of our various utility tasks
Figure 8: Changes in utility, confabulation, and writing defects from supervised initialization for the point-reconstruction NLA reruns. Shading shows pointwise 95% bootstrap intervals.
Figure 9: Matched predictive gain, mismatch penalty, and total hard-negative separation across the 12 classification datasets for point-reconstruction NLA (solid) and Flow-NLA (dashed). Curves are three-checkpoint means over raw markers; shading shows 95% bootstrap intervals.