Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.
Figures & tables
Figure 1: Overview of DAFI’s three-component interpretation target.
Figure 2: DAFI: (a) three interpretation metrics; (b) the agent loop with reusable skills.
Method
Input
Output
Functional
Tokens
score (%)
score (%)
score
(M) ↓
SAGE
78.3
—
—
0.796
Token Change
—
27.8
—
0.00484
Claude Code
Restricted ∗
84.0
35.0
2.8
3.082
Unrestricted ∗
94.0
54.8
4.0
13.296
Table 1: Baseline comparison. Input score is the harmonic mean of activation coverage and boundary rejection. Input and Output scores are percentage-scaled, whereas Functional score uses a 1–5 scale. Costs are millions of tokens per feature. ∗ denotes the 25-feature Claude Code evaluation; all other rows use 100 features. Dashes: inapplicable metrics. Blue: DAFI. ↓ : lower is better.
Evidence
Initial
Refined
Gain
Neuronpedia
82.1
91.4
+9.3
Short-context probing
72.7
88.5
+15.8
Table 2: Input scores before and after adaptive refinement with Neuronpedia or one-template short-context probing on 100 matched GemmaScope features. Initial and refined scores are reported as percentages; gain is their absolute difference in percentage points. Bold values are best; blue marks short-context probing.
Table 5: Representative Shift cases. * marks Neuronpedia’s np_acts-logits-general explanations (activation examples and top-logit evidence); unmarked rows use oai_token-act-pair explanations (activation examples only). Input, Output, and Functional are DAFI components; L/F denote layer/feature indices, and IDs link to Neuronpedia.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Failure signal
Diagnosis
Candidate repair
Rerun scope
Low input activation coverage
Positive prompts do not realize the proposed trigger
Narrow the hypothesis to the observed token pattern; rebuild positive prompts
Input design and scoring
Low input boundary rejection
The hypothesis also covers adjacent negative contexts
Relation is misstated or lacks sufficient evidence
Revise the functional hypothesis if evidence suffices; otherwise adjust context, intervention position or strength, then revise
Functional construction and judgment; intervention onward if new evidence is needed
Mixed regression
A repair improves one metric but weakens another
Restore the best prior trace; change repair family or stop
Failed component only
Appendix
Table 6: Metric-aware repair policy. The controller changes the smallest component supported by the observed failure and reuses all unaffected evidence.
Metric
Improved, n (%)
Unchanged, n (%)
Decreased, n (%)
Minimum input component
88 (35.2)
161 (64.4)
1 (0.4)
Output score
174 (69.6)
63 (25.2)
13 (5.2)
Functional score
138 (55.2)
96 (38.4)
16 (6.4)
Appendix
Table 7: Diagnostic changes from initial to selected adaptive interpretations for 250 features. The minimum input component is the smaller of activation coverage and boundary rejection.
Score
Criterion
1
No supported relation between the input and output hypotheses.
2
The proposed relation is contradicted by the available evidence.
3
The relation is plausible but only weakly grounded in the evidence.
4
A clear relation is supported by representative evidence from both endpoints.
5
The relation is strongly supported, specific, and free of salient contradictions.
Appendix
Table 8: Five-point rubric used by the functional judge.
Activation coverage
Boundary rejection
Output score
Functional score
Joint pass (%)
Layer
Init.
Adapt.
Init.
Adapt.
Init.
Adapt.
Init.
Adapt.
Init.
Adapt.
L0
0.920
0.960
0.936
0.960
0.230
0.682
2.920
3.940
10.0
84.0
L6
0.844
0.952
0.788
0.920
0.313
0.714
2.720
3.860
6.0
74.0
L12
0.840
0.948
0.804
0.900
0.276
0.664
2.940
3.920
6.0
74.0
L18
0.864
0.904
0.748
0.908
0.338
0.719
2.620
3.940
4.0
78.0
L24
0.864
0.940
0.736
0.940
0.468
0.697
3.680
4.280
14.0
90.0
Appendix
Table 9: Layer-wise Input-score components, scores, and joint pass rates for 250 features.
Features interpreted
0
20
40
60
80
100
120
Joint pass (%)
58.0
82.0
80.0
84.0
86.0
92.0
92.0
Mean agent-loop rounds
4.00
2.60
3.20
2.70
2.64
2.70
2.42
Appendix
Table 10: Held-out joint pass rate and mean refinement rounds during skill accumulation.
Initialization
Activation coverage (%)
Boundary rejection (%)
Initial
Final
Initial
Final
Neuronpedia
83.8
91.6
80.4
91.2
Short-context probing
68.2
84.8
77.8
92.6
Appendix
Table 11: Input-score component breakdown for Neuronpedia and one-template short-context-probing initialization on 100 matched GemmaScope features.
Component
Scans
Sequences
Input token positions
Shared initialization
1
255,994
511,988
Follow-up: full vocabulary
5
1,279,970
4,863,886
Follow-up: random sample
18
855,000
3,410,000
Follow-up: manual candidates
6
71
270
Follow-up subtotal
29
2,135,041
8,274,156
Short-context total
30
2,391,035
8,786,144
Appendix
Table 12: Input volume for short-context probing in the finalized 100-feature experiment. Initialization is shared across targets; follow-up scans count unique successful calls in the final trajectories.
Template
Candidates
Active
Max. act.
<bos>{token}
255,994
0
0.00
<bos>Despite {token} (random)
50,000
4
7.75
<bos>Despite {token} (full)
255,994
39
9.75
Appendix
Table 13: Context discovery for Layer 18, Feature 3450. Both full scans cover the same vocabulary.
Statistic
Value
Inter-annotator plausibility Spearman
0.865
Inter-annotator plausibility quadratic-weighted κ
0.841
Plausibility-rating agreement within one point (%)
92.5
Inter-annotator pass agreement (%)
92.5
Inter-annotator pass Cohen’s κ
0.826
Relation-label agreement (%)
70.0
Appendix
Table 14: Independent human validation on 40 stratified features. Both system comparisons use the same feature-level Functional scores and are reported separately for the two annotators. Agreement within one rating point and pass agreement compare the annotators; ratings of at least 4 pass.
Metric
Output score
Functional score
Pearson correlation
0.783
0.710
Spearman correlation
0.811
0.649
Mean absolute error
0.129
0.531
Pass-decision agreement (%)
87.1
80.4
Appendix
Table 15: Cross-model agreement between GPT-5.6-Sol and the original DeepSeek-V4-Pro judgments on the final traces of 250 GemmaScope features. Pass-decision agreement uses the fixed thresholds of 0.5 for Output score and 4 for Functional score.
Condition
Layer
n
Input score Activation
Input score Boundary
Output score
Functional score
Mean rounds
Cold
0
20
0.78
0.87
0.446
3.20
1.40
Cold
7
20
0.54
0.86
0.176
3.15
1.50
Cold
14
20
0.68
0.71
0.518
3.10
1.30
Cold
21
20
0.62
0.81
0.371
2.80
1.75
Cold
27
20
0.90
0.80
0.573
3.35
1.85
Cold
All
100
0.704
0.810
0.417
3.12
1.56
Appendix
Table 16: Input-score components and layer-level transfer results on Qwen-Scope. Each layer contains 20 sampled features; All gives the 100-feature aggregate.
Method
Activation coverage (%)
Boundary rejection (%)
SAGE
84.6
72.8
Token Change
—
—
Claude Code, restricted ∗
84.0
84.0
Claude Code, unrestricted ∗
97.6
90.7
DAFI
91.6
91.2
Appendix
Table 17: Input-score component breakdown for the main baseline comparison. ∗ denotes the 25-feature Claude Code evaluation; all other rows use 100 features. Dashes indicate inapplicable metrics.
Method
Joint passes
Input score (%)
Output score (%)
Functional score
Tokens / feature (M) ↓
Activation coverage
Boundary rejection
Claude Code, restricted
2/25
84.0
84.0
35.0
2.8
3.082
Claude Code, unrestricted
19/25
97.6
90.7
54.8
4.0
13.296
Appendix
Table 18: Results for the restricted and unrestricted Claude Code configurations on the same 25 features. Activation coverage, boundary rejection, and Output score are percentage-scaled; Output score is not a probability. Functional score uses a 1–5 scale, and costs are millions of tokens per feature. The restricted setting uses a round-start budget, not a hard cap.
Signal
Evidence
Role
MaxAct
Activating texts
Input initialization
VocabProj
Top unembedded tokens
Output prior only
Token Change
Intervention deltas
Output evidence
Functional interpretation
htin , htout , htc
Final judgment
Appendix
Table 19: Output-side signals used by related work and by our dual-end evaluator.
Version
Scoring idea
Failure addressed
V1
Top positives
Broad-label recall
V2
Positives and negatives
Vague output patterns
V3
Per-feature calibration
Layer and scale drift
Appendix
Table 20: Output-score revisions and the failures that motivated them.
Layer
DAFI
Token Change
L0
0.637
0.108
L6
0.651
0.270
L12
0.617
0.194
L18
0.696
0.355
L24
0.733
0.461
Overall
0.667
0.278
Appendix
Table 21: Layer-wise Output score comparison between DAFI and Token Change on 100 matched features, with 20 features per layer. SD is the population standard deviation across the five layer means; lower values indicate more consistent performance across layers.
Varied metric
Threshold
Initial
Adaptive
Gain
Input components
0.7
8.0
80.0
+72.0
Input components
0.8
8.0
80.0
+72.0
Input components
0.9
6.4
49.6
+43.2
Output score
0.4
11.2
80.8
+69.6
Output score
0.5
8.0
80.0
+72.0
Output score
0.6
4.4
62.8
+58.4
Appendix
Table 22: One-at-a-time threshold sensitivity over 250 features. The default input-component threshold is 0.8 for both activation coverage and boundary rejection; the default Output and Functional score thresholds are 0.5 and 4. Rates are percentages.
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several directions. Pointer-like features, by contrast, predominantly exhibit diffuse effects. Our results show that a feature can be interpretable and causally relevant without providing a stable direction for steering.
Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2
UKP Lab, Technical University of Darmstadt · Department of Electrical Engineering Indian Institute of Technology Delhi, India
Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested. Single-token features that activate on one vocabulary item provide the diagnostic case where ground truth permits direct comparison. We analyze 3.9M features across six models and three SAE families using zero-ablation at full layer depth. Single-token features cluster 4.7x tighter in decoder space and concentrate in early layers (Layer 0 in GPT2-Small; L0-L4 in Gemma). Ablating them yields Benjamini-Hochberg-significant logit reductions in 178 of 208 full-layer conditions, with depth controlling whether damage cascades downstream or shapes the output directly. Cross-family causal differences exceed within-family scale effects: on the same base model, GemmaScope and BatchTopK features remain causally anchored, while LlamaScope features are locally redundant. The target token's rank recovers to within 2x baseline 96-98% of the time after the same ablation, and a controlled activation-function comparison reverses sign within the same model, leaving training recipe as the residual candidate. Cross-family interpretability claims are therefore sensitive to training methodology, not just activation function or scale.
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale. We introduce SAEVerbalizer, a framework that injects SAE decoder directions into an LLM's representations and fine-tunes the LLM's downstream layers to generate natural-language explanations of the injected features. Once trained, the resulting verbalizer explains SAE features directly from decoder directions, addressing both limitations. Our experiments show that the learned verbalization capability generalizes to unseen features, transfers across separately trained SAE dictionaries, and, with a lightweight adapter, extends to SAE features from different LLMs. Intervention experiments show that injecting multiple directions yields an explanation combining their meanings, while reversing individual directions produces corresponding meaning shifts.
Weihan Meng, Hongzhu Guo, Yi Jing +5
Tsinghua University · Peking University · Fudan University