Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while input-side evidence collection typically relies on costly large-corpus scans. We introduce functional interpretation, which characterizes an SAE feature as a mapping from its activating input semantics to its output effects under intervention, and present Dual-End Agentic Feature Interpretation (DAFI), an agent that actively gathers evidence and refines input-side, output-side, and functional interpretations through component-specific feedback. Its short-context token probing enables on-demand activation evidence collection without a full corpus scan. On GemmaScope, DAFI improves Input score by 13.1 percentage points over SAGE and Output score by 38.9 points over Token Change, while being substantially more token-efficient than a general-purpose coding agent. Skills distilled from successful refinements raise the held-out joint pass rate from 58.0% to 92.0% and improve both interpretation quality and efficiency when transferred to a new model-SAE setting. Across features with reliable endpoint interpretations, 70.7% exhibit non-equivalent input and output semantics. On AxBench, DAFI also improves steering-feature selection over output-score filtering. Code is available at https://github.com/THUAIS-Lab/DAFI.
Figures & tables
Figure 1: Overview of DAFI’s three-component interpretation target.
Figure 2: DAFI: (a) three interpretation metrics; (b) the agent loop with reusable skills.
Method
Input
Output
Functional
Tokens
score (%)
score (%)
score
(M) ↓
SAGE
78.3
—
—
0.796
Token Change
—
27.8
—
0.00484
Claude Code
Restricted ∗
84.0
35.0
2.8
3.082
Unrestricted ∗
94.0
54.8
4.0
13.296
Table 1: Baseline comparison. Input score is the harmonic mean of activation coverage and boundary rejection. Input and Output scores are percentage-scaled, whereas Functional score uses a 1–5 scale. Costs are millions of tokens per feature. ∗ denotes the 25-feature Claude Code evaluation; all other rows use 100 features. Dashes: inapplicable metrics. Blue: DAFI. ↓ : lower is better.
Evidence
Initial
Refined
Gain
Neuronpedia
82.1
91.4
+9.3
Short-context probing
72.7
88.5
+15.8
Table 2: Input scores before and after adaptive refinement with Neuronpedia or one-template short-context probing on 100 matched GemmaScope features. Initial and refined scores are reported as percentages; gain is their absolute difference in percentage points. Bold values are best; blue marks short-context probing.
Table 5: Representative Shift cases. * marks Neuronpedia’s np_acts-logits-general explanations (activation examples and top-logit evidence); unmarked rows use oai_token-act-pair explanations (activation examples only). Input, Output, and Functional are DAFI components; L/F denote layer/feature indices, and IDs link to Neuronpedia.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Failure signal
Diagnosis
Candidate repair
Rerun scope
Low input activation coverage
Positive prompts do not realize the proposed trigger
Narrow the hypothesis to the observed token pattern; rebuild positive prompts
Input design and scoring
Low input boundary rejection
The hypothesis also covers adjacent negative contexts
Relation is misstated or lacks sufficient evidence
Revise the functional hypothesis if evidence suffices; otherwise adjust context, intervention position or strength, then revise
Functional construction and judgment; intervention onward if new evidence is needed
Mixed regression
A repair improves one metric but weakens another
Restore the best prior trace; change repair family or stop
Failed component only
Appendix
Table 6: Metric-aware repair policy. The controller changes the smallest component supported by the observed failure and reuses all unaffected evidence.
Metric
Improved, n (%)
Unchanged, n (%)
Decreased, n (%)
Minimum input component
88 (35.2)
161 (64.4)
1 (0.4)
Output score
174 (69.6)
63 (25.2)
13 (5.2)
Functional score
138 (55.2)
96 (38.4)
16 (6.4)
Appendix
Table 7: Diagnostic changes from initial to selected adaptive interpretations for 250 features. The minimum input component is the smaller of activation coverage and boundary rejection.
Score
Criterion
1
No supported relation between the input and output hypotheses.
2
The proposed relation is contradicted by the available evidence.
3
The relation is plausible but only weakly grounded in the evidence.
4
A clear relation is supported by representative evidence from both endpoints.
5
The relation is strongly supported, specific, and free of salient contradictions.
Appendix
Table 8: Five-point rubric used by the functional judge.
Activation coverage
Boundary rejection
Output score
Functional score
Joint pass (%)
Layer
Init.
Adapt.
Init.
Adapt.
Init.
Adapt.
Init.
Adapt.
Init.
Adapt.
L0
0.920
0.960
0.936
0.960
0.230
0.682
2.920
3.940
10.0
84.0
L6
0.844
0.952
0.788
0.920
0.313
0.714
2.720
3.860
6.0
74.0
L12
0.840
0.948
0.804
0.900
0.276
0.664
2.940
3.920
6.0
74.0
L18
0.864
0.904
0.748
0.908
0.338
0.719
2.620
3.940
4.0
78.0
L24
0.864
0.940
0.736
0.940
0.468
0.697
3.680
4.280
14.0
90.0
Appendix
Table 9: Layer-wise Input-score components, scores, and joint pass rates for 250 features.
Features interpreted
0
20
40
60
80
100
120
Joint pass (%)
58.0
82.0
80.0
84.0
86.0
92.0
92.0
Mean agent-loop rounds
4.00
2.60
3.20
2.70
2.64
2.70
2.42
Appendix
Table 10: Held-out joint pass rate and mean refinement rounds during skill accumulation.
Initialization
Activation coverage (%)
Boundary rejection (%)
Initial
Final
Initial
Final
Neuronpedia
83.8
91.6
80.4
91.2
Short-context probing
68.2
84.8
77.8
92.6
Appendix
Table 11: Input-score component breakdown for Neuronpedia and one-template short-context-probing initialization on 100 matched GemmaScope features.
Component
Scans
Sequences
Input token positions
Shared initialization
1
255,994
511,988
Follow-up: full vocabulary
5
1,279,970
4,863,886
Follow-up: random sample
18
855,000
3,410,000
Follow-up: manual candidates
6
71
270
Follow-up subtotal
29
2,135,041
8,274,156
Short-context total
30
2,391,035
8,786,144
Appendix
Table 12: Input volume for short-context probing in the finalized 100-feature experiment. Initialization is shared across targets; follow-up scans count unique successful calls in the final trajectories.
Template
Candidates
Active
Max. act.
<bos>{token}
255,994
0
0.00
<bos>Despite {token} (random)
50,000
4
7.75
<bos>Despite {token} (full)
255,994
39
9.75
Appendix
Table 13: Context discovery for Layer 18, Feature 3450. Both full scans cover the same vocabulary.
Statistic
Value
Inter-annotator plausibility Spearman
0.865
Inter-annotator plausibility quadratic-weighted κ
0.841
Plausibility-rating agreement within one point (%)
92.5
Inter-annotator pass agreement (%)
92.5
Inter-annotator pass Cohen’s κ
0.826
Relation-label agreement (%)
70.0
Appendix
Table 14: Independent human validation on 40 stratified features. Both system comparisons use the same feature-level Functional scores and are reported separately for the two annotators. Agreement within one rating point and pass agreement compare the annotators; ratings of at least 4 pass.
Metric
Output score
Functional score
Pearson correlation
0.783
0.710
Spearman correlation
0.811
0.649
Mean absolute error
0.129
0.531
Pass-decision agreement (%)
87.1
80.4
Appendix
Table 15: Cross-model agreement between GPT-5.6-Sol and the original DeepSeek-V4-Pro judgments on the final traces of 250 GemmaScope features. Pass-decision agreement uses the fixed thresholds of 0.5 for Output score and 4 for Functional score.
Condition
Layer
n
Input score Activation
Input score Boundary
Output score
Functional score
Mean rounds
Cold
0
20
0.78
0.87
0.446
3.20
1.40
Cold
7
20
0.54
0.86
0.176
3.15
1.50
Cold
14
20
0.68
0.71
0.518
3.10
1.30
Cold
21
20
0.62
0.81
0.371
2.80
1.75
Cold
27
20
0.90
0.80
0.573
3.35
1.85
Cold
All
100
0.704
0.810
0.417
3.12
1.56
Appendix
Table 16: Input-score components and layer-level transfer results on Qwen-Scope. Each layer contains 20 sampled features; All gives the 100-feature aggregate.
Method
Activation coverage (%)
Boundary rejection (%)
SAGE
84.6
72.8
Token Change
—
—
Claude Code, restricted ∗
84.0
84.0
Claude Code, unrestricted ∗
97.6
90.7
DAFI
91.6
91.2
Appendix
Table 17: Input-score component breakdown for the main baseline comparison. ∗ denotes the 25-feature Claude Code evaluation; all other rows use 100 features. Dashes indicate inapplicable metrics.
Method
Joint passes
Input score (%)
Output score (%)
Functional score
Tokens / feature (M) ↓
Activation coverage
Boundary rejection
Claude Code, restricted
2/25
84.0
84.0
35.0
2.8
3.082
Claude Code, unrestricted
19/25
97.6
90.7
54.8
4.0
13.296
Appendix
Table 18: Results for the restricted and unrestricted Claude Code configurations on the same 25 features. Activation coverage, boundary rejection, and Output score are percentage-scaled; Output score is not a probability. Functional score uses a 1–5 scale, and costs are millions of tokens per feature. The restricted setting uses a round-start budget, not a hard cap.
Signal
Evidence
Role
MaxAct
Activating texts
Input initialization
VocabProj
Top unembedded tokens
Output prior only
Token Change
Intervention deltas
Output evidence
Functional interpretation
htin , htout , htc
Final judgment
Appendix
Table 19: Output-side signals used by related work and by our dual-end evaluator.
Version
Scoring idea
Failure addressed
V1
Top positives
Broad-label recall
V2
Positives and negatives
Vague output patterns
V3
Per-feature calibration
Layer and scale drift
Appendix
Table 20: Output-score revisions and the failures that motivated them.
Layer
DAFI
Token Change
L0
0.637
0.108
L6
0.651
0.270
L12
0.617
0.194
L18
0.696
0.355
L24
0.733
0.461
Overall
0.667
0.278
Appendix
Table 21: Layer-wise Output score comparison between DAFI and Token Change on 100 matched features, with 20 features per layer. SD is the population standard deviation across the five layer means; lower values indicate more consistent performance across layers.
Varied metric
Threshold
Initial
Adaptive
Gain
Input components
0.7
8.0
80.0
+72.0
Input components
0.8
8.0
80.0
+72.0
Input components
0.9
6.4
49.6
+43.2
Output score
0.4
11.2
80.8
+69.6
Output score
0.5
8.0
80.0
+72.0
Output score
0.6
4.4
62.8
+58.4
Appendix
Table 22: One-at-a-time threshold sensitivity over 250 features. The default input-component threshold is 0.8 for both activation coverage and boundary rejection; the default Output and Functional score thresholds are 0.5 and 4. Rates are percentages.