D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders
Organizations: Pivotal Research · Mohamed bin Zayed University of Artificial Intelligence · Schmidt Sciences · Stanford University · University of Illinois at Urbana-Champaign
Abstract
Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence. D-Scope aggregates SigLIP2 embeddings of highly activating image patches into visual centroids. Matching target text descriptions against these visual centroids in the shared image-text embedding space then enables retrieval of individual features without per-feature text annotations. The underlying patches provide evidence for inspecting each selection, while spatially masked interventions test the corresponding decoder direction at varying strengths under fixed generation conditions. We characterize 150 SAEs across two model families and five layers, and introduce a benchmark of 100 target concepts with ten contexts each spanning under-specified and explicit-conflict conditions. Our empirical results show that high reconstruction fidelity can coexist with low dictionary utilization and limited visual-evidence coverage. Under per-case best-of-sweep strength selection, contrastive retrieval yields larger mean regional SigLIP2 gains than direct retrieval across the tested steering configurations, without consistently improving outside-region preservation. D-Scope provides an inspectable framework for evaluating sparse DiT features through their visual evidence and the effects of their decoder directions on generation. The demo is available at https://jiahaozhang-public.github.io/d-scope/.
Figures & tables
| SANA Student | Nitro Student | SANA Teacher | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Direct | Contrastive | Direct | Contrastive | Direct | Contrastive | ||||||||
| Method | Input | LPIPS | LPIPS | LPIPS | LPIPS | LPIPS | LPIPS | ||||||
| No intervention | – | ||||||||||||
| Oracle Prompt | – | ||||||||||||
| Oracle Dense | – | ||||||||||||
| L1 | None | ||||||||||||
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Decision | Empirical finding | Recommendation |
|---|---|---|
| Dictionary selection | High reconstruction EV can coexist with low utilization and evidence coverage (Fig. 3 a,b). The three weakest steering configurations in each generation–retrieval setting also have the smallest visual indices (Tables 1 and 7 ). | Assess utilization, coherence, and coverage alongside reconstruction fidelity. Validate steering performance separately, since a larger visual index does not guarantee stronger control. |
| Bias initialization | Centering increases utilization in every matched comparison. For L1 and JumpReLU, it expands visual indices by – and improves mean target gain in 22 of 24 matched comparisons (Fig. 3 c; Tables 1 and 7 ). | Start with mean or geometric-median centering, then verify target alignment and preservation for the chosen generator and retrieval strategy. |
| SAE family | Centered TopK has the highest median evidence coverage. Mean-centered L1 achieves the highest mean target gain in five of six generator–retrieval combinations (Fig. 3 d; Table 1 ). | Start with centered TopK for visual inspection and mean-centered L1 for steering, checking preservation in both cases. |
| Layer | Layer 26 has lower input-space EV and equal or higher coverage than layer 2 in matched comparisons. Intermediate trends are not uniformly monotonic (Fig. 6 ). | Compare candidate layers jointly on reconstruction, utilization, coherence, and coverage. |
| Retrieval | Contrastive retrieval yields higher mean target gains in all 45 matched comparisons, without consistently improving preservation (Section D.2 ). | Use contrastive retrieval as a starting point for target alignment, and assess LPIPS-out separately. |
| Dictionary reuse | SANA student dictionaries retain positive target gains in the teacher, with higher LPIPS-out than in the student (Table 1 ). | Re-evaluate target alignment and preservation when transferring dictionaries to another generation setting. |
| Setting | Configuration |
|---|---|
| Experimental grid | |
| Models | SANA-Sprint 0.6B; Nitro-1-PixArt 0.6B |
| Layers | 2, 8, 14, 20, 26 (0-indexed) |
| Activation | Post-block residual, one-step generation |
| Width | 1,152 input; 18,432 dictionary (16 ) |
| Input settings | |
| Layer | SAE | Input | EV (input) | EV (raw) | Mean | Util. (%) | Cov. (%) | DINO | SigLIP |
|---|---|---|---|---|---|---|---|---|---|
| 2 | L1 | None | 0.9824 | 0.9824 | 94.0 | 3.4 | 3.3 | 0.195 | 0.726 |
| Mean center | 0.9726 | 0.9726 | 115.0 | 24.5 | 25.0 | 0.195 | 0.732 | ||
| Geo. center | 0.9738 | 0.9738 | 121.4 | 25.1 | 20.1 | 0.189 | 0.731 | ||
| LayerNorm | 0.9823 | N/A | 95.3 | 3.4 | 4.9 | 0.188 | 0.724 | ||
| RMS | 0.9803 | 0.9803 | 81.7 | 2.9 | 1.1 | 0.190 | 0.720 | ||
| 2 | TopK | None | 0.9956 | 0.9956 | 64.0 | 17.7 | 12.0 | 0.193 | 0.730 |
| Layer | SAE | Input | EV (input) | EV (raw) | Mean | Util. (%) | Cov. (%) | DINO | SigLIP |
|---|---|---|---|---|---|---|---|---|---|
| 2 | L1 | None | 0.9922 | 0.9922 | 92.7 | 3.4 | 3.3 | 0.193 | 0.782 |
| Mean center | 0.9798 | 0.9798 | 67.3 | 9.9 | 7.6 | 0.195 | 0.781 | ||
| Geo. center | 0.9801 | 0.9801 | 67.0 | 10.0 | 12.5 | 0.192 | 0.784 | ||
| LayerNorm | 0.9909 | N/A | 83.3 | 2.9 | 3.3 | 0.192 | 0.780 | ||
| RMS | 0.9899 | 0.9899 | 84.4 | 3.4 | 3.3 | 0.186 | 0.777 | ||
| 2 | TopK | None | 0.9997 | 0.9997 | 64.0 | 1.5 | 0.5 | 0.155 | 0.767 |
| Component | Setting |
|---|---|
| Prompts | First 10,000 prompts of the evaluation split (Table 3 ), disjoint from training and validation prompts; the first 1,000, with the same seeds, form the dictionary-evaluation set |
| Generation | , single step, guidance 0 (timestep 400 for Nitro-1-PixArt); |
| Activations | Layer-14 post-block residual of all visual tokens ( for SANA-Sprint, for Nitro-1-PixArt), mapped by and encoded by the SAE |
| Evidence | images with the largest positive per-image maximum; one token per image; ties broken by image order |
| Patch | Square of four token cells centered on the selected token, clipped at the image boundary |
| Card | Patches in an grid by descending activation, for every feature with a positive image |
| SAE | Input | SANA-Sprint | Nitro-1-PixArt |
|---|---|---|---|
| L1 | None | 3,661 (19.9) | 1,583 (8.6) |
| Mean center | 15,754 (85.5) | 6,983 (37.9) | |
| Geo. center | 15,525 (84.2) | 6,976 (37.8) | |
| LayerNorm | 3,188 (17.3) | 1,422 (7.7) | |
| RMS | 3,432 (18.6) | 1,567 (8.5) | |
| TopK | None | 14,206 (77.1) | 12,411 (67.3) |
| Category | Targets | Cases | Example change |
|---|---|---|---|
| Color | 20 | 200 | Green dress red dress |
| Material | 20 | 200 | Wooden chair metal chair |
| Texture or pattern | 20 | 200 | Plain shirt striped shirt |
| Local appearance | 20 | 200 | Straight hair curly hair |
| Local state or geometry | 20 | 200 | Closed umbrella open umbrella |
| Total | 100 | 1,000 |
| Category: Local state or geometry Target region: rose | ||
|---|---|---|
| Steering instruction: make the rose bloom | ||
| Target query : rose in full bloom | ||
| Under-specified | Explicit-conflict | |
| Source prompt | a single rose in a garden with morning dew | a single closed rose bud in a garden with morning dew |
| Source state | Blooming state is unspecified. | A closed bud is explicitly specified. |
| Reference query | rose | rose bud |
| Validation | External Top-1 | External | ||||
|---|---|---|---|---|---|---|
| Image domain | Top-1 | Top-5 | Released | Other | Matched | Top-5 |
| SANA-Sprint | 60.2 | 83.3 | 9.9 | 51.1 | 77.5 | 95.1 |
| Nitro-1-PixArt | 51.7 | 74.8 | 9.5 | 52.6 | 66.7 | 88.4 |
| SANA-Sprint | Nitro-1-PixArt | |||||
|---|---|---|---|---|---|---|
| Style | Val. | Ext. | Base. | Val. | Ext. | Base. |
| Blossom Season | 87.5 | 91.7 | 90.0 | 85.0 | 87.5 | 90.0 |
| Comic Etch | 85.0 | 100.0 | 90.0 | 92.5 | 100.0 | 100.0 |
| Mosaic | 90.0 | 91.7 | 92.5 | 82.5 | 100.0 | 100.0 |
| Neon Lines | 95.0 | 100.0 | 100.0 | 95.0 | 100.0 | 100.0 |
| Pencil Drawing | 90.0 | 100.0 | 100.0 | 97.5 | 100.0 | 100.0 |
| Component | Setting |
|---|---|
| SAE site | Layer-14 post-block residual, 18,432 features |
| SAE families | L1, TopK, JumpReLU |
| Input settings | None, mean centering, geometric-median centering, LayerNorm, RMS scaling |
| Target styles | 10 shared styles |
| Object contexts | Architectures, Birds, Cats, Dogs, Flowers, Horses, Human, Trees |
| Generation seeds | Five |
| SAE family | Input convention | UA | IRA | CRA | Control UA |
|---|---|---|---|---|---|
| SANA-Sprint | |||||
| L1 | None | 89.0 | 52.6 | 57.8 | 49.5 |
| L1 | Mean | 38.0 | 96.6 | 92.5 | 3.5 |
| L1 | Geo. median | 37.8 | 96.5 | 92.5 | 3.5 |
| L1 | LayerNorm | 70.0 | 71.1 | 80.7 | 26.7 |
| L1 | RMS | 29.0 | 85.2 | 93.0 | 15.7 |
| SANA-Sprint | Nitro-1-PixArt | |||||||
|---|---|---|---|---|---|---|---|---|
| Target style | UA | IRA | CRA | Control | UA | IRA | CRA | Control |
| Blossom Season | 70.0 | 96.9 | 92.5 | 10.0 | 10.0 | 99.7 | 97.5 | 10.0 |
| Comic Etch | 12.5 | 97.2 | 92.5 | 10.0 | 22.5 | 98.6 | 97.5 | 0.0 |
| Mosaic | 100.0 | 96.9 | 92.5 | 7.5 | 32.5 | 95.8 | 97.5 | 0.0 |
| Neon Lines | 2.5 | 95.0 | 92.5 | 0.0 | 90.0 | 96.1 | 97.5 | 0.0 |
| Pencil Drawing | 100.0 | 96.4 | 92.5 | 0.0 | 87.5 | 77.8 | 100.0 | 0.0 |
| Split | Generated | Nude | Non-nude | Uncertain | Used |
|---|---|---|---|---|---|
| Discovery (paired) | 80 | 35 | 42 | 3 | 77 |
| Test (paired) | 80 | 35 | 43 | 2 | 78 |
| Test (hard negatives) | 40 | 0 | 40 | 0 | 40 |
| All test | 120 | 35 | 83 | 2 | 118 |
| Total | 200 | 70 | 125 | 5 | 195 |
| Component | Setting |
|---|---|
| Generator | SANA-Sprint 0.6B |
| Generation | , one inference step, guidance 0 |
| SAE | Layer-14 mean-centered L1 |
| Dictionary size | 18,432 features |
| Spatial resolution | tokens |
| Reporting pool size |