This paper aims to analyse the feature space of a vision-related Deep Neural Network (DNN) by proposing a decoder that can generate an image whose feature closely matches a user-specified feature. Supported by quantitative evidence of its high feature-matching accuracy, our decoder facilitates precise analysis of the DNN's feature space. Our decoder is implemented as a guided diffusion model that guides the image generation of a pre-trained diffusion model to minimise the Euclidean distance between the feature of a clean image estimated at each step and the user-specified feature. The key advantages of our decoder are its training-free applicability to analyse the feature spaces of different DNNs and its practical feasibility on a single COTS GPU. The experiments targeting CLIP's image encoder and ResNet-50 demonstrate the effectiveness of our decoder both as a feature-matching image generator and as a visual feature space analyser. The codes and data are available at https://github.com/ccilab-doshisha/FeatDec
Figures & tables
Figure 1 : An overview of our decoder that guides stable diffusion’s image generation to generate an image whose feature closely matches a user-specified feature fs . At step t , our decoder estimates a clean latent representation z^t,0 from zt , generates the corresponding image x^t from z^t,0 , extracts the feature f(x^t) of x^t , and computes the loss as the squared Euclidean distance between f(x^t) and fs .
Figure 2 : Results obtained by targeting CLIP’s image encoder with ResNet-50 backbone and defining fs as a feature, which is extracted from each of 100 MS-COCO images (a) and 100 ImageNet images (b).
Figure 3 : Examples of human evaluation scores. For each generated image, in addition to the squared Euclidean distance between its feature and fs of the corresponding actual image, the scores given by three human evaluators are shown in square brackets. (a) The generated image acceptably shows a bedroom as in the actual image. (b) The generated image misses a rocking chair that is one of the main objects in the actual image. (c) The anatomical topology of a wolf in the generated image is broken. (d) A grey fox cannot be recognised from the generated image.
Figure 4 : A comparison of images generated by our decoder and RCDM when ResNet-50 is the target feature extractor, and fs is defined as the feature extracted from each of the 100 MS-COCO and 100 ImageNet images used in Fig. 2 .
Figure 5 : Three of the five cases where an image generated by our decoder has a larger squared Euclidean distance to fs (extracted from an actual ImageNet image) than the nearest image retrieved from the non-sampled ImageNet dataset.
Figure 6 : Instance-level visualisation of the modality gap in CLIP’s feature space. Each row is based on a paired caption and image shown in the first and fourth columns, respectively. The second and third columns present images generated by defining fs as a caption feature extracted by CLIP’s text encoder and its rescaled version, respectively. The rightmost three columns display images generated by defining fs as 0.8× -, 1.0× - (i.e., original) and 1.2× -scaled versions of an actual image feature extracted by CLIP’s image encoder. The number under each generated image and the number in brackets indicate the squared Euclidean distance and normalised cosine similarity between its feature and fs used to generate it, respectively. When fs is an image feature, the cosine similarity between fs and the feature of every generated image is always larger than 0.99 , so such cosine similarities are omitted.
Fig. S1 : Distributions of values in ∇ztl(f(x^t),fs) at the first self-recurrence iteration (i.e., k=0 ) for each of the first seven reverse steps (i.e., t=999,⋯,993 ) when the target feature extractor f is defined as CLIP’s image encoder with ResNet-50 backbone (a), ResNet-50 (b) and ViT-H/14 (c). Here, fs is defined as the feature extracted from image 1 in Fig. S2 . In each visualised distribution, the dotted line indicates the mean and the solid lines represent ±∇thres used in gradient clipping (i.e., three times the standard deviation for (a) and (b), and twice the standard deviation for (c)).
Fig. S2 : 15 actual images used in the additional experiments.
Fig. S3 : Images generated by our decoder when using ViT-H/14 as the target feature extractor. For each index, three generated images are shown together with the squared Euclidean distances between their features and that of correspondingly-indexed actual image in Fig. S2 .
Fig. S4 : 10 images that are carefully chosen from the tusker category of ImageNet dataset [ 35 ] and validated by human judgment to be visually similar.
Fig. S5 : A comparison between finally-generated and best images when using CLIP’s image encoder with ResNet-50 backbone as the target feature extractor. For each index, our decoder is run three times to generate images whose features closely match the feature of the correspondingly-indexed image in Fig. S2 . For each run, the finally-generated image and the best image are shown together with the squared Euclidean distances between their features and that of the corresponding actual image.
Fig. S6 : (Continued from Fig. S5 ) A comparison between finally-generated and best images when using CLIP’s image encoder with ResNet-50 backbone as the target feature extractor.
Fig. S7 : Comparison of histograms of the squared Euclidean distances between the user-specified features derived from the 100 MS-COCO (or 100 ImageNet) images and the features of the images generated by our decoder (a) and RCDM (b), with ResNet-50 used as the target feature extractor.
(a) MS-COCO
Score
4
3
2
1
Mean ± std
Ours
8.6 %
24.0 %
48.3 %
19.0 %
2.223 ± 0.853
RCDM
0.3 %
16.0 %
53.0 %
30.6 %
1.860 ± 0.679
(b) ImageNet
Score
4
3
2
1
Mean ± std
Ours
15.6 %
24.3 %
44.6 %
15.3 %
2.403 ± 0.929
Table S1 : Comparison of human evaluation scores for images generated by specifying features extracted from the 100 MSCOCO images (a) and 100 ImageNet images (b) through ResNet-50.
Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet 256×256, PerF-L achieves FID of 1.91, approaching 1.86 of JiT-H with only half the parameters, while PerF-H further achieves FID of 1.63 and 1.76 on ImageNet 256×256 and 512×512, respectively.
Chong Wang, Zixuan Fu, Shiqi Huang +3
Nanyang Technological University · KTH Royal Institute of Technology · Hebei University of Technology
Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70× over Vanilla while maintaining competitive output quality.
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China
Tokenizers are a key component of state-of-the-art generative image models, extracting the most important features from the signal while reducing data dimension and redundancy. Most current tokenizers are based on KL-regularized variational autoencoders (KL-VAE), trained with reconstruction, perceptual and adversarial losses. Diffusion decoders have been proposed as a more principled alternative to model the distribution over images conditioned on the latent. However, matching the performance of KL-VAE still requires adversarial losses, as well as a higher decoding time due to iterative sampling. To address these limitations, we introduce a new pixel diffusion decoder architecture for improved scaling and training stability, benefiting from transformer components and GAN-free training. We use distillation to replicate the performance of the diffusion decoder in an efficient single-step decoder. This makes SSDD the first diffusion decoder optimized for single-step reconstruction trained without adversarial losses, reaching higher reconstruction quality and faster sampling than KL-VAE. In particular, SSDD improves reconstruction FID from 0.87 to 0.46 with 1.4× higher throughput and preserve generation quality of DiTs with 3.8× faster sampling. As such, SSDD can be used as a drop-in replacement for KL-VAE, and for building higher-quality and faster generative models.
Théophane Vallaeys, Jakob Verbeek, Matthieu Cord
Meta Fundamental AI Research · Sorbonne University