This paper aims to analyse the feature space of a vision-related Deep Neural Network (DNN) by proposing a decoder that can generate an image whose feature closely matches a user-specified feature. Supported by quantitative evidence of its high feature-matching accuracy, our decoder facilitates precise analysis of the DNN's feature space. Our decoder is implemented as a guided diffusion model that guides the image generation of a pre-trained diffusion model to minimise the Euclidean distance between the feature of a clean image estimated at each step and the user-specified feature. The key advantages of our decoder are its training-free applicability to analyse the feature spaces of different DNNs and its practical feasibility on a single COTS GPU. The experiments targeting CLIP's image encoder and ResNet-50 demonstrate the effectiveness of our decoder both as a feature-matching image generator and as a visual feature space analyser. The codes and data are available at https://github.com/ccilab-doshisha/FeatDec
Figures & tables
Figure 1 : An overview of our decoder that guides stable diffusion’s image generation to generate an image whose feature closely matches a user-specified feature fs . At step t , our decoder estimates a clean latent representation z^t,0 from zt , generates the corresponding image x^t from z^t,0 , extracts the feature f(x^t) of x^t , and computes the loss as the squared Euclidean distance between f(x^t) and fs .
Figure 2 : Results obtained by targeting CLIP’s image encoder with ResNet-50 backbone and defining fs as a feature, which is extracted from each of 100 MS-COCO images (a) and 100 ImageNet images (b).
Figure 3 : Examples of human evaluation scores. For each generated image, in addition to the squared Euclidean distance between its feature and fs of the corresponding actual image, the scores given by three human evaluators are shown in square brackets. (a) The generated image acceptably shows a bedroom as in the actual image. (b) The generated image misses a rocking chair that is one of the main objects in the actual image. (c) The anatomical topology of a wolf in the generated image is broken. (d) A grey fox cannot be recognised from the generated image.
Figure 4 : A comparison of images generated by our decoder and RCDM when ResNet-50 is the target feature extractor, and fs is defined as the feature extracted from each of the 100 MS-COCO and 100 ImageNet images used in Fig. 2 .
Figure 5 : Three of the five cases where an image generated by our decoder has a larger squared Euclidean distance to fs (extracted from an actual ImageNet image) than the nearest image retrieved from the non-sampled ImageNet dataset.
Figure 6 : Instance-level visualisation of the modality gap in CLIP’s feature space. Each row is based on a paired caption and image shown in the first and fourth columns, respectively. The second and third columns present images generated by defining fs as a caption feature extracted by CLIP’s text encoder and its rescaled version, respectively. The rightmost three columns display images generated by defining fs as 0.8× -, 1.0× - (i.e., original) and 1.2× -scaled versions of an actual image feature extracted by CLIP’s image encoder. The number under each generated image and the number in brackets indicate the squared Euclidean distance and normalised cosine similarity between its feature and fs used to generate it, respectively. When fs is an image feature, the cosine similarity between fs and the feature of every generated image is always larger than 0.99 , so such cosine similarities are omitted.
Fig. S1 : Distributions of values in ∇ztl(f(x^t),fs) at the first self-recurrence iteration (i.e., k=0 ) for each of the first seven reverse steps (i.e., t=999,⋯,993 ) when the target feature extractor f is defined as CLIP’s image encoder with ResNet-50 backbone (a), ResNet-50 (b) and ViT-H/14 (c). Here, fs is defined as the feature extracted from image 1 in Fig. S2 . In each visualised distribution, the dotted line indicates the mean and the solid lines represent ±∇thres used in gradient clipping (i.e., three times the standard deviation for (a) and (b), and twice the standard deviation for (c)).
Fig. S2 : 15 actual images used in the additional experiments.
Fig. S3 : Images generated by our decoder when using ViT-H/14 as the target feature extractor. For each index, three generated images are shown together with the squared Euclidean distances between their features and that of correspondingly-indexed actual image in Fig. S2 .
Fig. S4 : 10 images that are carefully chosen from the tusker category of ImageNet dataset [ 35 ] and validated by human judgment to be visually similar.
Fig. S5 : A comparison between finally-generated and best images when using CLIP’s image encoder with ResNet-50 backbone as the target feature extractor. For each index, our decoder is run three times to generate images whose features closely match the feature of the correspondingly-indexed image in Fig. S2 . For each run, the finally-generated image and the best image are shown together with the squared Euclidean distances between their features and that of the corresponding actual image.
Fig. S6 : (Continued from Fig. S5 ) A comparison between finally-generated and best images when using CLIP’s image encoder with ResNet-50 backbone as the target feature extractor.
Fig. S7 : Comparison of histograms of the squared Euclidean distances between the user-specified features derived from the 100 MS-COCO (or 100 ImageNet) images and the features of the images generated by our decoder (a) and RCDM (b), with ResNet-50 used as the target feature extractor.
(a) MS-COCO
Score
4
3
2
1
Mean ± std
Ours
8.6 %
24.0 %
48.3 %
19.0 %
2.223 ± 0.853
RCDM
0.3 %
16.0 %
53.0 %
30.6 %
1.860 ± 0.679
(b) ImageNet
Score
4
3
2
1
Mean ± std
Ours
15.6 %
24.3 %
44.6 %
15.3 %
2.403 ± 0.929
Table S1 : Comparison of human evaluation scores for images generated by specifying features extracted from the 100 MSCOCO images (a) and 100 ImageNet images (b) through ResNet-50.
Jul 30, 2026·Hanshuai Cui, Zhiqing Tang, Zhi Yao +3DrafterCorrection
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China