The Fréchet distance is a de facto standard for evaluating generative models across domains, appearing as FID for images and FVD for videos. It summarizes the discrepancy between generated and reference distributions in a single scalar, with lower values typically interpreted as better generation quality. However, this scalar view can obscure what drives the comparison. For example, in COCO dataset, increasing the number of diffusion sampling steps improves ImageReward scores yet worsens (increases) FID. Motivated by this mismatch, we seek to make the Fréchet distance more interpretable by uncovering where the discrepancy lies. To this end, we introduce directional Fréchet distance, the expected squared projection of the optimal transport displacement onto a given direction. Across our image, video, and protein case studies, we find that a small number of interpretable directions account for much of the distance. We use these directions to explain the FID increase in terms of semantic concepts represented by CLIP embeddings, quantify FVD's bias toward per-frame appearance, and revisit the interpretation of Protein FID. We open-source our codebase at https://github.com/yhlee-add/directional-fd.
Figures & tables
Figure 1: Sign reversal on SD1.5/COCO under two standard knobs. As the classifier-free guidance scale (left) or the number of sampling steps (right) increases, per-sample quality rises while FID worsens. A lower FID therefore does not track higher per-sample quality.
share
dictionary
surviving concepts (weight)
cos
e1
24.0%
CLIP
composition simplicity ( 0.55 ), depth of field ( 0.12 ), furniture ( 0.08 )
Table 1: Sparse concept fit of the top eigenvectors of M over the two concept dictionaries, for SD1.5, 50 steps, and COCO. Share is FD(ei)/FD , the weights are the largest surviving Lasso coefficients, and cos is the reconstruction cosine.
Figure 2: Reference and generated features projected onto the top three eigenvectors of M , using SD1.5, 50 steps, and COCO. Each point pairs a caption’s reference and generated projection. The reference (red) and generated (blue) marginal distributions are drawn inside, and the dashed line is y=x . For e1 , the points are slightly above y=x , indicating that the generated set is shifted toward +e1 . For e2 , the points concentrate vertically, demonstrating that the generated set has less variance, and vice versa for e3 .
Figure 3: Left : directional FD of the top three fixed eigenvectors against total FID across the eight-step COCO sweep, in the frozen 50 -step basis. Each point is one step setting, the solid line is the least-squares fit, and its slope is the fraction of the FID change that eigenvector carries, and the dashed line is the slope=1 an eigenvector carrying the entire change would follow. The top eigenvector e1 carries most of it (slope 0.66 , R2=0.96 ), while e2 and e3 stay nearly flat (slopes 0.03 and 0.02 ). Right : cumulative slope s1:i , the summed slopes of the top i eigenvectors, which rises above the whole change (dashed) over the head and falls back to it over the tail (the eigenvectors before and after the curve’s peak).
flicker set, M(R,T)
generator, M(R,G)
k=1
k=10
k=1
k=10
embedding, family
appearance
flicker
appearance
flicker
appearance
flicker
appearance
flicker
I3D, motion blur
0.325
0.023
0.663
0.273
0.022
0.005
0.288
0.150
I3D, elastic
0.372
0.211
0.648
0.495
0.027
0.012
0.288
0.107
VideoMAE, motion blur
0.061
0.470
0.181
0.612
0.013
0.019
0.095
0.096
VideoMAE, elastic
0.163
0.389
0.282
0.553
0.012
0.009
0.148
0.084
Table 2: Share of directional FD on the appearance and flicker bands (the top eigenvectors of M(R,S) and M(S,T) ), at the top k=1 and k=10 eigenvectors. The flicker-set panel evaluates M(R,T) at severity 3 , with bands averaged over the other four severities. The generator panel evaluates ModelScope’s generated set, zero-shot on the UCF-101 class names, on bands averaged over all five severities. The two bands overlap, so the shares of a row can sum past one.
Table 3: Sparse concept fit of the top eigenvectors of M over the geometry concept dictionary, and how the two sets differ along each, for La-Proteina at 400 steps against the ESM3 reference, in the 32 -dimensional PCA space of Protein FID. Share is FD(ei)/FD , the weights are Lasso coefficients, cos is the reconstruction cosine, d is the Cohen’s d of the reference and generated projections onto ei , and var is their variance ratio. We show the top surviving concepts here. Refer to Appendix G for details.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
dictionary
surviving concepts (weight)
cos
e1
CLIP
composition simplicity ( +0.55 ), depth of field ( +0.12 ), furniture ( +0.08 ), electronics ( +0.05 ), daytime ( +0.04 ), portrait versus scene ( −0.03 )
Table 4: Every concept whose weight rounds to at least 0.01 in the sparse concept fits of the top eigenvectors of M , for SD1.5, 50 steps, and COCO. Each weight is positive when the concept, or class logit, grows toward +ei , the generated end. The last column is the reconstruction cosine.
Figure 4: Cumulative slope β1:i=∑j≤iβj of the top i fixed eigenvectors against i . Here βi is the slope of the directional FD FD(ei) against the step count, in each reference’s own fixed 50 -step basis. The curves of COCO, Flickr, ImageNet, and MJHQ rise over the head before falling over the tail, while the face curves descend throughout. Each endpoint is the total slope βFD , which is positive for COCO, Flickr, and ImageNet, negative for the faces, and near zero for MJHQ.
embedding, family
set
cos
union
flicker ∣ appearance
appearance ∣ flicker
I3D, motion blur
flicker set
0.74
0.746
0.083
0.473
I3D, elastic
flicker set
0.80
0.833
0.185
0.338
VideoMAE, motion blur
flicker set
0.61
0.786
0.605
0.174
VideoMAE, elastic
flicker set
0.69
0.741
0.459
0.188
I3D, motion blur
generated
0.74
0.384
0.096
0.234
I3D, elastic
generated
0.80
0.371
0.083
0.265
Appendix
Table 5: Overlap of the k=10 appearance and flicker bands. The flicker set is M(R,T) at severity 3 with bands computed from the other severities, and the generated set is M(R,G) with bands computed from all severities. Columns give the largest principal-angle cosine between the bands, the share of their union, and the share of the flicker band orthogonalized against the appearance band (flicker ∣ appearance) and the reverse.
severity
1
2
3
4
5
I3D ratio
1.41
1.29
1.18
1.11
1.07
VideoMAE ratio
3.56
3.63
3.23
2.87
2.77
I3D flicker share, k=1
0.053
0.028
0.023
0.028
0.031
VideoMAE flicker share, k=1
0.235
0.371
0.470
0.475
0.374
Appendix
Table 6: Scalar ratio and top flicker-eigenvector share across the five motion blur severities, with bands computed from the other severities.
Table 7: Full sparse concept fits of the top eigenvectors of M for La-Proteina at 400 steps, in the 32 -dimensional PCA space of Protein FID (Table 3 , penalty 10−3 ) and in the raw 1536 -dimensional ESM3 space (penalty 10−5 ). Every surviving concept is listed with its weight, and the other columns are as in Table 3 .
We show that Fréchet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation space. Our idea is simple: decouple the population size for FD estimation (e.g., 50k) from the batch size for gradient computation (e.g., 1024). We term this approach FD-loss. Optimizing FD-loss reveals several surprising findings. First, post-training a base generator with FD-loss in different representation spaces consistently improves visual quality. Under the Inception feature space, a one-step generator achieves0.72 FID on ImageNet 256x256. Second, the same FD-loss repurposes multi-step generators into strong one-step generators without teacher distillation, adversarial training or per-sample targets. Third, FID can misrank visual quality: modern representations can yield better samples despite worse Inception FID. This motivates FDrk, a multi-representation metric. We hope this work will encourage further exploration of distributional distances in diverse representation spaces as both training objectives and evaluation metrics for generative models.
Fréchet Inception Distance (FID) is widely used to evaluate image generators, yet lower FID does not always correspond to better sample quality. We show that this mismatch depends in part on the geometry of the reference dataset. In a controlled study across six datasets, distributional density and effective rank significantly explain how FID changes as sample quality improves. Concentrated datasets tend to yield more favorable FID trends, whereas more dispersed datasets can make FID worsen despite better samples. Attribution to precision and recall and ablations with alternative feature spaces and distances support the same conclusion. These results suggest that distributional metrics should be interpreted together with the geometry of the reference dataset for more reliable benchmarking.
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. We attribute this failure to the static pretrained feature spaces used by existing Fréchet losses. These feature spaces provide incomplete and fixed views of the differences between real and generated distributions. To address this limitation, we propose Adversarial Fréchet Distance (AdvFD), which complements the static representation targets in FD-Loss with a calibrated adversarially learned representation. AdvFD augments the original static Fréchet objective with a learnable representation that adversarially maximizes the Fréchet discrepancy between real and generated samples, while the generator minimizes the same discrepancy in the resulting adaptive feature space. To prevent the adversarial representation from trivially increasing the objective through feature amplification, we further introduce real-feature whitening, which normalizes its scale and covariance geometry and stabilizes the min--max optimization. Extensive experiments show that AdvFD consistently improves one-step generator post-training across both JiT and pMF backbones and across different model scales.