Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's performs as well as or slightly better than dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. First-order matching extends naturally to 1D token-sequence latents and across modalities, where distilling a text encoder into an image autoencoder still improves diffusability; a relational objective based only on each image's nearest neighbours improves it as well. Code and blog post are available at https://github.com/AdrienRR/structure-agnostic-distillation and https://kyutai.org/blog/2026-09-28-structure-agnostic-distillation/.
Figures & tables
Baselines
Structure-agnostic (ours)
Setting
Bottleneck
Teacher
No Distill.
VF ( Yao and Wang, 2025 )
Pool-Align
CKA
Soft-KL
A: shared grid
2D grid
DINOv2 ( Oquab et al., 2024 )
9.52
6.04
5.77
7.06
6.51
B: no grid
1D seq
DINOv2 ( Oquab et al., 2024 )
26.49
n/a
15.62
18.13
17.87
C: cross-modal
2D grid
BGE-large ( Xiao et al., 2024 )
9.52
n/a
6.97
10.41
8.51
Table 1: Generation quality (gFID, ↓ ), without classifier-free guidance (no per-arm tuning); for the 2D grid the ranking is preserved under VA-VAE’s guidance recipe (Appendix A ). Autoencoders are trained for 50 epochs on 256×256 ImageNet and the diffusion priors for 64 epochs (Appendix A ). Rows are settings A–C (bottleneck structure × teacher model). The structure-agnostic methods ( Pool-Align , CKA , Soft-KL ) apply in every setting; the position-wise baseline VF needs a shared latent–teacher grid, so it is inapplicable ( n/a ) in the 1D (B) and cross-modal (C) settings. Within a row only the distillation method changes, with architecture and budget fixed, so differences reflect the objective, not compute; absolute gFID differs across bottleneck structures, so we compare only within a setting. Best per row in bold .
Figure 2: Distillation speeds up diffusion training. Per-checkpoint gFID ( 50 k samples, without classifier-free guidance) against DiT training steps for the 2D-grid (left) and 1D-sequence (right) settings. Every distilled latent converges to a lower final gFID than No-Distillation and reaches low gFID in fewer steps; the horizontal marker reports how many fewer steps Pool-Align needs to reach No-Distillation ’s final gFID.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Autoencoder (tokenizer), all arms
Optimizer
Adam, β=(0.5,0.9) , weight decay 0
Learning rate
10−4 , constant (no warmup or schedule)
Global batch size
256 ( 8 per GPU ×32 GPUs)
Epochs
50
Precision
fp32
Gradient clipping
global norm 1.0
Appendix
Table 2: Full training and evaluation configuration. The autoencoder trains first, then the diffusion prior on its frozen latents. Data is ImageNet-1k at 256×256 , normalized to [−1,1] , with random-crop augmentation (no horizontal flip). The weight target r is the gradient-norm-balancing factor of Appendix A.2 .
Figure 3: Position-wise alignment carries a spatial gradient; pooled latents track content. Top-3 PCA (as RGB) of one image’s representation for the DINOv2 teacher, the undistilled latent, and each distilled latent. The undistilled latent is nearly featureless; the position-wise VA-VAE ( VF ) latent carries a smooth horizontal spatial gradient, whereas the pooled latents show no such gradient and instead pick out the foreground object from the background, tracking content rather than residual information from positional embeddings in the teacher’s representations.
Figure 4: Samples from Pool-Align on the 2D grid (DINOv2 teacher, setting A): autoencoder trained for 50 epochs, diffusion prior for 64 epochs ( 80 k steps); sampled in 250 Euler steps with classifier-free guidance (recipe in Appendix A ), which corresponds to gFID 1.99 .
Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.
Jean-Marie Lemercier, Tomas Geffner, Karsten Kreis +3
We demonstrate that in knowledge distillation for diffusion models, the teacher network's highly complex denoising process - stemming from its substantially larger capacity - poses a significant challenge for the student model to faithfully mimic. To address this problem, we propose a coarse-to-fine distillation framework with LInear FiTtingbased distillation (LIFT) and Piecewise Local Adaptive Coefficient Estimation (PLACE). First, LIFT decomposes the objective into a "coarse" alignment and a "fine" refinement. The student is then trained on coarse alignment before proceeding to hard refinement. Second, PLACE extends LIFT to address spatially non-uniform errors by partitioning outputs into error-based groups, providing locally adaptive guidance. Our experiments show that LIFT and PLACE is effective across diffusion spaces (image/latent), backbones (U-Net/DiT), tasks (unconditional/conditional), datasets, and even extends to flow-based models such as MMDiT (SD3). Furthermore, under extreme compression with a 1.3M-parameter student (only 1.6% of the teacher), conventional KD fails to provide sufficient guidance for stable training, with FID scores often degrading to 50-200+, but our method remains stably convergent and achieves an FID of 15.73.
Hyunsoo Han, Sangyeop Yeo, Jaejun Yoo
Ulsan National Institute of Science and Technology (UNIST)
Discrete diffusion models excel at visual synthesis but rely on slow, iterative decoding. Existing single-step distillation methods attempt to bypass this bottleneck, either by training auxiliary score networks that effectively double compute, or by introducing specialized parameterizations and multi-stage pipelines that fragment optimization. In this paper, we introduce Fixed-Point Distillation (FPD), an end-to-end framework that constructs local correction targets by partially corrupting the student's one-step draft and refining it with a single teacher step. To compute the training objective in a semantically meaningful space, we lift discrete tokens into continuous features and apply a multi-bandwidth drift loss that iteratively accumulates these corrections. To backpropagate through the discrete bottleneck, we employ a straight-through estimator that feeds exact hard-sampled tokens to the teacher and decoder during the forward pass, ensuring that training and inference operate on the same codebook manifold, while routing continuous gradients back to the student logits. This fully differentiable pathway additionally accommodates an optional unconditional adversarial objective to enhance perceptual realism. Evaluations on both class- and text-conditional generation validate the effectiveness of our framework. FPD achieves competitive visual fidelity and structural alignment within a single inference step, narrowing the gap to multi-step teachers while outperforming existing discrete distillation baselines.