The unsupervised discovery of features that are both semantically meaningful and stable across runs remains a central challenge in representation learning. We introduce entropy-ordered flows (EOFlows), a normalizing flow (NF) framework that augments standard maximum likelihood training with an orthogonality regularizer on the decoder Jacobian. The regularizer is rooted in Independent Mechanism Analysis and encourages geometric disentanglement, and a stochastic estimator makes it tractable at image scale (CelebA at D=2352 and 12288). Learned features form near-orthogonal curvilinear coordinates and can be ordered by their explained (manifold) entropy after training, analogous to the ranking by explained variance in PCA, which turns EOFlows into a non-linear generalization of PCA. EOFlows identify an order of magnitude more stable features than existing methods, and these features emerge in distinguishable categories (global, local, and generic) and support tentative semantic interpretations. The local features have strikingly sparse support in pixel space, although our method never enforces this. Retaining only the most important, i.e. highest entropy, features turns the bijective flow into an autoencoder with adjustable bottleneck, rivaling the rate-distortion performance of dedicated autoencoders.
Figures & tables
Figure 1: Archetypes of the features learned by EOFlows on CelebA, sorted by decreasing entropy. Global features involve the entire image, e.g. head pose. Local features affect only a small number of pixels, e.g. eye gaze. Generic features resemble windowed Fourier transforms when a subregion (e.g. the background) lacks semantic structure. Finally, features with low entropy contain only noise.
Figure 2: Induced coordinates. A standard normal latent (left) is pushed forward by linear (center) and non-linear (right) decoders, which induce affine and curvilinear coordinates in the data space. Both decoders of each pair fit the same distribution, but only the ones on the right, PCA and EOFlows, have orthogonal coordinate lines U1(z2) and U2(z1) that align with the data geometry.
Figure 3: Denoising of held-out images by one Tweedie step and through the top- C bottleneck of an EOFlow ( λ=1 ).
Figure 4: Entropy spectra, compression and denoising rate-distortion for representative models. (a) Manifold entropy spectra Hi define the ordering of latent dimensions. EOFlows and NDFlows concentrate their entropy in a few hundred dimensions; EOFlows flatten towards the noise entropy Hσϵ , separating high-entropy core from low-entropy detail dims. The plain NF spreads it over all dims, and FactorVAE (best VAE) uses at most 20 active latents. (b) Denoising error of held-out images reconstructed through the top- C entropy-ordered bottleneck; the plain NF is omitted, as its error is far above the plotted range. Additionally to models in (a): denoising autoencoders trained separately for each C (App. F.6 ), and Tweedie denoising of the flows without a bottleneck (dotted).
Figure 5: Stable features across seeds. (a, b) Sketch: coordinate lines Ui of a reference model (gray) against one that reproduces them (a) and one that does not (b), quantified by the alignment of their tangent vectors Ji . (c) Left: matched Ui stability (Sec. 4 ) of every coordinate, sorted in decreasing order (median over seed pairs). EOFlows keep gaining stable coordinates up to 500 epochs, whereas NDFlows and FactorVAE, trained just as long, do not. The plain NF has no stable coordinate, and noisy PCA (dashed) has hundreds of them, by construction. Budgets and seeds in Table F.1 . Right: archetypes of 7 representative EOFlow features in three random seeds (all six: App. G.7 ).
Figure 6: Equivariance of EOFlows.
Figure 7: Sweeping the orthogonality penalty λ from NF to PCA. Each NF/EOFlow trained with 3 – 4 seeds for 300 epochs, 6 noisy PCA fits. (a) Negative log-likelihood on training and held-out data; the gap between them is overfitting. (b) Hadamard gap LH/D . (c) Coordinates reproduced per seed pair (matched Ui stability ≥0.5 ). (d) Sparsity via mean and minimum ENZ. (e) Non-linearity of coordinates (one seed, App. F.7.4 ). (f) gFID of prior samples after Tweedie denoising.
CelebA (fully connected)
dSprites (conv.)
Shapes3D (conv.)
model
stability ↑
orthog- onality ↓
sparsity min/mean ↓
TAD ↑
MIG ↑
DCI ↑
MIG ↑
DCI ↑
EOFlow λ=1
121 ±5
0.041 ±.001
0.03 / 0.20
0.72 ±.01
0.14 ±.10
0.07 ±.04
0.15 ±.03
0.52 ±.03
EOFlow λ=0.1
25.0 ±3.6
0.17 ±<.01
0.09/0.32
1.00 ±.01
0.26 ±.01
0.19 ±<.01
0.17 ±.07
0.41 ±.09
NDFlow ν=1
8.5 ±1.8
0.29
0.30/0.40
0.24 ±.01
–
–
–
–
NDFlow ν=0.1
1.7 ±1.2
0.46 ±<.01
0.32/0.41
0.58 ±.07
–
–
–
–
NF
0
1.6 ±<.1
0.37/0.41
0.00 ±.00
–
–
–
–
Table 1: Reproducible, sparse and disentangled. Mean ±sd over seeds. CelebA at 28×28 : 300 epochs (AnnealedVAE: 100 ), fully-connected architectures; stability and sparsity as in Sec. 4 ; TAD = attribute disentanglement score ( Yeats et al., 2022 ) (App. G.6 ). Orthogonality = Hadamard gap of the core coordinates (App. F.7.1 ). dSprites and Shapes3D: 64×64 , 100 epochs, convolutional architectures (App. G.3 ). PCA fits orthogonal and reproducible components by construction (in parentheses). On the disentanglement benchmarks we compare EOFlows with the VAEs only.
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1: Variance of the stochastic Hadamard estimate along the λ sweep (last 50 epochs of training; lines are means over 3 – 6 seeds per λ , shaded bands ± one standard deviation, mostly narrower than the line). (a) Per-step variance of the estimate of LH/D in the logged training loss (blue) and as predicted by the variance formula at the same weights, from full Jacobians at 64 noise-inflated training images (dashed). (b) The same variance relative to the squared mean, i.e. to the penalty itself. (c) Per-step variance of the two terms of the training loss.
model
epochs
seeds
notes
EOFlow λ=1
500
103, 105, 106, 108–110
EOFlow λ=1
300
103, 105, 106
EOFlow λ=0.1
500
103, 105, 106
EOFlow λ=0.1
300
103, 105, 106
snapshots after 301 – 302 epochs
EOFlow λ=0.01
300
103, 105, 106
EOFlow λ=3
300
103, 106, 108, 109
Appendix
Table F.1: Trained models behind the reported numbers (CelebA 28×28 , σϵ=0.05 ; flows at batch size 512 and learning rate 6⋅10−4 unless noted). Cells at 300 and 500 epochs of the same λ or ν are the same trainings, evaluated at both lengths.
Figure F.1: Twelve random CelebA training images, clean (top) and inflated with Gaussian noise of σϵ=0.05 as seen by the flows (bottom; one noise draw, redrawn at every training step).
C=10
20
50
100
200
500
small AE ( 9 M), denoising (Fig. 4 b)
11.95
7.98
4.76
3.17
2.02
1.65
small AE ( 9 M), pure reconstruction
11.96
7.98
4.76
3.17
2.03
1.61
large AE ( 122 M), denoising
18.58
13.41
8.68
6.12
4.21
3.14
large AE ( 122 M), pure reconstruction
18.48
13.35
8.59
6.00
4.07
3.18
EOFlow λ=1
14.66
9.40
4.96
3.00
1.77
0.99
Appendix
Table F.2: Held-out denoising error (MSE, ×10−3 ; noisy input, clean target) of the autoencoders, one model per C and seed 101 , and of EOFlows ( λ=1 , mean over three seeds) through their top- C bottleneck.
model
C=32
C=300
C=D
EOFlow λ=1
0.035 ±<.001
0.041 ±.001
0.014 ±<.001
EOFlow λ=0.1
0.13 ±<.01
0.17 ±<.01
0.058 ±.002
NDFlow ν=1
0.20
0.29
0.20 ±<.01
NDFlow ν=0.1
0.38 ±.01
0.46 ±<.01
0.44 ±<.01
NF
1.2 ±<.1
1.6 ±<.1
1.6 ±<.1
FactorVAE
0.38 ±.07
–
–
Appendix
Table F.3: Orthogonality of all models: Hadamard gap per coordinate (Eq. F.2 ) over the C core coordinates, mean ±sd over seeds (NDFlow ν=1 at C≤300 : one seed). Table 1 combines C=300 for the flows and PCA with C=32 , i.e. all latents, for the VAEs; C=D for EOFlows and the NF is Fig. 7 b.
Figure F.2: Stability of coordinate grids (a, the protocol of Fig. 5 c) and correlation of latent codes (b, squared Pearson correlation on 10,000 held-out images, Hungarian-matched), sorted per seed pair (median, min–max band); same models and seed pairs as Fig. 5 .
Figure F.3: Geometry of the coordinates against the entropy rank i for all 2352 coordinates of the EOFlow ( λ=1 ) and of the plain NF (same seed, 300 epochs each) at 512 training images: (a) non-linearity across images (zero iff the lines Ui through all images are parallel), (b) mean norm of the Jacobian column Ji . Dots are single coordinates, lines running medians over rank bins.
arm
NLL
LH/D
FID
stable features per seed pair
ref (12 learnable mixings)
−1.291
0.012
29
54 / 46 / 40
last1 (final mixing learnable)
−1.285
0.017
28
9 / 11 / 7
frozen (no learnable mixing)
−1.262
0.030
42
2 / 3 / 5
nomix (permutations only)
−1.251
0.027
37
1827 / 1828 / 1824
noperm (couplings only)
−1.247
0.025
38
1834 / 1838 / 1843
Appendix
Table G.1: Architecture ablation on CelebA, 100 epochs, three seeds per arm (means; two ref seeds reach NLL −1.294 and FID 27 , the third −1.284 and 33 ; all other arms agree across seeds to the last digit). NLL: held-out negative log-likelihood per dimension of noise-inflated images; LH/D : Hadamard gap; FID at C=100 for prior samples; stable: latent dimensions with matched squared cosine similarity ≥0.5 , given for the three seed pairs. The stable counts of the two mixing-free arms are set in italics: they count the pixel basis, not learned features (see text).
Figure G.1: Stability of the latent dimensions across seeds for the five arms of the architecture ablation, three seed pairs per arm. (a) Matched squared cosine similarity of every latent dimension, sorted; the crossing with the dotted line at 0.5 gives the number of stable features. (b) The same quantity along the entropy order of the first seed (mean over ten neighboring dimensions); the two mixing-free arms (dashed) are stable only beyond rank ≈300 , in the pixel-basis tail.
Figure G.2: Archetypes of the five arms of the architecture ablation: the 100 highest-entropy archetypes of one seed of each arm (difference images xi+4−xi−4 as in Appendix G.10 , each normalized to [0,1] , entropy order row-major).
Figure G.3: Entropy spectra of the five arms of the architecture ablation, mean ± standard deviation over three seeds; the two mixing-free arms are dashed, and the dotted line is the noise entropy Hσϵ . The spectra coincide over the leading ranks, so the entropy ordering is comparable across the arms.
Figure G.4: Convolutional EOFlow on CelebA 64×64 ( D=12288 , σϵ=0.1 , 100 epochs). (a) Sorted manifold entropy spectra of the EOFlow ( λ=1 ), the same flow trained by maximum likelihood alone (NF, λ=0 ) and PCA of the noisy data (dashed), with the noise entropy Hσϵ (dotted). (b) Archetypes of the 100 highest-entropy EOFlow coordinates in entropy order (difference images g(+4ei)−g(−4ei) from the latent origin, each normalized). (c) EOFlow prior samples decoded through the C=100 leading coordinates and (d) the same latent draws decoded at full D after one Tweedie step.
λ=1
λ=0
PCA
NLL (nats/dim)
−0.785
−0.797
−0.767
entropy range, all latents (nats)
4.81
3.05
4.79
entropy range, 768 coarsest (nats)
4.70
1.18
–
LH/D
0.0079
0.123
0
MSE, C=100
0.0054
0.0730
–
MSE, C=768
0.0019
0.0021
–
Appendix
Table G.2: Convolutional EOFlow on CelebA 64×64 ( D=12288 , σϵ=0.1 , 100 epochs, one seed each) and PCA of the noisy data. NLL: held-out negative log-likelihood per dimension; entropy range: spread of the manifold entropies Hi over all latents and over the 768 coarsest ones; MSE: denoising error of 2000 held-out images through the C leading coordinates (noisy input: 0.0100 ).
Figure G.5: Entropy spectra on ground-truth benchmarks (four seeds each, σϵ=0.2 , λ=0.1 , 100 epochs). Sorted manifold entropy spectra Hi of EOFlows on (a) dSprites and (b) Shapes3D, the spectrum of PCA of the noisy data (dashed) and the noise entropy Hσϵ (dotted). The vertical line marks the number of ground-truth factors. EOFlows concentrate the entropy in a few coordinates and drop to the noise level at a sharp knee, which PCA lacks.
dSprites (conv.)
Shapes3D (conv.)
model
MIG ↑
DCI ↑
FactorVAE ↑
MIG ↑
DCI ↑
FactorVAE ↑
EOFlow λ=1
0.14 ±.10
0.07 ±.04
0.65 ±.13
0.15 ±.03
0.52 ±.03
0.80 ±.01
EOFlow λ=0.1
0.26 ±.01
0.19 ±<.01
0.81 ±.02
0.17 ±.07
0.41 ±.09
0.70 ±.02
FactorVAE
0.15 ±.07
0.14 ±.03
0.76 ±.09
0.13 ±.03
0.14 ±.05
0.73 ±.04
β -TCVAE
0.26 ±.09
0.25 ±.03
0.78 ±.07
0.31 ±.12
0.42 ±.06
0.71 ±.09
β -VAE
0.06 ±.02
0.06 ±.01
0.64 ±.06
0.22 ±.07
0.25 ±.08
0.83 ±.15
Appendix
Table G.3: Disentanglement on dSprites and Shapes3D ( 64×64 , 100 epochs). Mean ±sd over seeds. MIG ( Chen et al., 2018 ) , DCI disentanglement ( Eastwood & Williams, 2018 ) and the FactorVAE score ( Kim & Mnih, 2018 ) on 10 latents per model (VAEs: all 10 ; EOFlows and noisy PCA: the 10 highest-entropy coordinates), from clean held-out images. VAEs at literature settings ( β -VAE β=4 , AnnealedVAE C=25 and γ=1000 , FactorVAE γ=10 , β -TCVAE β=6 ); EOFlows with σϵ=0.2 . Seeds (dSprites / Shapes3D): EOFlow λ=1 3 / 3, EOFlow λ=0.1 4 / 4, FactorVAE 3 / 3, β -TCVAE 6 / 4, β -VAE 4 / 3, AnnealedVAE 3 / 3.
Figure G.6: Archetypes: EOFlow trained on shuffled pixels ( λ=1 , 100 ep, seed 101 ). Top: the 100 highest-entropy latents after un-shuffling with PT , in the format of App. G.10 (number above each triplet is the entropy rank). Bottom: the five leading latents in the basis the model was trained in, each image normalized to [0,1] .
Figure G.7: FID of prior samples through the top- C bottleneck (no Tweedie step; mean ± sd over seeds) for EOFlows at 300 epochs and PCA, and of full-prior samples after one Tweedie step (dotted lines, including the plain NF).
model
epochs
TAD
plain NF ( λ=0 )
300
0.00 ± 0.00 (3)
EOFlow λ=0.01
300
0.48 ± 0.21 (3)
EOFlow λ=0.1
300
1.00 ± 0.01 (3)
EOFlow λ=1
300
0.72 ± 0.01 (3)
EOFlow λ=3
300
0.60 ± 0.05 (4)
EOFlow λ=10
300
0.45 ± 0.03 (4)
Appendix
Table G.4: Attribute-based disentanglement (TAD) ( Yeats et al., 2022 ) on the CelebA validation split (higher is better; mean ± sd over seeds). Flows: same architecture and recipe; VAEs: best tuned cell of each family (App. F.4 ); number of seeds in parentheses. The weakly regularized λ=0.1 scores highest; λ=1 ties the best VAE with far less seed scatter.
Figure G.8: The 100 most stable features of EOFlow λ=1 after 500 epochs in all six seeds. The latents of the reference seed 105 are ranked by their mean matched stability over the five other seeds, column by column; each row shows the difference images xi+4−xi−4 of the matched latents in all six seeds (reference first, sign flips resolved). Left of each row: the entropy rank of the feature in the reference seed (#) and the mean ± sd of its matched cos2 over the five seed pairs.
Figure G.9: Prior samples of every model from the same 20 prior codes: decoded from the full prior, after one Tweedie step on these samples, and through bottlenecks that keep the C=100 or C=300 highest-entropy latents ( 300 epochs, seed 103 ). The plain NF has no meaningful latent order, so its bottlenecks keep arbitrary latents. Continued on the next page.
Figure G.10: Prior samples (continued): PCA (Gaussian samples with the covariance of the noisy data, their posterior mean, and samples of the top C components), NDFlows ( 300 epochs, seed 103 unless noted) and the VAEs of App. F.4 ( 300 epochs, seed 101 unless noted), which decode the first 32 entries of each code and give full-prior samples only.
Figure G.11: Reconstructions through the bottleneck, plain NF ( λ=0 ).
Figure G.12: Reconstructions through the bottleneck, EOFlow λ=0.01 .
Figure G.13: Reconstructions through the bottleneck, EOFlow λ=0.1 .
Figure G.14: Reconstructions through the bottleneck, EOFlow λ=1 .
Figure G.15: Reconstructions through the bottleneck, EOFlow λ=3 .
Figure G.16: Reconstructions through the bottleneck, EOFlow λ=10 .
Figure G.17: Reconstructions through the bottleneck, PCA (projection onto the leading C components; Tweedie = posterior mean of the Gaussian model).
Figure G.18: Archetypes: plain NF, λ=0 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.19: Archetypes: EOFlow λ=0.01 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.20: Archetypes: EOFlow λ=0.1 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.21: Archetypes: EOFlow λ=1 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.22: Archetypes: EOFlow λ=3 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.23: Archetypes: EOFlow λ=10 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.24: Archetypes: PCA. For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.25: Archetypes: NDFlow ν=0.1 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.26: Archetypes: NDFlow ν=1 (300 ep, seed 103). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.27: Archetypes: NDFlow ν=10 (300 ep, seed 106). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.28: Archetypes: FactorVAE γ=10 (300 ep, seed 101). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.29: Archetypes: β -TCVAE β=4 (300 ep, seed 101). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.30: Archetypes: β -VAE β=4 (300 ep, seed 101). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).
Figure G.31: Archetypes: AnnealedVAE C=25 (100 ep, seed 101). For every latent i (number above each triplet, ordered by decreasing manifold entropy): decoded image at zi=+4 (top), at zi=−4 (middle) and their difference image, normalized per tile (bottom).