Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.com/kaiyuyue/sphere2}{github.com/kaiyuyue/sphere2}.
Figures & tables
Figure 1: Curated 1-step generation without CFG on ImageNet 512×512 by Sphere Encoder 2.
Figure 2: Sphere Encoder vs. Sphere Encoder 2. Uncurated generation on ImageNet 256×256 from Sphere Encoder ( Yue et al., 2026 ) (top) and Sphere Encoder 2 (bottom) with same classes.
Figure 3: Sphere Encoder 2. The encoder E maps an image to a clean latent z^(0∘) on the sphere S , which is then rotated toward a random direction by an angle α∈[0,90∘] in Eq. 4 . The cutoff angle αcutoff splits the arc into a reconstruction regime ( α⩽αcutoff ) and a generation regime ( α⩾αcutoff ). The decoder D serves both: near the pole it reconstructs the input, and at the equator z^(90∘)=F(e) it generates a new image from a random latent in Eq. 2 .
Figure 5: Cutoff angle αcutoff , which splits the rotation arc into the reconstruction and generation regimes. Generation improves as the cutoff moves toward the equator until 70∘ .
Figure 6: Reconstruction and generation regimes at different angles from 0∘ to 90∘ on Oxford Flowers 256×256 . On the same rotation arc with αcutoff=70∘ , the decoder reconstructs the input below the cutoff and generates new images above it.
Figure 7: Score matching loss. Fake features from the frozen extractor ϕ are denoised by the real and fake score models (each a two-layer transformer EDM denoiser ( Karras et al., 2022 ) ), and their difference gives the direction g that pulls fake features toward real ones. During transport, both score models are frozen and run once each; only the decoder D is updated by g through Lscore .
Model
Params
Guid.
NFE
FDr 6 ↓ @ K steps
gFID ↓ @ K steps
1
2
4
6
1
2
4
6
Image Size 256×256
Sphere ( 2026 )
948M
CFG
K×2
10.59
5.44
3.96
3.92
25.12
14.08
11.25
10.63
Sphere Latent ( 2026 )
130M
CFG
K×2
–
–
–
–
–
12.22
8.61
7.85
Sphere2-B
159M
–
K
2.67
2.23
2.17
2.19
7.62
7.45
7.50
7.55
Sphere2-L
487M
–
K
2.19
1.85
1.82
1.83
9.36
8.33
7.98
7.93
Table 1: Main comparison on Oxford Flowers . Compared with Sphere Encoder ( Yue et al., 2026 ) and Sphere Latent Encoder ( Do et al., 2026 ) . Params for Sphere2 cover both encoder and decoder; for latent-space models they cover the generator only. Sphere2 has no CFG in sampling.
Figure 8: Uncurated 4-step generation without CFG on Oxford Flowers 512×512 by Sphere2-L.
Model
Params
Guid.
Total GFLOPs
256×256
512×512
GFLOPs ×
NFE
FDr 6 ↓
gFID ↓
FDr 6 ↓
gFID ↓
Latent space
SiT-XL/2 ( 2024 )
675M
CFG
119 ×
250×2
8.44
2.06
–
2.62
SiT-XL/2 + REPA ( 2025 )
675M
CFG
119 ×
250×2
5.45
1.42
–
2.08
LightningDiT-XL/2 ( 2025 )
675M
CFG
119 ×
250×2
4.57
1.35
–
–
DDT-XL/2 ( 2026b )
675M
CFG
119 ×
250×2
5.70
1.26
–
1.28
Table 2: Main comparison on ImageNet 256×256 and 512×512 . Params and GFLOPs for Sphere2 cover both encoder and decoder; for latent-space models they cover the generator only and exclude the VAE decoder. GFLOPs are measured for a single forward, so Total GFLOPs = GFLOPs × NFE. Guidance: CFG is classifier-free guidance applied at sampling, so NFE is doubled; AG is AutoGuidance ( Karras et al., 2024 ) , which doubles NFE with a smaller guiding model; CFG ∗ denotes guidance folded into the training target for single-forward sampling. Among pixel-space models, taking each model at its best setting, the best, second and third FDr 6 per resolution are shaded 00 , 00 and 00 . † marks Sphere2 models trained with the LFD-lite loss in Sec. B.1 .
Figure 9: Generation quality vs. sampling compute on ImageNet 256×256 among pixel-space models. Circles are pixel-space baselines and stars are Sphere2 at 1, 4 (without CFG) and 4×2 NFE (with CFG), joined per model. The halo area around each marker is proportional to parameters.
Model
Sampler
Guid.
bsz
img/s ↑
NFE
ms/NFE
50K imgs (min) ↓
Sphere2-L
loop × 1
–
1
53.4
1
18.7
16
loop × 2
CFG
1
14.5
4 + 2 enc.
17.3
58
loop × 4
CFG
1
6.1
8 + 6 enc.
20.7
137
loop × 1
–
16
424.8
1
2.4
2
loop × 2
CFG
16
145.4
4 + 2 enc.
1.7
6
loop × 4
CFG
16
62.8
8 + 6 enc.
2.0
14
Table 3: Sampling speed at 512×512 on a single NVIDIA H100. Timings are end-to-end, averaged over runs after warmup (std below 1% ), at batch size 1 and 16. For Sphere2, “loop ×K ” denotes K sampling steps of Alg. 1 . NFE counts decoder forwards; “ +m enc.” adds the encoder forwards used by spherical CFG, and “ +49 guide” the guiding-model forwards of AutoGuidance. Sphere2 and JiT run in bf16, pMF and RAE-DiT in fp32, following their official implementations.
Figure 10: Curated 2-step generation without CFG on ImageNet 512×512 by Sphere Encoder 2.
Figure 11: Uncurated samples on ImageNet 512×512 . Left: RAE-DiT DH -XL/2 with NFE=50×2 and Guidance=1.5 . Right: Sphere2-L with NFE=4 and no CFG. Best viewed zoomed in.
Figure 12: Uncurated samples on ImageNet 512×512 . Left: JiT-L/32 with NFE=50×4 and CFG=2.5 . Right: Sphere2-L with NFE=2 and no CFG. Best viewed zoomed in.
Figure 13: Uncurated samples on ImageNet 512×512 . Left: pMF-L/32 with NFE=1 and CFG=7.5 . Right: Sphere2-L with NFE=1 and no CFG. Best viewed zoomed in.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A14: Class interpolation on ImageNet 512×512 by Sphere2-L with 4-step sampling and no CFG. Each row fixes the random point and slerps between two class embeddings, from the left class to the right one. Rows are ordered from similars pairs at the top to dramatic pairs at the bottom.
Figure A15: Interpolation from non-ImageNet images by Sphere2-L trained on ImageNet 512×512 . The first column shows input images outside the ImageNet classes: a raccoon, a matcha cake, and grapes. The second and third columns decode the clean latent of each input conditioned on α=0∘ and α=85∘ , respectively. The remaining columns unroll the refinement loop of Alg. 1 , which carries the input toward a target class: raccoon to red panda (class 387), matcha cake to espresso (class 967), and grapes to indigo bunting (class 14).
Steps K
LFD-lite
gFID ↓
FDr 6 ↓
FDr 6 ↓ per feature space
Incep.
ConvNeXt
DINOv2
MAE
SigLIP
CLIP
Sphere2-B
-
5.20
14.53
3.49
2.99
13.18
15.08
28.01
24.42
1
✓
3.43
12.76
2.30
2.59
11.55
13.64
24.41
22.03
-
4.72
9.48
3.17
2.67
8.83
10.80
16.09
15.35
2
✓
2.69
8.43
1.81
2.36
7.43
9.62
14.83
14.51
Appendix
Table A4: Effect of LFD-lite per feature space on ImageNet 256×256 . FDr 6 at K sampling steps, trained with and without LFD-lite , broken down into its six feature spaces.
Figure A16: Effect of LFD-lite for Sphere2-L on ImageNet 256×256 . Left and middle: FID and FDr 6 with and without the loss at different sampling steps. Right: relative FDr reduction in each of the six feature spaces.
Figure A17: Distribution shift is invisible but measurable. 4-step samples on ImageNet 256×256 from Sphere2-B trained with and without LFD-lite are visually indistinguishable, yet the loss aligns their feature statistics with those of real data in the Inception space where gFID is measured. More uncurated samples for side by side comparison are in Fig. A18 for 1 step and Fig. A19 for 4 steps.
Figure A18: Uncurated samples on ImageNet 256×256 . Left: Sphere2-B without LFD-lite and NFE=1 . Right: Sphere2-B with LFD-lite and NFE=1 .
Figure A19: Uncurated samples on ImageNet 256×256 . Left: Sphere2-B without LFD-lite and NFE=4 . Right: Sphere2-B with LFD-lite and NFE=4 .
Figure A20: Angle distribution in training. Left: sampling curves of α for different k . Middle and right: generation quality of Sphere2-B on Oxford Flowers 256×256 for each k . Uniform sampling ( k=0 ) is the best choice.
Figure A21: Sampling angle α with 1-step generation using Sphere2-B trained on ImageNet 256×256 . Sampling angle decides the generation diversity by forcing the decoder away from the mean.
Model
α
Steps
Metric
Guidance angle γ , 256×256
Guidance angle γ , 512×512
0∘
5∘
10∘
15∘
20∘
30∘
40∘
50∘
0∘
5∘
10∘
15∘
20∘
30∘
40∘
50∘
Sphere2-B
85.5∘
2
FDr 6 ↓
9.48
9.03
8.77
8.63
8.55
8.55
8.67
8.89
10.28
9.85
9.56
9.45
9.33
9.26
9.35
9.52
gFID ↓
4.72
5.30
5.80
6.30
6.78
7.66
8.34
9.04
5.64
6.11
6.46
7.06
7.46
8.24
9.16
9.81
4
FDr 6 ↓
8.58
8.08
7.89
7.83
7.91
8.20
8.65
9.17
9.35
8.92
8.69
8.61
8.66
8.91
9.33
9.85
gFID ↓
4.89
5.58
6.33
7.00
7.66
8.66
9.54
10.02
6.04
6.62
7.36
7.94
8.50
9.54
10.51
11.23
Sphere2-L
84∘
2
FDr 6 ↓
6.42
6.31
6.28
6.30
6.39
6.62
6.94
7.32
7.50
7.37
7.28
7.30
7.31
7.44
7.61
7.84
Appendix
Table A6: Guidance angle γ of spherical CFG on ImageNet 256×256 and 512×512 . Sampling angle α is fixed per model, listed as 256/512 when it differs between resolutions.
Figure A23: Guidance angle γ of spherical CFG at 4 sampling steps on ImageNet 256×256 and 512×512 in Tab. A6 . The hollow marker on each curve is the FDr 6 -optimal γ , which the CFG rows of Tab. 2 use. Sampling angle α is 85.5∘ for Sphere2-B and 84∘ for Sphere2-L at both resolutions. See Fig. A24 for generated images along the γ sweep.
Figure A24: Spherical CFG for 2-step generation on ImageNet 256×256 with Sphere2-B. Each row is one random point decoded at increasing guidance angle γ . Guidance fixes the bad structures and strengthen the class semantics.
Encoder
d=N×D
Compression ratio
FDr 6 ↓
gFID ↓
1
2
4
1
2
4
DINOv3-small
162×384
2
3.92
3.14
3.04
10.29
8.24
7.51
DINOv3-small
162×256
3
4.53
3.77
3.61
17.67
13.64
12.27
DINOv3-base
162×768
1
4.28
3.58
3.51
11.47
9.28
8.69
DINOv3-base
162×384
2
5.17
4.48
4.46
18.18
15.67
15.78
SigLIP2-base
162×768
1
5.82
4.37
4.18
15.15
13.73
15.15
Appendix
Table A7: Frozen latent space of the encoder E with different pretrained encoders and latent dimensions on Oxford-Flowers 256×256 .
Feature ϕ
Layers
Params
FDr 6 ↓
gFID ↓
1
4
1
4
DINO-S/8
5 11
25.19M
6.16
4.63
24.86
17.17
3 7 11
37.79M
5.95
4.49
27.23
17.68
5 8 11
37.79M
6.20
4.72
24.28
17.80
2 5 8 11
50.38M
6.48
5.12
27.20
20.28
ConvNeXt V2-N
1 2
24.16M
6.51
4.88
26.80
18.60
Appendix
Table A8: Feature extractor ϕ and its layers used by the score loss on Oxford Flowers 256×256 . Each score model is a two-layer transformer with width 384. Params counts total parameters of all score models. Params of DINO-S/8 is 21.67M, and ConvNeXt V2-N fcmae is 14.98M.
Feat. ϕ
Transport (qmean,qstd)
FDr 6 ↓
gFID ↓
1
4
1
4
DINO-S/8
(−1.2,1.0)
4.54
3.93
13.69
10.89
(−1.5,1.0)
4.57
3.94
13.35
10.54
(−1.8,1.0)
4.61
3.96
13.19
10.37
(−1.8,0.8)
4.97
4.28
14.46
11.02
(−1.2,1.0)
3.82
3.09
11.40
8.28
Appendix
Table A9: Sigma distribution for transport procedure of the score loss on Flowers 256×256 .
Sphere2-B
Sphere2-L
architecture
encoder E
DINOv3-small (frozen, 29M)
latent
N=16×16 , D=384 , d=98,304
decoder D
ViT-B
ViT-L
depth
12
24
hidden dim
768
1024
Appendix
Table A10: Configurations of experiments. Values in a merged cell are shared by both models.
Model
timm identifier
Arch.
Dim
Input
Objective
Pooling
Inception-v3 ( Szegedy et al., 2016 )
inception_v3 (torch-fidelity)
CNN
2048
299
supervised
global avg pool
ConvNeXt-v2 ( Woo et al., 2023 )
convnextv2_base.fcmae_ft_in22k_in1k
CNN
1024
224
self-supervised
global avg pool
MAE ( He et al., 2022 )
vit_large_patch16_224.mae
ViT
1024
224
reconstructive
CLS token
DINOv2 ( Oquab et al., 2024 )
vit_large_patch14_dinov2.lvd142m
ViT
1024
256
contrastive
CLS token
SigLIP2 ( Tschannen et al., 2025 )
vit_so400m_patch16_siglip_256.v2_webli
ViT
1152
224
vision-language
CLS token
CLIP ( Radford et al., 2021 )
vit_large_patch14_clip_224.openai
ViT
1024
256
vision-language
CLS token
Appendix
Table A11: Representation models of FDr 6 at 256×256 . The set follows the setup of FDr 6 ( Yang et al., 2026 ) , which is originally for ImageNet 256×256 .
Model
timm identifier
Arch.
Dim
Input
Objective
Pooling
Inception-v3 ( Szegedy et al., 2016 )
inception_v3 (torch-fidelity)
CNN
2048
299
supervised
global avg pool
ConvNeXt-v2 ( Woo et al., 2023 )
convnextv2_base.fcmae_ft_in22k_in1k_384
CNN
1024
384
self-supervised
global avg pool
MAE ( He et al., 2022 )
vit_large_patch16_224.mae
ViT
1024
224
reconstructive
CLS token
DINOv2 ( Oquab et al., 2024 )
vit_large_patch14_dinov2.lvd142m
ViT
1024
518
contrastive
CLS token
SigLIP2 ( Tschannen et al., 2025 )
vit_so400m_patch16_siglip_512.v2_webli
ViT
1152
512
vision-language
CLS token
CLIP ( Radford et al., 2021 )
vit_large_patch14_clip_336.openai
ViT
1024
336
vision-language
CLS token
Appendix
Table A12: Representation models of FDr 6 at 512×512 . Each family uses the highest-resolution checkpoint. Inception-v3 and MAE have no such checkpoint and keep the ones from Tab. A11 .