cs.CVOct 1, 2026

Sphere Encoder 2

Authors: Kaiyu Yue, Sean McLeish, Ruchit Rawal, Brian Bartoldson, Menglin Jia, Tom Goldstein

Organizations: University of Maryland · Lawrence Livermore National Laboratory · Cornell University

Abstract

Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.com/kaiyuyue/sphere2}{github.com/kaiyuyue/sphere2}.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Efficient Image Synthesis with Sphere Latent Encoder

    May 15, 2026Tung Do, Thuan Hoang Nguyen, Hao LiLarge-Scale Image GenerationLatent Diffusion Model

  2. HiRAE: Hierarchical Representation Autoencoding with Residual Budgets

    Sep 29, 2026Xuanyu Zhu, Yan Bai, Yang Shi +5Autoregressive Image Generation

  3. IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

    Jun 9, 2026Yitong Chen, Zijie Diao, Junke Wang +5Autoregressive Image GenerationMasked Autoencoders