cs.LGFeb 10, 2026

ELROND: Exploring and decomposing intrinsic capabilities of diffusion models

Authors: Paweł Skierś, Emilia Kaczmarczyk, Tomasz Trzciński, Kamil Deja

Organizations: Warsaw University of Technology · IDEAS Research Institute

Abstract

A single text prompt passed to a diffusion model yields a wide range of visual outputs determined solely by a stochastic process, leaving users with no direct control over which semantic variations appear. Exploring this range is difficult: random search offers no guarantee of covering it, while prompt editing is coarse, as even a small change in wording can substantially alter the generated image. We argue that systematic exploration instead requires recovering how the model itself organizes the conditional distributions it can produce. We formalize this structure as a generative manifold, and present ELROND, a method for recovering its tangent space at a given conditioning. To that end, we collect gradients obtained by backpropagating the differences between stochastic realizations of a fixed prompt, and decompose them into interpretable directions using Principal Component Analysis or a Sparse Autoencoder. We show that our method recovers this subspace accurately in a controlled setting and validate it behaviorally on large-scale models. We also demonstrate that recovered structure enables broader exploration of the model's capabilities than methods relying on external representations, and mitigates mode collapse in distilled models without retraining.

Explore similar work

CardsList
  1. PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

    Sep 28, 2026Vighnesh Subramaniam, Boris Katz, Brian Cheung +3Diffusion SamplingVideo Generation

  2. Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds

    Aug 4, 2026Ning Zhu, An Chen, Mengfei Zhao +4Text-To-Image Diffusion ModelsDenoising Trajectory