Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of the resulting endpoints. We introduce JIVE, a training-free framework that enhances generative diversity by injecting velocity perturbations aligned with the leading right singular subspace of the generator's endpoint Jacobian. By leveraging this local geometric structure, JIVE provably maximizes endpoint diversity while preserving sample quality. To maintain practical efficiency, we compute these perturbation directions via matrix-free iterations rooted in classical numerical linear algebra, requiring only a small computational overhead. Across different benchmarks, JIVE boosts both pixel and feature-level diversity in few-step and one-step generation.
Figures & tables
Figure 1: At an intermediate trajectory point xt , JIVE injects a local velocity perturbation sampled from the leading right singular subspace of the endpoint Jacobian . This explicitly maximizes endpoint diversity , as theoretically guaranteed by Theorem 1 .
Figure 2: Qualitative comparison from matched initial latents. Visual comparison of eight samples generated with 4-step FLUX.1-schnell using the same initial seeds for the prompt “ A scientist ”. The samples of the original method (unmodified) have similar compositions, whereas JIVE produces distinct subjects and backgrounds.
Figure 3: In a toy example, we compare unmodified flow matching (Fig.(a)), a batch repulsion method (OSCAR) (Fig.(b)) and JIVE (Fig.(c)). Eight samples flow from the center toward sixteen Gaussian modes arranged on a ring. Blue arrows are the velocity field, orange arrows are velocity perturbations, and gray curves are the trajectories of the samples. In panel (a), many points flow to the same few modes. In panel (b), batch repulsion pushes trajectories further apart, but the repulsion direction does not maximize endpoint spread. In panel (c), JIVE finds perturbations maximizing endpoint diversity and pushes points towards new modes. See Appendix A.1 for details.
Figure 4: From local variation to set-level diversity. Three endpoints (blue) are perturbed within centered neighborhoods (orange). The rightmost curves show the original and perturbed endpoint densities. They have the same endpoint mean, but the perturbed density is broader, which increases the expected pairwise spread.
Methods
Seconds
Time ×
FLOPs ×
Mem. (GB)
One forward pass
0.307
1.00
1.00
22.3
JVP
0.617
2.01
2.00
22.3
Jt⊤Jt
1.302
4.24
5.04
35.4
Table 1: Computational efficiency of the proposed methods ( Jt⊤Jt vs. JVP) with one matrix iteration. This is evaluated on FLUX.1-schnell at 512×512 resolution on an A100.
Method
F-Vendi ↑
P-Vendi ↑
Pairwise L2↑
CLIP ↑
CLIP-IQA ↑
HPSv2 ↑
Standard ( n=504 )
(Jt⊤Jt) Unmodified sampler
3.76 ± 0.10
5.00 ± 0.05
323.69 ± 2.80
30.25 ± 0.13
0.668 ± 0.003
28.68 ± 0.14
(Jt⊤Jt) OSCAR
3.80 ± 0.10
4.98 ± 0.05
327.43 ± 2.81
30.27 ± 0.13
0.664 ± 0.003
28.66 ± 0.14
(Jt⊤Jt) JIVE(JVP)
4.73 ± 0.11
5.49 ± 0.03
455.83 ± 2.12
30.00 ± 0.13
0.608 ± 0.003
25.98 ± 0.14
(Jt⊤Jt) JIVE( Jt⊤Jt )
5.14 ± 0.11
5.60 ± 0.02
458.20 ± 1.75
29.83 ± 0.12
0.602 ± 0.003
25.62 ± 0.13
(Jt⊤Jt) CADS ( τ2=1.2 )
6.10 ± 0.16
5.24 ± 0.04
344.42 ± 2.14
28.96 ± 0.15
0.678 ± 0.003
27.70 ± 0.15
Table 2: Diversity and plug-in composition on PartiPrompts. We evaluate both JIVE(JVP) and JIVE( Jt⊤Jt ) on FLUX.1-schnell at four evaluation steps. We use prompts from PartiPrompts at 16 images per prompt. Reported values are mean ± SE over prompts. The norms used are 12 on both JIVE variants when alone, and 4 when combined with CADS ( τ2=1.2 ).
Figure 5: Diversity-fidelity tradeoffs on GenEval. We show tradeoff curves for CLIPScore and HPSv2 versus Vendi-DINO, averaged over all 553 GenEval prompts at 4 images per prompt. Additional tradeoff curves for PickScore and GenEval score are in Appendix Figure 9 .
Method
F-Vendi ↑
P-Vendi ↑
CLIP ↑
HPSv2 ↑
PickScore ↑
GenEval ↑
FLUX.1-schnell (1 step, 512×512 )
Unmodified sampler
1.940 ± 0.019
1.683 ± 0.013
33.189 ± 0.138
30.540 ± 0.109
23.330 ± 0.044
0.700
CADS ( τ2=1.2 )
2.243 ± 0.024
1.779 ± 0.013
32.497 ± 0.134
29.770 ± 0.120
23.026 ± 0.043
0.583
OSCAR
2.073 ± 0.021
1.680 ± 0.013
30.648 ± 0.113
29.037 ± 0.118
22.776 ± 0.045
0.667
Ours (JVP, FD)
2.200 ± 0.021
1.964 ± 0.018
32.932 ± 0.132
28.890 ± 0.119
22.945 ± 0.044
0.684
Ours ( Jt⊤Jt , FD)
2.272 ± 0.023
2.080 ± 0.016
32.909 ± 0.132
28.858 ± 0.115
22.910 ± 0.043
0.672
Table 3: Diversity-fidelity comparison on GenEval (553 prompts, 4 images per prompt). Both JIVE variants use one direction-estimation iteration with finite differences. JVP/ Jt⊤Jt use norms 26/8 on FLUX and 26/16 on Turbo, respectively. Bold indicates the highest diversity score within each model.
Figure 6: Diverse generation from inverted latents. Controlled ODE inversion maps a stroke image to a latent. Denoising this latent returns a single reconstruction, while applying a JIVE( J⊤J ) perturbation after the first denoising step yields distinct outputs from the same latent.
Dataset
Method
Controller
Norm
F-Vendi ↑
P-Vendi ↑
LPIPS d ↑
Source L2↓
DINO f ↑
KID × 1k ↓
classroom
baseline
η=0.5
N/A
4.67 ± 0.17
2.82 ± 0.06
0.684 ± 0.002
121.68 ± 1.34
0.051 ± 0.002
53.95
JIVE( Jt⊤Jt )
η=0.5
8
5.37 ± 0.16
2.98 ± 0.06
0.716 ± 0.002
126.12 ± 1.34
0.057 ± 0.002
54.90
kitchen
baseline
η=0.5
N/A
6.78 ± 0.18
2.33 ± 0.07
0.680 ± 0.003
114.09 ± 0.96
0.018 ± 0.003
40.62
JIVE( Jt⊤Jt )
η=0.5
8
7.40 ± 0.16
2.46 ± 0.07
0.723 ± 0.002
121.47 ± 1.02
0.023 ± 0.003
39.99
conference room
baseline
η=0.5
N/A
4.99 ± 0.16
2.69 ± 0.11
0.654 ± 0.002
119.87 ± 1.60
0.085 ± 0.004
35.12
JIVE( Jt⊤Jt )
η=0.5
8
5.80 ± 0.18
2.88 ± 0.12
0.696 ± 0.003
126.15 ± 1.67
0.086 ± 0.003
25.50
Table 4: Diverse edits from inverted latents. RF-Inversion on FLUX.1-dev (30 steps), with JIVE( Jt⊤Jt ) injected after the first reverse-step controller update ( ηt=0 , norm 8). F-Vendi uses DINOv2-base features while P-Vendi uses pixels. Source L2 is measured relative to the source stroke image. Values are mean ± SE over 32 sources. Orange denotes JIVE.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Eigenvector and singular vector alignment across models. The histograms show absolute cosine similarity of the leading eigenvector with the leading right and left singular vectors at t=0 ( n=500 per model). Blue/red denote right/left singular vectors throughout. Alignment is strong in multi-step models (left) and weaker in distilled models (right), especially with the right singular vector.
Figure 8: The leading singular values decay rapidly. Spectra for two distilled models, normalized by the largest singular value. The spectrum decays rapidly in both models, demonstrating that alignment with top directions allows for more endpoint displacement per unit perturbation
Model
Direction
E1
E2
E4
E8
E16
∥Jw∥2/σrms
FLUX.1-schnell
Random
1.5e-5
3.1e-5
6.1e-5
1.2e-4
2.4e-4
1.0
JVP
0.033
0.050
0.072
0.102
0.127
30.8
Jt⊤Jt
0.624
0.803
0.915
0.969
0.989
198.3
SD3.5-Large-Turbo
Random
3.8e-6
7.6e-6
1.5e-5
3.1e-5
6.1e-5
1.0
JVP
0.074
0.108
0.152
0.226
0.277
102.7
Jt⊤Jt
0.648
0.816
0.916
0.981
0.992
321.6
Appendix
Table 5: Energy Ek=∑i=1k∣⟨δ,vi⟩∣2 in the top- k right singular subspace for unit δ , and gain relative to σrms=(∑iσi2/d)1/2 . The random row is an analytic value ( k/d ). Both JVP and Jt⊤Jt are evaluated with one iteration.
Configuration
Seconds
Time × Base
FLOPs × Base
Peak Mem. (GB)
Baseline
0.307
1.00
1.00
22.3
JVP FD (forward)
0.617
2.01
2.00
22.3
JVP FD (central)
0.619
2.02
2.00
22.3
JVP exact
1.733
5.64
3.00
23.5
Jt⊤Jt FD (forward)
1.302
4.24
5.04
35.4
Jt⊤Jt exact
2.425
7.90
6.03
35.4
Appendix
Table 6: Runtime and memory overhead of Jacobian-vector and Jt⊤Jt products on FLUX.1-schnell at 512×512 resolution on an A100. Only the transformer is resident on the GPU. FD is finite difference.
Diversity
Fidelity
Method
F-Vendi ↑
P-Vendi ↑
CLIP ↑
HPSv2 ↑
PickScore ↑
GenEval ↑
FLUX.1-schnell (1 step, 512×512 )
Deterministic
1.940 ± 0.019
1.683 ± 0.013
33.189 ± 0.138
30.540 ± 0.109
23.330 ± 0.044
0.700
Extra denoising steps (5)
2.028 ± 0.021
1.710 ± 0.013
32.934 ± 0.134
30.483 ± 0.109
23.249 ± 0.043
0.686
Ours (JVP, FD)
2.200 ± 0.021
1.964 ± 0.018
32.932 ± 0.132
28.890 ± 0.119
22.945 ± 0.044
0.684
Ours ( Jt⊤Jt , FD)
2.272 ± 0.023
2.080 ± 0.016
32.909 ± 0.132
28.858 ± 0.115
22.910 ± 0.043
0.672
Appendix
Table 7: Comparing the low-cost version of JIVE with using extra denoising steps. Experiments were conducted using GenEval prompts at 4 images per prompt.
Figure 9: PickScore and GenEval score tradeoffs on GenEval. Companion to Figure 5 with the same settings: PickScore (left) and GenEval score (right) versus Vendi-DINO, averaged over 553 prompts. On both FLUX.1-schnell and SD3.5-Large-Turbo, both low-cost variants of JIVE have a better diversity-GenEval tradeoff than CADS or OSCAR.
Figure 10: Finite-difference and exact JVP estimates produce similar diversity-quality frontiers. Points are evaluated on Flux.1-schnell at 1 step using GenEval prompts at 4 images per prompt
Figure 11: Finite-difference and exact J⊤J estimates produce similar diversity-quality frontiers. Points are evaluated on Flux.1-schnell at 1 step using GenEval prompts at 4 images per prompt
Figure 12: Jacobian-guided perturbation versus random perturbation. Random directions yield almost none of the diversity gain, so the improvement comes from the choice of direction rather than from injected norm. Points are evaluated on Flux.1-schnell at 1 step using GenEval prompts at 4 images per prompt
Figure 13: Qualitative diversity across twelve prompts. Each prompt block shows six matched samples from the unperturbed sampler in the upper row and JIVE in the lower row. Both rows use the same initial seeds. Across varied subjects and compositions, JIVE produces broader changes in appearance, viewpoint, and style while preserving the prompt content. Prompts 1-4 are shown here, and prompts 5-12 continue on the following pages.
Figure 14: Qualitative diversity across twelve prompts (continued). Prompts 5-8. Upper rows show unperturbed samples and lower rows show JIVE samples generated from the same initial seeds.
Figure 15: Qualitative diversity across twelve prompts (continued). Prompts 9-12. Upper rows show unperturbed samples and lower rows show JIVE samples generated from the same initial seeds.
Table 8: Category-level results on the curated prompt set. Each entry is the mean over 10 prompts and 128 images per arm. Panels (a) and (b) cover four categories each. Shading marks the primary Feature Vendi criteria. Boldface denotes the best result within each category and tier. For all metrics, higher values are better.
Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse. Existing strategies for enhancing diversity predominantly focus on intervening during the generation trajectory. We identify a critical oversight that the standard Gaussian initialization often causes trajectories to collapse into dominant modes because it is agnostic to the guidance potential landscape. In this work, we formulate selecting the initial noise from a guidance potential posterior, which effectively re-weights the prior towards diversity-rich regions. To sample from this distribution efficiently, we introduce Diversity-inducing Initialization (DivIn), which leverages Langevin dynamics to actively navigate the initialization landscape, steering initial noise away from collapsing regions while anchoring them to the valid data manifold. Our method serves as an inference-time diversity enhancement compatible with both diffusion and flow matching models. Extensive experiments show that DivIn exhibits a superior performance in both class-to-image and text-to-image scenarios. Furthermore, we highlight that as DivIn is orthogonal to trajectory-based methods, combining them significantly expands the diversity-quality Pareto frontier beyond what either achieves in isolation.
Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low-density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training-free framework that refines intermediate states before each solver step. Our key finding reveals that the severity of this misguidance is quantified by the divergence of the marginal velocity field that is readily computable during inference with a well-optimized model. FDS exploits this signal to steer states toward less ambiguous regions. As a plug-and-play framework compatible with standard solvers and off-the-shelf flow backbones, FDS consistently improves fidelity across various generation tasks including text-to-image synthesis, and inverse problems.
State-of-the-art flow models generate stunning images from text or image prompts. However, they suffer from diversity collapse when generating multiple samples under the same conditioning. Existing methods address this issue via either latent guidance, which has limited effectiveness, or sample selection, which relies on external reward models that incur significant inference-time overhead. In this work, we introduce an efficient, training-free self-guidance mechanism to mitigate diversity collapse without requiring additional reward models. Specifically, we disperse the internal features of the flow model during batch generation with feature self-guidance. Further, to keep the features close to the manifold, we introduce a manifold regularization step that projects these dispersed features back onto the data manifold, ensuring diverse generation without sacrificing alignment with the input conditions. Our method integrates seamlessly as a plug-and-play module into pretrained flow models, adding only a marginal inference cost. Experiments demonstrate significant improvements in diversity while preserving fidelity across several conditional flow models, including multi-step and few-step text-to-image, depth-to-image, and reference image generation.
Pradhaan S Bhat, Rishubh Parihar, Abhijnya Bhat +1