Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of the resulting endpoints. We introduce JIVE, a training-free framework that enhances generative diversity by injecting velocity perturbations aligned with the leading right singular subspace of the generator's endpoint Jacobian. By leveraging this local geometric structure, JIVE provably maximizes endpoint diversity while preserving sample quality. To maintain practical efficiency, we compute these perturbation directions via matrix-free iterations rooted in classical numerical linear algebra, requiring only a small computational overhead. Across different benchmarks, JIVE boosts both pixel and feature-level diversity in few-step and one-step generation.
Figures & tables
Figure 1: At an intermediate trajectory point xt , JIVE injects a local velocity perturbation sampled from the leading right singular subspace of the endpoint Jacobian . This explicitly maximizes endpoint diversity , as theoretically guaranteed by Theorem 1 .
Figure 2: Qualitative comparison from matched initial latents. Visual comparison of eight samples generated with 4-step FLUX.1-schnell using the same initial seeds for the prompt “ A scientist ”. The samples of the original method (unmodified) have similar compositions, whereas JIVE produces distinct subjects and backgrounds.
Figure 3: In a toy example, we compare unmodified flow matching (Fig.(a)), a batch repulsion method (OSCAR) (Fig.(b)) and JIVE (Fig.(c)). Eight samples flow from the center toward sixteen Gaussian modes arranged on a ring. Blue arrows are the velocity field, orange arrows are velocity perturbations, and gray curves are the trajectories of the samples. In panel (a), many points flow to the same few modes. In panel (b), batch repulsion pushes trajectories further apart, but the repulsion direction does not maximize endpoint spread. In panel (c), JIVE finds perturbations maximizing endpoint diversity and pushes points towards new modes. See Appendix A.1 for details.
Figure 4: From local variation to set-level diversity. Three endpoints (blue) are perturbed within centered neighborhoods (orange). The rightmost curves show the original and perturbed endpoint densities. They have the same endpoint mean, but the perturbed density is broader, which increases the expected pairwise spread.
Methods
Seconds
Time ×
FLOPs ×
Mem. (GB)
One forward pass
0.307
1.00
1.00
22.3
JVP
0.617
2.01
2.00
22.3
Jt⊤Jt
1.302
4.24
5.04
35.4
Table 1: Computational efficiency of the proposed methods ( Jt⊤Jt vs. JVP) with one matrix iteration. This is evaluated on FLUX.1-schnell at 512×512 resolution on an A100.
Method
F-Vendi ↑
P-Vendi ↑
Pairwise L2↑
CLIP ↑
CLIP-IQA ↑
HPSv2 ↑
Standard ( n=504 )
(Jt⊤Jt) Unmodified sampler
3.76 ± 0.10
5.00 ± 0.05
323.69 ± 2.80
30.25 ± 0.13
0.668 ± 0.003
28.68 ± 0.14
(Jt⊤Jt) OSCAR
3.80 ± 0.10
4.98 ± 0.05
327.43 ± 2.81
30.27 ± 0.13
0.664 ± 0.003
28.66 ± 0.14
(Jt⊤Jt) JIVE(JVP)
4.73 ± 0.11
5.49 ± 0.03
455.83 ± 2.12
30.00 ± 0.13
0.608 ± 0.003
25.98 ± 0.14
(Jt⊤Jt) JIVE( Jt⊤Jt )
5.14 ± 0.11
5.60 ± 0.02
458.20 ± 1.75
29.83 ± 0.12
0.602 ± 0.003
25.62 ± 0.13
(Jt⊤Jt) CADS ( τ2=1.2 )
6.10 ± 0.16
5.24 ± 0.04
344.42 ± 2.14
28.96 ± 0.15
0.678 ± 0.003
27.70 ± 0.15
Table 2: Diversity and plug-in composition on PartiPrompts. We evaluate both JIVE(JVP) and JIVE( Jt⊤Jt ) on FLUX.1-schnell at four evaluation steps. We use prompts from PartiPrompts at 16 images per prompt. Reported values are mean ± SE over prompts. The norms used are 12 on both JIVE variants when alone, and 4 when combined with CADS ( τ2=1.2 ).
Figure 5: Diversity-fidelity tradeoffs on GenEval. We show tradeoff curves for CLIPScore and HPSv2 versus Vendi-DINO, averaged over all 553 GenEval prompts at 4 images per prompt. Additional tradeoff curves for PickScore and GenEval score are in Appendix Figure 9 .
Method
F-Vendi ↑
P-Vendi ↑
CLIP ↑
HPSv2 ↑
PickScore ↑
GenEval ↑
FLUX.1-schnell (1 step, 512×512 )
Unmodified sampler
1.940 ± 0.019
1.683 ± 0.013
33.189 ± 0.138
30.540 ± 0.109
23.330 ± 0.044
0.700
CADS ( τ2=1.2 )
2.243 ± 0.024
1.779 ± 0.013
32.497 ± 0.134
29.770 ± 0.120
23.026 ± 0.043
0.583
OSCAR
2.073 ± 0.021
1.680 ± 0.013
30.648 ± 0.113
29.037 ± 0.118
22.776 ± 0.045
0.667
Ours (JVP, FD)
2.200 ± 0.021
1.964 ± 0.018
32.932 ± 0.132
28.890 ± 0.119
22.945 ± 0.044
0.684
Ours ( Jt⊤Jt , FD)
2.272 ± 0.023
2.080 ± 0.016
32.909 ± 0.132
28.858 ± 0.115
22.910 ± 0.043
0.672
Table 3: Diversity-fidelity comparison on GenEval (553 prompts, 4 images per prompt). Both JIVE variants use one direction-estimation iteration with finite differences. JVP/ Jt⊤Jt use norms 26/8 on FLUX and 26/16 on Turbo, respectively. Bold indicates the highest diversity score within each model.
Figure 6: Diverse generation from inverted latents. Controlled ODE inversion maps a stroke image to a latent. Denoising this latent returns a single reconstruction, while applying a JIVE( J⊤J ) perturbation after the first denoising step yields distinct outputs from the same latent.
Dataset
Method
Controller
Norm
F-Vendi ↑
P-Vendi ↑
LPIPS d ↑
Source L2↓
DINO f ↑
KID × 1k ↓
classroom
baseline
η=0.5
N/A
4.67 ± 0.17
2.82 ± 0.06
0.684 ± 0.002
121.68 ± 1.34
0.051 ± 0.002
53.95
JIVE( Jt⊤Jt )
η=0.5
8
5.37 ± 0.16
2.98 ± 0.06
0.716 ± 0.002
126.12 ± 1.34
0.057 ± 0.002
54.90
kitchen
baseline
η=0.5
N/A
6.78 ± 0.18
2.33 ± 0.07
0.680 ± 0.003
114.09 ± 0.96
0.018 ± 0.003
40.62
JIVE( Jt⊤Jt )
η=0.5
8
7.40 ± 0.16
2.46 ± 0.07
0.723 ± 0.002
121.47 ± 1.02
0.023 ± 0.003
39.99
conference room
baseline
η=0.5
N/A
4.99 ± 0.16
2.69 ± 0.11
0.654 ± 0.002
119.87 ± 1.60
0.085 ± 0.004
35.12
JIVE( Jt⊤Jt )
η=0.5
8
5.80 ± 0.18
2.88 ± 0.12
0.696 ± 0.003
126.15 ± 1.67
0.086 ± 0.003
25.50
Table 4: Diverse edits from inverted latents. RF-Inversion on FLUX.1-dev (30 steps), with JIVE( Jt⊤Jt ) injected after the first reverse-step controller update ( ηt=0 , norm 8). F-Vendi uses DINOv2-base features while P-Vendi uses pixels. Source L2 is measured relative to the source stroke image. Values are mean ± SE over 32 sources. Orange denotes JIVE.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Eigenvector and singular vector alignment across models. The histograms show absolute cosine similarity of the leading eigenvector with the leading right and left singular vectors at t=0 ( n=500 per model). Blue/red denote right/left singular vectors throughout. Alignment is strong in multi-step models (left) and weaker in distilled models (right), especially with the right singular vector.
Figure 8: The leading singular values decay rapidly. Spectra for two distilled models, normalized by the largest singular value. The spectrum decays rapidly in both models, demonstrating that alignment with top directions allows for more endpoint displacement per unit perturbation
Model
Direction
E1
E2
E4
E8
E16
∥Jw∥2/σrms
FLUX.1-schnell
Random
1.5e-5
3.1e-5
6.1e-5
1.2e-4
2.4e-4
1.0
JVP
0.033
0.050
0.072
0.102
0.127
30.8
Jt⊤Jt
0.624
0.803
0.915
0.969
0.989
198.3
SD3.5-Large-Turbo
Random
3.8e-6
7.6e-6
1.5e-5
3.1e-5
6.1e-5
1.0
JVP
0.074
0.108
0.152
0.226
0.277
102.7
Jt⊤Jt
0.648
0.816
0.916
0.981
0.992
321.6
Appendix
Table 5: Energy Ek=∑i=1k∣⟨δ,vi⟩∣2 in the top- k right singular subspace for unit δ , and gain relative to σrms=(∑iσi2/d)1/2 . The random row is an analytic value ( k/d ). Both JVP and Jt⊤Jt are evaluated with one iteration.
Configuration
Seconds
Time × Base
FLOPs × Base
Peak Mem. (GB)
Baseline
0.307
1.00
1.00
22.3
JVP FD (forward)
0.617
2.01
2.00
22.3
JVP FD (central)
0.619
2.02
2.00
22.3
JVP exact
1.733
5.64
3.00
23.5
Jt⊤Jt FD (forward)
1.302
4.24
5.04
35.4
Jt⊤Jt exact
2.425
7.90
6.03
35.4
Appendix
Table 6: Runtime and memory overhead of Jacobian-vector and Jt⊤Jt products on FLUX.1-schnell at 512×512 resolution on an A100. Only the transformer is resident on the GPU. FD is finite difference.
Diversity
Fidelity
Method
F-Vendi ↑
P-Vendi ↑
CLIP ↑
HPSv2 ↑
PickScore ↑
GenEval ↑
FLUX.1-schnell (1 step, 512×512 )
Deterministic
1.940 ± 0.019
1.683 ± 0.013
33.189 ± 0.138
30.540 ± 0.109
23.330 ± 0.044
0.700
Extra denoising steps (5)
2.028 ± 0.021
1.710 ± 0.013
32.934 ± 0.134
30.483 ± 0.109
23.249 ± 0.043
0.686
Ours (JVP, FD)
2.200 ± 0.021
1.964 ± 0.018
32.932 ± 0.132
28.890 ± 0.119
22.945 ± 0.044
0.684
Ours ( Jt⊤Jt , FD)
2.272 ± 0.023
2.080 ± 0.016
32.909 ± 0.132
28.858 ± 0.115
22.910 ± 0.043
0.672
Appendix
Table 7: Comparing the low-cost version of JIVE with using extra denoising steps. Experiments were conducted using GenEval prompts at 4 images per prompt.
Figure 9: PickScore and GenEval score tradeoffs on GenEval. Companion to Figure 5 with the same settings: PickScore (left) and GenEval score (right) versus Vendi-DINO, averaged over 553 prompts. On both FLUX.1-schnell and SD3.5-Large-Turbo, both low-cost variants of JIVE have a better diversity-GenEval tradeoff than CADS or OSCAR.
Figure 10: Finite-difference and exact JVP estimates produce similar diversity-quality frontiers. Points are evaluated on Flux.1-schnell at 1 step using GenEval prompts at 4 images per prompt
Figure 11: Finite-difference and exact J⊤J estimates produce similar diversity-quality frontiers. Points are evaluated on Flux.1-schnell at 1 step using GenEval prompts at 4 images per prompt
Figure 12: Jacobian-guided perturbation versus random perturbation. Random directions yield almost none of the diversity gain, so the improvement comes from the choice of direction rather than from injected norm. Points are evaluated on Flux.1-schnell at 1 step using GenEval prompts at 4 images per prompt
Figure 13: Qualitative diversity across twelve prompts. Each prompt block shows six matched samples from the unperturbed sampler in the upper row and JIVE in the lower row. Both rows use the same initial seeds. Across varied subjects and compositions, JIVE produces broader changes in appearance, viewpoint, and style while preserving the prompt content. Prompts 1-4 are shown here, and prompts 5-12 continue on the following pages.
Figure 14: Qualitative diversity across twelve prompts (continued). Prompts 5-8. Upper rows show unperturbed samples and lower rows show JIVE samples generated from the same initial seeds.
Figure 15: Qualitative diversity across twelve prompts (continued). Prompts 9-12. Upper rows show unperturbed samples and lower rows show JIVE samples generated from the same initial seeds.
Table 8: Category-level results on the curated prompt set. Each entry is the mean over 10 prompts and 128 images per arm. Panels (a) and (b) cover four categories each. Shading marks the primary Feature Vendi criteria. Boldface denotes the best result within each category and tier. For all metrics, higher values are better.