Self-training generative models - the continued improvement of a model using its own outputs - is becoming increasingly important as high-quality training data becomes scarce. However, naively finetuning on model-generated samples leads to degradation through model collapse and the model autophagy disorder. Negative-guidance self-training methods turn this degradation into a useful signal, using a model finetuned on its own outputs to guide the original model toward improved generation. Existing methods, however, take the negative signal in standard model outputs as given. We instead ask whether this signal can be explicitly strengthened. We introduce Geometrically Modified Outputs (GMOs), which reweight the singular values of the generator's input-output Jacobian to increase the influence of its leading singular directions. This geometric modification amplifies the mode-seeking behavior and distortions of standard outputs, providing a stronger and more targeted negative signal for self-training. Across a range of one-step generative models, GMOs consistently improve the performance of negative-guidance methods, including Neon and SIMS, compared with using standard model outputs.
Figures & tables
Figure 1: Geometrically Modified Outputs (GMOs) improve the performance of one-step generative models through self-training. Here we present examples of GMOs for the IMM ( Zhou et al., 2025 ) one-step generative model trained on Imagenet256 ( Krizhevsky et al., 2012 ) . In the first panel, we show the model’s standard outputs; in the center panel, we show the perturbations of these outputs that yield the corresponding GMOs in the right panel (see Figure 11 for more examples). Using GMOs in self-training algorithms more effectively improves generative models.
Figure 2: The geometry of a generative model collapses under naïve self-training. Here we finetune a MeanFlow SiT-B/2 model pre-trained on ImageNet256 for 50 epochs on 30,000 of its own outputs. Throughout training, we monitor the FID ( Heusel et al., 2017 ) of the model (left), and its effective rank on a collection of 16 fixed latent vectors (right).
Figure 3: Here we consider computing spectral statistics from the input-output Jacobians of a MeanFlow one-step generator ( Geng et al., 2025 ) with a UNet architecture ( Ronneberger et al., 2015 ; Song & Ermon, 2019 ) trained on CIFAR10 ( Krizhevsky & Hinton, 2009 ) . Ten latent vectors are sampled, and E , σ1 , u1 , and v1 are either computed exactly or using approximations. Power iteration is implemented 30 times to obtain estimates for σ^1 , u^1 , and v^1 . Similarly, 100 samples are used to generate the Hutchinson estimate E^ . In the first panel, we record \nicefrac∣σ^1−σ1∣σ1 . In the second panel, we record 1−\nicefrac⟨u^1,u1⟩∥u^1∥2∥u1∥2 , and similarly for v1 . In the third panel, we record \nicefracE^−EE . In the fourth panel, we record the time taken to exactly compute the Jacobian and its singular value decomposition (Exact) and the time taken to perform 20 power iterations and 20 Hutchinson approximations (Approximate).
Figure 4: GMOs provide meaningful perturbations to a model’s outputs that improve the effectiveness of Neon. Here, we consider applying the Neon self-training algorithm to the MeanFlow ( Geng et al., 2025 ) SiT-B/2 model ( Ma et al., 2024 ) trained on ImageNet256 ( Krizhevsky et al., 2012 ) . We generated collections of 30,000 GMOs for alpha values 0.0 (standard outputs), 0.05 , and 0.2 . We then finetuned the pre-trained checkpoint with a maximum compute budget of 1.8×106 . At regular checkpoints, we apply Neon with different weight-merging parameters w in the range [0,2] . In the first and second panels, we show the minimum FID value achieved at each fine-tuning checkpoint with the corresponding weight-merging parameter w , respectively. In the third panel, we visually compare standard outputs to GMOs and the Gaussian baseline (see Figure 15 for more examples).
Figure 5: Neon with GMOs improves image quality and diversity. For the IMM 1-step generator ( Zhou et al., 2025 ) trained on ImageNet256 ( Krizhevsky et al., 2012 ) , we plot FID, precision, recall, density, and coverage ( Heusel et al., 2017 ; Kynkäänniemi et al., 2019 ) as a function of the GMO coefficient α . Here, α equal to zero corresponds to Neon ( Alemohammad et al., 2026 ) applied to standard outputs, while α>0 corresponds to Neon applied to α -GMOs. Each curve reports the per- α optimum over fine-tuning budget B and merge weight w . For more details, refer to Appendix D.2 .
Figure 6: GMOs are transferable across generative models of different sizes and inference strategies. In the left panel, we consider using ImageNet256 GMOs from a SiT-B/2 MeanFlow model to improve a SiT-L/2 MeanFlow model using Neon. We compare this to Neon applied directly to the SiT-L/2 model’s outputs and GMOs. A computational fine tuning budget of 1.2×106 is used in every case, and GMOs are generated with α equal to 0.1 . In the right panel, we consider using CIFAR10 GMOs from a one-step IMM model to improve a 2-step IMM model using Neon. We compare this to Neon applied directly to the two-step IMM model’s outputs and GMOs. For more experimental detail on this particular experiment, refer to Appendix D.1 .
Architecture
Base FID
Neon w/ Standard Outputs
Neon w/ GMOs
B
w
FID
B
w
α
FID
IMM
8.34
2.5×106(6.1×10−3%)
1.6
7.32(±0.07)
2.87×106(7.0×10−3%)
1.6
0.2
6.25(±0.03)
MeanFlow SiT-B/2
6.08
7.2×105(0.25%)
0.8
5.70(±0.01)
4.8×105(0.17%)
1.2
0.05
5.60(±0.01)
MeanFlow SiT-L/2
3.97
1.8×106(0.63%)
0.5
3.74(±0.02)
1.8×106(0.63%)
0.3
0.1
3.70(±0.02)
AlphaFlow SiT-B/2
5.55
4.8×105(0.16%)
0.6
5.36(±0.02)
4.8×105(0.16%)
0.8
0.1
5.16(±0.02)
AlphaFlow SiT-XL/2
2.93
3×105(0.1%)
1.2
2.64(±0.02)
3×105(0.1%)
1.1
0.05
2.59(±0.02)
Table 1: The utilization of GMOs in the Neon self-training algorithm yields better generative models. Here, we implement the Neon self-training algorithm on various one-step generative models trained on ImageNet256. We consider Neon on the standard output of these generators, or on the corresponding GMOs, across various compute levels (as a percentage of pre-training compute), weight-merging parameters w , and α values for GMOs (refer to Appendix D.2 for more details). Among the best-performing configurations, we report the average and standard deviation of the FID across five random seeds and 50,000 output samples. We provide a repository here containing the model checkpoints obtained using GMOs.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: GMOs amplify the “mode-seeking” nature of model outputs and their errors. Here we consider a collection of 2048 GMOs for α values 0.0 , 0.1 , 0.2 , 0.6 , and 1.0 generated using AlphaFlow ( Zhang et al., 2026 ) SiT-XL/2 ( Ma et al., 2024 ) model trained on ImageNet256 ( Krizhevsky et al., 2012 ) . In the top row, we visualize a t-SNE projection ( van der Maaten & Hinton, 2008 ) of the CLIP embeddings ( Radford et al., 2021 ) of the GMOs. In the bottom row, we visualize the GMOs for a specific latent vector.
Model
Standard Neon (% of pretraining)
Neon+GMOs (% of pretraining)
IMM
0.007%
0.01%
AlphaFlow-XL/2
0.10%
0.51%
AlphaFlow-B/2
0.16%
0.57%
MeanFlow-B/2
0.26%
0.61%
MeanFlow-L/2
0.63%
1.06%
Appendix
Table 2: Total pipeline cost as a percentage of pretraining compute, including the upfront cost of GMO generation.
Figure 8: GMOs can reduce the finetuning required to achieve optimal results using Neon. Here we apply Neon to a UNet generative model ( Ronneberger et al., 2015 ) training on CIFAR10 ( Krizhevsky & Hinton, 2009 ) using the MeanFlow framework ( Geng et al., 2025 ) . We monitor the value of FID for a range of compute levels B and weight-merging parameters w . With red markers, we indicate which combination of these hyperparameters yields the model with the lowest FID score.
Figure 9: The complete set of evaluation curves for the experiment shown in the right panel of Figure 6 , and described in Appendix D.1 .
Method
FID ↓
Base IMM
8.342±0.058
SIMS, standard outputs
7.591±0.057
SIMS, GMOs ( α=0.1 )
6.920±0.048
Appendix
Table 3: SIMS-style guidance on IMM ImageNet256. FID is reported as mean ± sample standard deviation across five evaluations of fixed checkpoints, using 50,000 samples per evaluation. Both guided variants use ω=1.6 .
Architecture
Method
Precision
Recall
Density
Coverage
IMM (DiT-XL/2)
Base
.585±.002
.647±.001
.652±.001
.662±.003
Neon-std
.598±.002
.654±.002
.695±.001
.693±.002
GMOs
.615±.003
.647±.002
.728±.003
.714±.002
MeanFlow SiT-B/2
Base
.717±.002
.452±.002
1.081±.006
.749±.001
Neon-std
.720±.001
.452±.003
1.110±.006
.755±.001
GMOs
.725±.001
.451±.003
1.124±.007
.759±.002
Appendix
Table 4: Precision, recall, density, and coverage (5 seeds, mean ± std). Density and coverage are maintained or improved from standard-outputs Neon to GMOs on every architecture.
Architecture
Base FID
Neon-std FID
GMOs FID
Welch p
GMOs’ reduction, % of Neon’s
IMM (DiT-XL/2)
8.34
7.32±.07
6.25±.03
∼10−9
96%
MeanFlow SiT-B/2
6.08
5.70±.01
5.60±.01
2.6×10−7
26%
MeanFlow SiT-L/2
3.97
3.74±.02
3.70±.02
0.013
17%
AlphaFlow SiT-B/2
5.55
5.36±.02
5.16±.02
2.6×10−7
105%
AlphaFlow SiT-XL/2
2.93
2.64±.02
2.59±.02
0.004
17%
Appendix
Table 5: Statistical significance of the GMO improvement over standard-outputs Neon (5 seeds, mean ± std). The final column reports the additional FID reduction of GMOs as a percentage of the reduction Neon itself achieves over the base model.
Figure 10: Hyperparameter sweeps described in Appendix D.2 that yielded the results in Table 1 .
Figure 11: Additional examples of those shown in Figure 1 at different GMO α levels.
Figure 12: Qualitative comparison on IMM ImageNet-256 (Part I).
Figure 12: Qualitative comparison on IMM ImageNet-256 (Part II).
Figure 13: In Figure 3 , we quantitatively demonstrate that power iteration can be used to accurately approximate the singular vectors of the input-output Jacobians of generative models. Here, we qualitatively support this by showing the exact (top) and approximate top left singular vectors of these input-output Jacobians.
Figure 14: Here we provide additional examples to that illustrated in the bottom row of Figure 7 .
Figure 15: Here we provide additional examples to complement the third panel of Figure 4 .