The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a self-consuming training loop. Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to isolated models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding networked self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system's long-term stability and diversity are shaped by each model's access to real data, cross-model data consumption, and the structure of the interaction graph.
Figures & tables
Figure 1 : From networked self-consuming models to graph abstraction. (1) Left: A generative ecosystem in which each model is trained on its own real-data distribution and a weighted mixture of synthetic data produced by other models. (2) Right: Abstraction of the system as a directed, weighted interaction graph that captures effective synthetic-data propagation among models. (3) Top right: An example local retraining rule, where Model 3 is trained at iteration t on p3data+λ3∑j=1Kw3jpθjt−1 .
Figure 2 : Stability experiments when all models trained on full CIFAR-10. Each row corresponds to a different interaction graph. For each row, the left figure illustrates the direction and strength of synthetic data flow, while the right figures report the FID between generated samples and real data for each model. Results are shown under different synthetic mixing strengths, with αi=α shared across models.
Model
Model Type
Stability
Diversity
A
VLB Diffusion
{0,⋯,9}
{0,⋯,9}
B
OT-CFM
{3,4}
C
iCFM
{5,6}
D
Hybrid Diffusion
{7,8,9}
Table 1 : Models and real-data assignments. The same data for all models is used for stability; the heterogeneous assignment is used for diversity experiments.
Figure 3 : Diversity across models under heterogeneous data with different interaction graphs. Systems are defined in Fig. 2 . Each figure reports the evolution of diversity D across retraining rounds under different synthetic mixing ratios αi=α shared across models.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
System
∥W∥2
ρ(J)
remp
ρ(D)
maxCγ(C)
Sys. 1 (isolated)
1.00
0.900
0.900
0.900
0.900
Sys. 2 (weak ring)
1.00
0.606
0.606
0.700
0.450
Sys. 3 (strong ring)
1.00
0.582
0.582
0.674
0.604
Sys. 4 (dense)
1.00
0.600
0.600
0.900
0.225
Appendix
Table 2: Closed-form spectral validation at λ=1 ( s=0.5 ). The measured asymptotic rate remp matches the exact Jacobian spectral radius ρ(J) , while ρ(D) upper-bounds ρ(J) . The last column reports the largest cycle gain and verifies the lower bound in Prop. 3.10 .
System
scD
scJ
Sys. 1 (isolated)
0.556
0.556
Sys. 2 (weak ring)
0.714
0.825
Sys. 3 (strong ring)
0.742
0.860
Sys. 4 (dense)
0.556
0.833
Appendix
Table 3: Critical synthetic-data ratio s=λ/(1+λ) . The comparison-matrix threshold scD is a sufficient stability certificate and is therefore no larger than the exact Jacobian threshold scJ in all tested systems.
#
Model
Architecture
Stability
Diveristy
Retrain Steps
A
DDPM
MLP (hidden dimensions 128)
t∈{0,⋯,7}
t∈{0,4}
100
B
CFM
MLP
t∈{0,⋯,7}
t∈{2,6}
50
C
GAN
MLP
t∈{0,⋯,7}
t∈{3,5}
30
D
DDPM
MLP (hidden dimensions 256)
t∈{0,⋯,7}
t∈{1,7}
100
Appendix
Table 4: Model architectures and data access.
Figure 4: Stability experiments when all models trained on full dataset. Each row corresponds to a different interaction graph. For each row, the left figure illustrates the direction and strength of synthetic data flow, while the right figures report the Wasserstein distance ( WD ) between generated samples and real data for each model. Results are shown under different synthetic mixing strengths, with αi=α shared across models.
Figure 5: Stability experiments when all models trained on full dataset (continued).
Figure 6: Generated samples when all models trained on full dataset with different interaction graphs. All models use αi=α=1.0 .
Figure 7: Generated samples when all models trained on full dataset. All models use αi=α=0.2 .
Figure 8: Diversity across models under heterogeneous data with different interaction graphs. Each figure reports the system-level diversity metric D across retraining rounds under different synthetic mixing strengths α .
Figure 9: Generated samples under the heterogeneous data with different interaction graphs. All models use αi=α=1.0 .
Figure 10: Generated samples under the heterogeneous data with different interaction graphs. All models use αi=α=0.2 .
#
Model
Variants
Stability
Diversity
Retrain Steps
A
DDPM
VLB
class {0,⋯,9}
class {0,⋯,9}
100
B
CFM
OT-CFM
class {0,⋯,9}
class {3,4}
600
C
CFM
iCFM
class {0,⋯,9}
class {5,6}
600
D
DDPM
hybrid
class {0,⋯,9}
class {7,8,9}
100
Appendix
Table 5: Model architectures and data access.
Figure 11: Stability experiments when all models trained on full dataset. Each row corresponds to a different interaction graph. For each row, the left figure illustrates the direction and strength of synthetic data flow, while the right figures report the recall between generated samples and real data for each model. Results are shown under different synthetic mixing strengths, with αi=α shared across models.
Figure 12: Stability experiments when all models trained on full dataset. Each row corresponds to a different interaction graph. For each row, the left figure illustrates the direction and strength of synthetic data flow, while the right figures report the precision between generated samples and real data for each model. Results are shown under different synthetic mixing strengths, with αi=α shared across models.
Figure 13: Generated samples under the heterogeneous real data setting with different interaction graphs. All models use αi=α=1.0 .
Figure 14: Generated samples under the heterogeneous real data setting with different interaction graphs. All models use αi=α=0.2 .
Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective. Prior work suggests that such collapse is unavoidable without adding real data into the mix. We revisit this conclusion from an alignment perspective and show that collapse can be mitigated through curation based on multiple reward functions. We formalize the dynamics of recursive training under heterogeneous preferences and prove that, under certain conditions, the model converges to a stable distribution that allocates probability mass across competing high-reward regions. The limiting distribution preserves diversity and provably satisfies a weighted Nash bargaining solution, offering a formal interpretation of value aggregation in synthetic retraining loops.
Ali Falahati, Mohammad Mohammadi Amiri, Kate Larson +1
University of Waterloo · 2Rensselaer Polytechnic Institute.
Generative models are increasingly trained in self-consuming iterative loops, where users curate preferred samples from model-generated candidates and the curated samples are used to train future generations of the model. Prior work has largely assumed fixed user preferences, but in practice exposure to model outputs gradually reshapes what users perceive as desirable, creating a feedback loop in which model distributions and user preferences co-evolve. We take a first step toward understanding the long-term behavior of such coupled dynamics. We show that when training relies entirely on user-curated synthetic data, iterative curation amplifies initial biases and drives the system toward one of multiple singleton equilibria in which the instance holding an initial advantage eventually dominates. In contrast, injecting reference data into training at a sufficiently large rate fundamentally changes the dynamics and yields a unique globally attracting equilibrium. Building on this insight, we study how reference-data injection can be used to control long-term outcomes, and propose an efficient algorithm that jointly selects a reference distribution and its mixing weight to steer the coupled system toward equilibria that preserve desired attributes while minimizing data collection costs.
As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive degradation or collapse. In this article, we provide the first positive, rigorous theoretical results, to the best of our knowledge, showing that under model-agnostic mild conditions, the model converges to the true data-generating distribution. The convergence rate is the minimum of the model's intrinsic rate and the fraction of real data at each training iteration, revealing a phase transition between data-limited and model-limited regimes. We further show that, for biased real data, correcting the bias prevents the persistence and amplification of early bias over training iteration. Extensive experiments across simulations, real images and texts validate our theoretical framework, establishing quantitative conditions for long-term AI stability in contaminated environments.
Kevin Wang, Hongqian Niu, Didong Li
Department of Biostatistics, University of North Carolina at Chapel Hill