Research
Stability and Diversity of Networked Self-Consuming Generative Ecosystems
Overview Research area: Machine learning theory — specifically the dynamics of generative models trained on synthetic data produced by other generative models (self-consuming training loops), analyzed

- arXiv
- 2610.09409
- Published
- 2026-10-07
- Authors
- Xiukun Wei, Yang Zhang, Xueru Zhang
AI summary
Overview
Research area: Machine learning theory — specifically the dynamics of generative models trained on synthetic data produced by other generative models (self-consuming training loops), analyzed through spectral graph theory.
Technical level: Advanced. The paper builds on fixed-point analysis, Jacobians, spectral radius, nonnegative matrix theory, algebraic connectivity, and Wasserstein-distance-based regularity assumptions. The experiments are standard, but the theory requires comfort with linear algebra and optimization.
One-sentence scope: The paper introduces a directed, weighted-graph framework for many interacting self-consuming generative models and derives conditions under which the system converges, plus bounds on how much inter-model diversity the interactions destroy.
What This Paper Is About
Prior work on "self-consuming" generative models — models retrained on data that includes their own synthetic output — mostly studies a single isolated model, or at most two interacting models with simplified dynamics. In real deployments, many models feed synthetic data into each other through complex pathways (for example, one model's captions or instruction data training another model), so the ecosystem is genuinely networked.
This paper models that ecosystem as a directed, weighted graph where nodes are generative models and edge weights describe what fraction of a model's synthetic training data comes from each peer. The goal is to characterize when the whole network's retraining dynamics settle down to a stable equilibrium, and how much the models' outputs converge toward one another as a result.
Key Contributions
-
A unified theoretical framework for networked self-consuming models. Generative models are represented as nodes in a directed, weighted graph with an arbitrary, row-stochastic interaction matrix W (self-loops allowed), heterogeneous model classes (Θ_i ≠ Θ_j), and heterogeneous real-data distributions (p_i^data varies across models).
-
Graph-structured convergence and stability analysis. The paper derives local stability conditions around an arbitrary fixed point that separate a per-model local amplification factor from a graph-level global propagation effect, generalizing the single-model stability threshold to arbitrary directed and weighted graphs.
-
Topology-dependent diversity analysis. It characterizes long-term diversity at the fixed point as a graph-structured linear smoothing of the real-data representations, giving an upper bound on system output diversity and showing that synthetic-data consumption contracts inter-model heterogeneity in a way that depends on the interaction topology.
-
Empirical validation. The predictions are tested on a synthetic 8-Gaussian dataset and on CIFAR-10 (60,000 images, 10 classes) across multiple generative model classes and several interaction topologies, with the observations reported as aligning with the analysis.
Main Findings
-
Stability is a system-level property, not just a local one. A single self-consuming model can fail to converge when its synthetic-data weight λ_i is large; in a network, stability also depends on the interaction graph W. Theorem 3.7 shows that if α_i > L_i ε_i for every model and ρ(D) < 1, where D := diag(β_1, …, β_K) W, then the fixed point Θ̄ is locally asymptotically stable, and iterates converge at asymptotic linear rate limsup ‖Θ^t − Θ̄‖_2^{1/t} ≤ ρ(J) ≤ ρ(D).
-
The local amplification factor β_i combines mixing weight and local curvature. β_i := λ_i M_i / (α_i + λ_i(α_i − L_i ε_i)), where M_i is the strongest cross-model sensitivity between model i and its parents, α_i is the strong-concavity constant, and ε_i is the distributional gap between model i's real data and its parents' output distributions.
-
A tractable sufficient condition. Corollary 3.8: if β_max ‖W‖_2 < 1, then ρ(D) < 1 and the fixed point is locally asymptotically stable; equivalently, when M_i ‖W‖_2 > α_i − L_i ε_i for all i, stability is guaranteed if λ_i < α_i / (M_i ‖W‖_2 − (α_i − L_i ε_i)). As cross-model sensitivity M_i and the graph's spectral norm ‖W‖_2 grow, the maximum allowable synthetic-data weight shrinks.
-
The norm-based condition can be conservative, so cycle structure matters. The spectral norm aggregates all interactions into one scalar, so graphs with the same ‖W‖_2 can behave differently. Proposition 3.10 shows ρ(D) ≥ max_m γ(C_m), where γ(C) is the geometric mean of the D-entries around a simple directed cycle — so a single dominant cycle can bottleneck the system. Corollary 3.11: if ρ(D) < 1, every simple directed cycle must satisfy ∏ β_i w_ij < 1, and every self-loop must satisfy β_i w_ii < 1.
-
Cycle length alone does not determine stability. If every model along a cycle contributes the same per-step gain (β_i w_ij = b), the cycle gain γ(C) = b is independent of the cycle's length.
-
Fixed-point outputs are a linear smoothing of real-data centroids. Proposition 4.4: F̄ = (I + ΛL)^{-1}[F_data + (I + Λ)R̄], where Λ = diag(λ_1, …, λ_K), L := I − W, and R̄ collects the per-model feature discrepancies. The operator (I + ΛL)^{-1} performs the smoothing.
-
An explicit diversity bound with a collapse coefficient. Theorem 4.5: D_out ≤ 2 ‖(I + ΛL)^{-1}‖_2^2 [D_data + (1/|V|) Σ_i (1+λ_i)^2 ‖r̄_i‖^2]. The factor ‖(I + ΛL)^{-1}‖_2^2 acts as a collapse coefficient mapping real-data diversity to output diversity along the non-consensus subspace.
-
No cross-model interaction preserves diversity. If Λ = 0 or the graph has no cross edges (w_ij = 0 for all i ≠ j), the operator does not average across models; and when each model's output matches its training mixture in feature space (r̄_i = 0 for all i), the bound reduces to D_out = D_data.
-
Stronger synthetic mixing has competing effects. Larger λ_i increases the averaging of (I + ΛL)^{-1}, pulling representations together, but it also amplifies the term (1 + λ_i)^2 ‖r̄_i‖^2 because each model leans more on imperfect synthetic data.
-
Imperfect representation adds topology-independent diversity. The discrepancy term (1/|V|) Σ_i (1+λ_i)^2 ‖r̄_i‖^2 contributes to D_out regardless of W, reflecting model error rather than meaningful inter-model disagreement. The paper bounds ‖r̄_i‖ ≤ (1 + 2λ_i)/(1 + λ_i) · ε_i.
-
Connectivity accelerates diversity collapse in the symmetric case. Corollary 4.6: for undirected W = W^T, homogeneous λ_i ≡ λ, and r̄_i = 0, D_out ≤ 1/(1 + λμ_2)^2 · D_data, where μ_2 is the algebraic connectivity (second smallest eigenvalue of L). Higher connectivity gives a smaller bound and stronger collapse, while a small μ_2 (weakly connected communities) slows collapse even in a densely connected network.
-
Prior work is extended, not replaced. The stability condition strictly generalizes the threshold of earlier single-model analyses, and the framework accommodates arbitrary directed and weighted graphs rather than the two-model or linearized settings studied previously.
Methodology in Plain English
Each generative model in the network is trained at each round on a mix of its own real data and synthetic samples drawn from its "parent" models in the previous round. The mixing is controlled by two things: how much synthetic data the model uses (λ_i), and how that synthetic data is split among parents (the row of W for that model). Models are fine-tuned from their previous parameters rather than retrained from scratch, which is why the analysis selects the maximizer closest to the previous iterate.
To study long-term behavior, the authors write the whole network's update as a single map 𝒢 from the stacked parameters to updated stacked parameters. A fixed point is a parameter vector unchanged by 𝒢. Because proving such a fixed point exists in general is hard for the coupled high-dimensional map, they instead analyze local behavior around any fixed point that does exist, using the Jacobian of 𝒢 there.
Under standard regularity assumptions (each model's real-data log-likelihood is strongly concave locally with constant α_i, and Hessians change in a controlled way — bounded by L_i — when the data distribution shifts), the Jacobian is bounded by a nonnegative matrix D built from the per-model amplification factors and the interaction graph. Checking whether the spectral radius ρ(D) is below 1 then gives a stability certificate, and cycle-based quantities give finer, topology-aware conditions.
For diversity, since models may have different architectures and parameter dimensions, the authors compare models in a shared feature space via a 1-Lipschitz map φ, and define system diversity as the variance of the per-model mean feature vectors. Solving the fixed-point relationship yields a linear system whose solution shows that outputs are a graph-driven smoothing of real-data centroids, which produces the diversity upper bound and the cleaner undirected-graph corollary in terms of algebraic connectivity.
Experiments use a synthetic 8-Gaussian mixture model and CIFAR-10. Four generative models are instantiated from two paradigms — denoising diffusion probabilistic models and continuous flow matching — with two variants each (VLB Diffusion, OT-CFM, iCFM, Hybrid Diffusion), initialized from released pretrained checkpoints with default hyperparameters; the synthetic experiments additionally include a GAN-based model. Every iteration, each model generates 10,000 samples and is fine-tuned on a mixed dataset of 10,000 samples, with the synthetic share set by α_i = λ_i/(λ_i + 1). Four interaction graphs are designed for controlled comparisons: System 1 is an isolated-retraining baseline; Systems 1 and 2 share the same spectral radius ‖W‖ = 1 but differ in structure; Systems 2 and 3 share structure but differ in edge weights; System 4 has shorter directed cycles. Distribution quality is measured with FID against real data using standard Inception feature embeddings, with precision and recall reported as a supplement, and diversity is measured with 512-dimensional features from an ImageNet-pretrained ResNet-18 with the final classification layer removed.
Why This Matters
Impact on research. The paper shifts the study of self-consuming generative models from isolated or two-model settings to arbitrary directed, weighted networks. It supplies a reusable analytical template — comparison matrix, spectral radius, cycle gains, smoothing operator, collapse coefficient — that connects self-consuming training dynamics to classical tools from nonnegative matrix and spectral graph theory, and it separates per-model effects from topology effects in a way prior work did not.
Real-world applications (examples discussed in the paper):
- Large language model fine-tuning pipelines that use synthetic instruction-response pairs, as with Orca, which is fine-tuned using large-scale synthetic instruction-response pairs generated by ChatGPT and GPT-4.
- Small-model training that leans on synthetic data from larger LLMs, as with Phi-3, which relies heavily on synthetic data generated by larger LLMs.
- Text-to-image training pipelines that use a model to rewrite training captions, as with DALL·E 3, which employs a dedicated captioning model to rewrite noisy prompts.
- Multi-model image generation pipelines such as PixArt-α, which uses LLaVA to refine its training data.
Industry relevance. Any organization operating multiple generative models — platforms hosting many models whose outputs end up in downstream training sets, or teams that fine-tune one model on another's output — faces the tradeoffs this paper formalizes: mixing more synthetic data can erode model-to-model diversity, and the interaction graph determines how quickly that erosion happens. The stability thresholds give a concrete way to reason about how much synthetic data is safe before retraining loops become unstable, and the diversity result warns that high connectivity between models can be self-defeating even when each model starts from different real data.
Future Directions
-
Global fixed-point existence. The paper deliberately studies local dynamics around an arbitrary fixed point rather than proving that fixed points exist in general, noting that existence typically requires global conditions making the retraining map a continuous self-map on a compact parameter space. Establishing and characterizing general fixed points for coupled multi-model systems remains open.
-
Diversity analysis beyond the symmetric case. The clean bound in terms of algebraic connectivity holds only for undirected graphs with homogeneous synthetic weight and vanishing feature discrepancy. Extending the topology-dependent interpretation to general directed, weighted, heterogeneous networks — where the global operator norm obscures structural features — is a natural next step.
-
Sources of error abstracted away by the theory. The assumptions abstract away finite-data effects and imperfect optimization. The experiments are framed as testing whether the qualitative theory survives realistic training conditions, which implies further work on how sampling noise, limited model expressivity, and optimization error feed into the discrepancy term r̄_i.
-
Richer graph structures. The experiments isolate graph structure, edge weights, and cycle configuration through four controlled systems, with additional graphs explored in an appendix. Characterizing behavior across broader classes of cycles, community structures, and heterogeneous topologies is a logical extension.
Target Audience
Researchers working on model collapse, synthetic data contamination, and the long-term behavior of generative model training loops; theorists interested in multi-agent or networked learning dynamics and spectral graph analysis; and practitioners at organizations that train or fine-tune multiple generative models whose outputs circulate into each other's training pipelines — particularly in LLM instruction tuning and text-to-image generation. Readers need a strong mathematical background to follow the derivations, though the qualitative messages about mixing strength, connectivity, and diversity are accessible to a broader technical audience.
Authors’ abstract
The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a self-consuming training loop. Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to isolated models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding networked self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system's long-term stability and diversity are shaped by each model's access to real data, cross-model data consumption, and the structure of the interaction graph.