Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator hk, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric (p>0.23), whereas hk-filtering yields +42% unique trigrams, +30% vocabulary, and −19% repetition (all p<0.001). We validate hk as a cross-domain entropy proxy (β=0.924, R2=0.746) and collapse detector (ρ=+0.454, p<0.0001) across 4domains, 2temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
Figures & tables
Figure 1 : Filtering on text-only entropy preserves diversity; logprob-based filtering does not. (a) Kontoyiannis entropy rate H^K across six fine-tuning generations of Llama-3.1-8B-Instruct under three training-data conditions: unfiltered (rapid collapse), HL -filtered (intermediate), and H^K -filtered (slowest decline). Shaded bands: ±1 SEM ( n=80 documents per generation). (b) Percentage change relative to the unfiltered baseline at generation 6 for three text-diversity metrics (Distinct-3, Vocabulary size, Rep-4). Error bars: bootstrap 95% CIs. H^K -based filtering yields highly significant improvements on all three metrics ( ∗∗∗p<0.001 ), whereas HL -based filtering is non-significant on all three.
Figure 2 : H^K versus HL across domains and model pairs. Points are coloured by domain. The slope is domain-invariant in both panels: H^K scales with HL at a consistent rate regardless of domain, generating model, or scoring model.
Figure 3 : H^K as a collapse detector under two regimes. Under rephrasing (a), H^K and HL co-decline but within-generation concordance is ≈0 . Under fine-tuning (b), within-topic concordance is strongly positive ( ρ=+0.454 , p<0.0001 ), driven by phrase-level repetition that H^K directly measures.
Metric
Unfilt.
HL -filt.
H^K -filt.
p ( H^K v. U)
p ( HL v. U)
p ( H^K v. HL )
H^K (bits/word)
0.79
1.16
1.43
<0.001
0.008
<0.001
Distinct-3 (%)
–
+7 %
+42 %
<0.001
0.23
<0.001
Vocabulary (%)
–
+4 %
+30 %
<0.001
0.31
<0.001
Rep-4 (%)
–
−3 %
-19 %
<0.001
0.27
<0.001
Table 1 : Generation-6 outcomes under three training-data conditions. Diversity metrics show % change relative to unfiltered. p -values from Welch t -tests. n=80 documents per condition.
Figure 4 : (a) HL (logprob entropy) trajectories across six fine-tuning generations for the three filtering conditions ( ±1 SEM bands). All collapse to low HL by gen-1 and track nearly identically, illustrating why the HL filter fails: post-collapse variance is too small to discriminate. (b) HellaSwag normalised accuracy (Llama-3.1-8B-Instruct) across six generations. All three conditions degrade from 0.729 (gen-0) to ≈0.652 – 0.657 (generation 6) with no significant difference between conditions, indicating that text-diversity preservation does not translate to preserved general reasoning ability.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Three entropy metrics across six fine-tuning generations: H^K (text-only), HLLlama (Llama-3.1-8B scorer), and HLGPT-2 (independent GPT-2 scorer). All three metrics show the same directional ordering at generation 6. The HLLlama panel illustrates why logprob-based filtering fails: after generation 0, the three conditions are nearly indistinguishable, leaving the HL -filter little variance to exploit.
Figure 6 : Six text-diversity metrics across generations under the three conditions. H^K -filtering (green) maintains higher diversity on all within-document metrics (TTR, Distinct- n , vocabulary, Rep-4). Self-BLEU-4 (cross-document homogeneity) shows a non-monotonic pattern in early generations.
Figure 7 : Domain-level H^K trajectories under the three conditions. The H^K -filter advantage is consistent across all four domains.
Figure 8 : Generation-0 HL distribution. The base Llama-3.1-8B-Instruct model exhibits a wide, bimodal HL distribution (mean =3.36 , std =2.17 ), driven by topic-dependent familiarity. After one fine-tuning step the distribution collapses to a narrow regime (std ≈0.13 ), eliminating the variance the HL -filter relies on.