No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse
Organizations: Adelaide Data Science Centre School of Mathematical Sciences Adelaide University Adelaide SA 5005, Australia
Abstract
Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator , computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric (), whereas -filtering yields unique trigrams, vocabulary, and repetition (all ). We validate as a cross-domain entropy proxy (, ) and collapse detector (, ) across 4domains, 2temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
Figures & tables
| Metric | Unfilt. | -filt. | -filt. | ( v. U) | ( v. U) | ( v. ) |
|---|---|---|---|---|---|---|
| (bits/word) | 0.79 | 1.16 | 1.43 | |||
| Distinct-3 (%) | – | % | +42 % | |||
| Vocabulary (%) | – | % | +30 % | |||
| Rep-4 (%) | – | % | -19 % |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.