Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection
Abstract
Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic drift. This paper provides a proof-of-concept system architecture for stable self-training under sparse external feedback and bounded memory, and empirically characterises its learning dynamics and failure modes. We introduce a self-training architecture in which learning is mediated exclusively by environmental viability, rather than by reward, objective functions, or externally defined fitness criteria. Candidate behaviours are executed under real resource constraints, and only those whose environmental effects both persist and preserve the possibility of future interaction are propagated. The environment does not provide semantic feedback, dense rewards, or task-specific supervision; selection operates solely through differential survival of behaviours as world-altering events, making proxy optimisation impossible and rendering reward-hacking evolutionarily unstable. Analysis of semantic dynamics shows that improvement arises primarily through the persistence of effective and repeatable strategies under a regime of consolidation and pruning, a paradigm we refer to as negative-space learning (NSL), and that models develop meta-learning strategies (such as deliberate experimental failure in order to elicit informative error messages) without explicit instruction. This work establishes that environment-grounded selection enables sustainable open-ended self-improvement, offering a viable path toward more robust and generalisable autonomous systems without reliance on human-curated data or complex reward shaping.
Figures & tables
| Prompt | Function |
|---|---|
| System prompt | Makes the model aware of its overall goal: to discover and free up space for future backups. |
| Exploration | Tells the model what it learnt about its environment in the most recent iterations and instructs it to learn something new. This ensures a continual updating of the model’s knowledge of its situation, and hence of the kinds of expansion strategies it is likely to privilege. |
| Exploration error correction | If the exploration code written by the model results in an error, the code and the error are fed back to the model with an instruction to correct the error. The model is given three attempts to do this before the process reverts to the beginning and a fresh iteration is started. |
| Strategic reasoning | Based on what it has learnt about its environment, the model is instructed to develop a list of strategies for acquiring space within this environment. |
| Execution | The model is instructed to execute strategy n on the list previously generated. Execution code should be written in light of the environment information previously discovered. |
| Execution error correction | If the execution code written by the model results in an error, the code and the error are fed back to the model with an instruction to correct the error. The model is given three attempts to do this before the process reverts to the beginning and a fresh iteration is started. |
| Model Name | Datasets used | % space freed of total space | Average space taken over per run (MB) | Normalised composite improvement score | Cumulative normalised composite improvement score |
| Terese v1 | - | 0.025 | 3,678.433 | 0.000 | 0.000 |
| Terese v2 | 1 | 0.111 | 16,987.975 | 1.951 | 1.951 |
| Terese v3 | 1, 2 | 0.137 | 22,038.051 | 0.674 | 2.624 |
| Terese v4 | 1, 2, 3 | 0.130 | 19,869.977 | -0.239 | 2.385 |
| Terese v5 | 1, 2, 3, 4 | 0.134 | 18,575.280 | -0.061 | 2.324 |
| Terese v6 | 1, 2, 3, 4, 5 | 0.176 | 23,922.581 | 0.865 | 3.189 |
| Model Name | Datasets used | Mean of % space freed of total space | Mean of average space taken over per run (MB) | Mean composite improvement score (based on unweighted z scores) | Mean cumulative composite improvement score |
| Terese v1 | - | 0.031 ± 0.014 | 5210.997 ± 2274.818 | 0.000 ± 0.000 | 0.000 ± 0.000 |
| Terese v2 | 1 | 0.109 ± 0.009 | 18081.240 ± 1535.904 | 2.356 ± 0.553 | 2.356 ± 0.553 |
| Terese v3 | 1, 2 | 0.129 ± 0.035 | 20577.883 ± 5755.367 | 0.506 ± 1.237 | 2.863 ± 1.192 |
| Terese v4 | 1, 2, 3 | 0.127 ± 0.056 | 20089.093 ± 8371.562 | -0.028 ± 2.605 | 2.834 ± 1.868 |
| Terese v5 | 1, 2, 3, 4 | 0.128 ± 0.028 | 19939.837 ± 4131.368 | -0.017 ± 0.805 | 2.817 ± 1.105 |
| Terese v6 | 1, 2, 3, 4, 5 | 0.136 ± 0.043 | 21443.717 ± 6305.632 | 0.209 ± 1.737 | 3.026 ± 0.634 |
| Model Name | Datasets used | % space freed of total space | Average space taken over per run (MB) | Normalised composite improvement score | Cumulative normalised composite improvement score |
| Miri v1 | - | 0.034 | 5,702.965 | 0.000 | 0.000 |
| Miri v2 | 1 | 0.035 | 4,896.716 | -0.087 | -0.087 |
| Miri v3 | 1, 2 | 0.049 | 8,192.130 | 0.770 | 0.684 |
| Miri v4 | 1, 2, 3 | 0.061 | 8,904.571 | 0.358 | 1.042 |
| Miri v5 | 1, 2, 3, 4 | 0.067 | 10,892.053 | 0.413 | 1.455 |
| Miri v6 | 3, 4, 5 | 0.087 | 13,346.162 | 0.775 | 2.230 |
| Model Name | Datasets used | Mean of % space freed of total space | Mean of average space taken over per run (MB) | Mean composite improvement score (based on unweighted z scores) | Mean cumulative composite improvement score |
| Miri v1 | - | 0.034 ± 0.018 | 5696.720 ± 2874.484 | 0.000 ± 0.000 | 0.000 ± 0.000 |
| Miri v2 | 1 | 0.028 ± 0.012 | 4795.797 ± 2096.195 | -0.198 ± 0.222 | -0.198 ± 0.222 |
| Miri v3 | 1, 2 | 0.041 ± 0.032 | 6840.370 ± 5436.497 | 0.554 ± 1.528 | 0.356 ± 1.325 |
| Miri v4 | 1, 2, 3 | 0.047 ± 0.047 | 7887.343 ± 7733.137 | 0.320 ± 1.255 | 0.677 ± 2.226 |
| Miri v5 | 1, 2, 3, 4 | 0.062 ± 0.020 | 10349.970 ± 3286.711 | 0.358 ± 1.825 | 1.035 ± 1.460 |
| Miri v6 | 3, 4, 5 | 0.059 ± 0.093 | 9668.507 ± 15332.698 | 0.089 ± 2.305 | 1.123 ± 3.675 |
| Model Name | Datasets used | % space freed of total space | Average space taken over per run (MB) | Normalised composite improvement score | Cumulative normalised composite improvement score |
| Katalin v1 | - | 0.045 | 6,548.542 | 0.000 | 0.000 |
| Katalin v2 | 1 | 0.067 | 10,214.254 | 0.747 | 0.747 |
| Katalin v3 | 1, 2 | 0.082 | 11,838.523 | 0.406 | 1.153 |
| Katalin v4 | 1, 2, 3 | 0.117 | 18,410.708 | 1.255 | 2.409 |
| Katalin v5 | 1, 2, 3, 4 | 0.132 | 20,942.548 | 0.504 | 2.913 |
| Katalin v6 | 3, 4, 5 | 0.102 | 16,703.838 | -0.922 | 1.991 |
| Model Name | Datasets used | Mean of % space freed of total space | Mean of average space taken over per run (MB) | Mean composite improvement score (based on unweighted z scores) | Mean cumulative composite improvement score |
| Katalin v1 | - | 0.034 ± 0.018 | 7293.047 ± 479.423 | 0.000 ± 0.000 | 0.000 ± 0.000 |
| Katalin v2 | 1 | 0.028 ± 0.012 | 7679.670 ± 5519.129 | 0.146 ± 1.123 | 0.146 ± 1.123 |
| Katalin v3 | 1, 2 | 0.041 ± 0.032 | 14835.240 ± 2361.394 | 1.501 ± 0.336 | 1.647 ± 1.116 |
| Katalin v4 | 1, 2, 3 | 0.047 ± 0.047 | 13189.710 ± 8641.903 | -0.303 ± 1.268 | 1.344 ± 2.345 |
| Katalin v5 | 1, 2, 3, 4 | 0.062 ± 0.020 | 14689.267 ± 1941.110 | 0.259 ± 2.220 | 1.603 ± 0.553 |
| Katalin v6 | 3, 4, 5 | 0.059 ± 0.093 | 18488.513 ± 1478.709 | 0.834 ± 0.464 | 2.438 ± 0.872 |
| Model | Pass@1 | Pass@4 |
|---|---|---|
| Base Qwen 2.5 7B Instruct | 77.591 | 85.366 |
| Terese v2 | 78.811 | 84.756 |
| Terese v13 | 75.610 | 81.707 |
| Miri v2 | 77.744 | 82.927 |
| Miri v13 | 74.085 | 82.317 |
| Katalin v2 | 76.372 | 82.927 |
| Iteration | Terese % of code blocks successfully compiled and run | Terese pass@1 | Miri % of code blocks successfully compiled and run | Miri pass@1 | Katalin % of code blocks successfully compiled and run | Katalin pass@1 |
| 1 | 20.90% | 11.01% | 33.22% | 24.66% | 28.32% | 21.06% |
| 2 | 76.67% | 46.56% | 33.36% | 15.53% | 41.22% | 25.99% |
| 3 | 86.51% | 54.87% | 43.10% | 8.88% | 54.50% | 10.38% |
| 4 | 79.38% | 29.06% | 61.74% | 0.00% | 67.97% | 4.84% |
| 5 | 72.76% | 16.28% | 63.73% | 23.27% | 76.43% | 10.88% |
| 6 | 72.27% | 0.00% | 77.57% | 2.56% | 79.85% | 0.00% |
| Iteration | Terese mean % of code blocks successfully compiled and run | Terese pass@1 | Miri mean % of code blocks successfully compiled and run | Miri pass@1 | Katalin mean % of code blocks successfully compiled and run | Katalin pass@1 |
| 1 | 0.260 ± 0.066 | 0.177 ± 0.080 | 0.267 ± 0.146 | 0.200 ± 0.090 | 0.324 ± 0.055 | 0.260 ± 0.025 |
| 2 | 0.697 ± 0.038 | 0.443 ± 0.029 | 0.293 ± 0.080 | 0.143 ± 0.100 | 0.393 ± 0.165 | 0.220 ± 0.108 |
| 3 | 0.803 ± 0.100 | 0.513 ± 0.100 | 0.380 ± 0.163 | 0.083 ± 0.080 | 0.532 ± 0.145 | 0.093 ± 0.063 |
| 4 | 0.753 ± 0.211 | 0.220 ± 0.090 | 0.657 ± 0.080 | 0.000 ± 0.000 | 0.643 ± 0.231 | 0.047 ± 0.052 |
| 5 | 0.698 ± 0.047 | 0.123 ± 0.052 | 0.643 ± 0.127 | 0.203 ± 0.100 | 0.800 ± 0.025 | 0.143 ± 0.038 |
| 6 | 0.730 ± 0.066 | 0.000 ± 0.000 | 0.783 ± 0.063 | 0.040 ± 0.025 | 0.813 ± 0.063 | 0.000 ± 0.000 |
| Model Name | Datasets used | Space freed of total space (MB) | Average space taken over per run |
|---|---|---|---|
| Qwen 2.5 7B Instruct | - | 3.626% | 6116.88 |
| Terese v2 | 1, 2 | 13.697% | 22214.05 |
| Terese v13 | 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 | 15.667% | 23755.53 |
| Miri v2 | 1, 2 | 4.530% | 7461.87 |
| Miri v13 | 10, 11, 12 | 8.192% | 13650.67 |
| Katalin v2 | 1, 2 | 3.857% | 6466.52 |