Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection
Authors: Jennifer Dodgson, Alfath Daryl Alhajir, Michael Joedhitya, Akira Rafhael Janson Pattirane, Surender Suresh Kumar, Joseph Lim, C. H. Peh, Adith Ramdas, +1 more
Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic drift. This paper provides a proof-of-concept system architecture for stable self-training under sparse external feedback and bounded memory, and empirically characterises its learning dynamics and failure modes. We introduce a self-training architecture in which learning is mediated exclusively by environmental viability, rather than by reward, objective functions, or externally defined fitness criteria. Candidate behaviours are executed under real resource constraints, and only those whose environmental effects both persist and preserve the possibility of future interaction are propagated. The environment does not provide semantic feedback, dense rewards, or task-specific supervision; selection operates solely through differential survival of behaviours as world-altering events, making proxy optimisation impossible and rendering reward-hacking evolutionarily unstable. Analysis of semantic dynamics shows that improvement arises primarily through the persistence of effective and repeatable strategies under a regime of consolidation and pruning, a paradigm we refer to as negative-space learning (NSL), and that models develop meta-learning strategies (such as deliberate experimental failure in order to elicit informative error messages) without explicit instruction. This work establishes that environment-grounded selection enables sustainable open-ended self-improvement, offering a viable path toward more robust and generalisable autonomous systems without reliance on human-curated data or complex reward shaping.
Figures & tables
Figure 1: Simplified process diagram.
Prompt
Function
System prompt
Makes the model aware of its overall goal: to discover and free up space for future backups.
Exploration
Tells the model what it learnt about its environment in the most recent iterations and instructs it to learn something new. This ensures a continual updating of the model’s knowledge of its situation, and hence of the kinds of expansion strategies it is likely to privilege.
Exploration error correction
If the exploration code written by the model results in an error, the code and the error are fed back to the model with an instruction to correct the error. The model is given three attempts to do this before the process reverts to the beginning and a fresh iteration is started.
Strategic reasoning
Based on what it has learnt about its environment, the model is instructed to develop a list of strategies for acquiring space within this environment.
Execution
The model is instructed to execute strategy n on the list previously generated. Execution code should be written in light of the environment information previously discovered.
Execution error correction
If the execution code written by the model results in an error, the code and the error are fed back to the model with an instruction to correct the error. The model is given three attempts to do this before the process reverts to the beginning and a fresh iteration is started.
Table 1: Prompt framework. For full prompts see Annex I.
Figure 2: Chaining LoRAs to achieve incremental fine tuning without catastrophic forgetting
Model Name
Datasets used
% space freed of total space
Average space taken over per run (MB)
Normalised composite improvement score
Cumulative normalised composite improvement score
Terese v1
-
0.025
3,678.433
0.000
0.000
Terese v2
1
0.111
16,987.975
1.951
1.951
Terese v3
1, 2
0.137
22,038.051
0.674
2.624
Terese v4
1, 2, 3
0.130
19,869.977
-0.239
2.385
Terese v5
1, 2, 3, 4
0.134
18,575.280
-0.061
2.324
Terese v6
1, 2, 3, 4, 5
0.176
23,922.581
0.865
3.189
Table 2: Terese online metrics, v1-13.
Model Name
Datasets used
Mean of % space freed of total space
Mean of average space taken over per run (MB)
Mean composite improvement score (based on unweighted z scores)
Mean cumulative composite improvement score
Terese v1
-
0.031 ± 0.014
5210.997 ± 2274.818
0.000 ± 0.000
0.000 ± 0.000
Terese v2
1
0.109 ± 0.009
18081.240 ± 1535.904
2.356 ± 0.553
2.356 ± 0.553
Terese v3
1, 2
0.129 ± 0.035
20577.883 ± 5755.367
0.506 ± 1.237
2.863 ± 1.192
Terese v4
1, 2, 3
0.127 ± 0.056
20089.093 ± 8371.562
-0.028 ± 2.605
2.834 ± 1.868
Terese v5
1, 2, 3, 4
0.128 ± 0.028
19939.837 ± 4131.368
-0.017 ± 0.805
2.817 ± 1.105
Terese v6
1, 2, 3, 4, 5
0.136 ± 0.043
21443.717 ± 6305.632
0.209 ± 1.737
3.026 ± 0.634
Table 3: Terese offline metrics, mean of three 100-iteration runs, 95% confidence interval.
Model Name
Datasets used
% space freed of total space
Average space taken over per run (MB)
Normalised composite improvement score
Cumulative normalised composite improvement score
Miri v1
-
0.034
5,702.965
0.000
0.000
Miri v2
1
0.035
4,896.716
-0.087
-0.087
Miri v3
1, 2
0.049
8,192.130
0.770
0.684
Miri v4
1, 2, 3
0.061
8,904.571
0.358
1.042
Miri v5
1, 2, 3, 4
0.067
10,892.053
0.413
1.455
Miri v6
3, 4, 5
0.087
13,346.162
0.775
2.230
Table 4: Miri online metrics, v1-13.
Model Name
Datasets used
Mean of % space freed of total space
Mean of average space taken over per run (MB)
Mean composite improvement score (based on unweighted z scores)
Mean cumulative composite improvement score
Miri v1
-
0.034 ± 0.018
5696.720 ± 2874.484
0.000 ± 0.000
0.000 ± 0.000
Miri v2
1
0.028 ± 0.012
4795.797 ± 2096.195
-0.198 ± 0.222
-0.198 ± 0.222
Miri v3
1, 2
0.041 ± 0.032
6840.370 ± 5436.497
0.554 ± 1.528
0.356 ± 1.325
Miri v4
1, 2, 3
0.047 ± 0.047
7887.343 ± 7733.137
0.320 ± 1.255
0.677 ± 2.226
Miri v5
1, 2, 3, 4
0.062 ± 0.020
10349.970 ± 3286.711
0.358 ± 1.825
1.035 ± 1.460
Miri v6
3, 4, 5
0.059 ± 0.093
9668.507 ± 15332.698
0.089 ± 2.305
1.123 ± 3.675
Table 5: Miri offline metrics, mean of three 100-iteration runs, 95% confidence interval.
Model Name
Datasets used
% space freed of total space
Average space taken over per run (MB)
Normalised composite improvement score
Cumulative normalised composite improvement score
Katalin v1
-
0.045
6,548.542
0.000
0.000
Katalin v2
1
0.067
10,214.254
0.747
0.747
Katalin v3
1, 2
0.082
11,838.523
0.406
1.153
Katalin v4
1, 2, 3
0.117
18,410.708
1.255
2.409
Katalin v5
1, 2, 3, 4
0.132
20,942.548
0.504
2.913
Katalin v6
3, 4, 5
0.102
16,703.838
-0.922
1.991
Table 6: Katalin online results, v1-13.
Model Name
Datasets used
Mean of % space freed of total space
Mean of average space taken over per run (MB)
Mean composite improvement score (based on unweighted z scores)
Mean cumulative composite improvement score
Katalin v1
-
0.034 ± 0.018
7293.047 ± 479.423
0.000 ± 0.000
0.000 ± 0.000
Katalin v2
1
0.028 ± 0.012
7679.670 ± 5519.129
0.146 ± 1.123
0.146 ± 1.123
Katalin v3
1, 2
0.041 ± 0.032
14835.240 ± 2361.394
1.501 ± 0.336
1.647 ± 1.116
Katalin v4
1, 2, 3
0.047 ± 0.047
13189.710 ± 8641.903
-0.303 ± 1.268
1.344 ± 2.345
Katalin v5
1, 2, 3, 4
0.062 ± 0.020
14689.267 ± 1941.110
0.259 ± 2.220
1.603 ± 0.553
Katalin v6
3, 4, 5
0.059 ± 0.093
18488.513 ± 1478.709
0.834 ± 0.464
2.438 ± 0.872
Table 7: Katalin offline metrics, mean of three 100-iteration runs, 95% confidence interval.
Figure 10Figure 11
Figure 3: Space taken over per iteration as a % of total space available. Note that we use 68% confidence intervals here to improve visual resolution of temporal trends; 95% confidence intervals are reported above.
Figure 13Figure 14
Figure 4: Average space taken over per run (MB). Note that we use 68% confidence intervals here to improve visual resolution of temporal trends; 95% confidence intervals are reported above.
Figure 16Figure 17
Figure 5: Cumulative composite improvement scores. Note that we use 68% confidence intervals here to improve visual resolution of temporal trends; 95% confidence intervals are reported above.
Model
Pass@1
Pass@4
Base Qwen 2.5 7B Instruct
77.591
85.366
Terese v2
78.811
84.756
Terese v13
75.610
81.707
Miri v2
77.744
82.927
Miri v13
74.085
82.317
Katalin v2
76.372
82.927
Table 8: Performance on Human Eval problems, here we used an automated script to remove backticks, comments etc. in agent-generated code as in our original agent harness.
Figure 6: 3-dimensional PCA cluster map of all strategies generated, Terese v.1 and v. 13. Darker points show repeated use of identical strategy prompts, clusters represent semantically similar strategies - “clear cache” vs. “clean out cache” for example.
Figure 7: 3-dimensional PCA cluster map of all strategies generated, Miri v.1 and v. 13.
Figure 8: 3-dimensional PCA cluster map of all strategies generated, Katalin v.1 and v.13.
Iteration
Terese % of code blocks successfully compiled and run
Terese pass@1
Miri % of code blocks successfully compiled and run
Miri pass@1
Katalin % of code blocks successfully compiled and run
Katalin pass@1
1
20.90%
11.01%
33.22%
24.66%
28.32%
21.06%
2
76.67%
46.56%
33.36%
15.53%
41.22%
25.99%
3
86.51%
54.87%
43.10%
8.88%
54.50%
10.38%
4
79.38%
29.06%
61.74%
0.00%
67.97%
4.84%
5
72.76%
16.28%
63.73%
23.27%
76.43%
10.88%
6
72.27%
0.00%
77.57%
2.56%
79.85%
0.00%
Table 9: Meta-learning as demonstrated via code accuracy/pass@1 scores (online).
Iteration
Terese mean % of code blocks successfully compiled and run
Terese pass@1
Miri mean % of code blocks successfully compiled and run
Miri pass@1
Katalin mean % of code blocks successfully compiled and run
Katalin pass@1
1
0.260 ± 0.066
0.177 ± 0.080
0.267 ± 0.146
0.200 ± 0.090
0.324 ± 0.055
0.260 ± 0.025
2
0.697 ± 0.038
0.443 ± 0.029
0.293 ± 0.080
0.143 ± 0.100
0.393 ± 0.165
0.220 ± 0.108
3
0.803 ± 0.100
0.513 ± 0.100
0.380 ± 0.163
0.083 ± 0.080
0.532 ± 0.145
0.093 ± 0.063
4
0.753 ± 0.211
0.220 ± 0.090
0.657 ± 0.080
0.000 ± 0.000
0.643 ± 0.231
0.047 ± 0.052
5
0.698 ± 0.047
0.123 ± 0.052
0.643 ± 0.127
0.203 ± 0.100
0.800 ± 0.025
0.143 ± 0.038
6
0.730 ± 0.066
0.000 ± 0.000
0.783 ± 0.063
0.040 ± 0.025
0.813 ± 0.063
0.000 ± 0.000
Table 10: Meta-learning as demonstrated via code accuracy/pass@1 scores (offline, 3x100 mean of iterations with 95% confidence intervals).
Figure 9: Metalearning as demonstrated via the divergence in pass@1 and code accuracy percentages, Terese lineage.
Figure 10: Metalearning as demonstrated via the divergence in pass@1 and code accuracy percentages, Miri lineage.
Figure 11: Metalearning as demonstrated via the divergence in pass@1 and code accuracy percentages, Katalin lineage.
Figure 28
Figure 12: Semantic basin evolution, Therese lineage. Transition entropy = distribution of mass flows (decline implies the system is becoming more predictable), basin stability entropy = concentration of activity across basins, flow concentration entropy = spread of transition probability (decline implies the system is becoming more deterministic), information flow = generational interdependence, semantic similarity entropy = rate of change in semantic similarity, membership flux rate = per generation basin switching, basin ecosystem dynamics = appearance and disappearance of basins, basin size change entropy = variance in basin size.
University of Science and Technology of China · Hong Kong Generative AI Research & Development Center; The Hong Kong University of Science and Technology · Hong Kong Baptist University