cs.AIApr 7, 2025

Generalising from Self-Produced Data: Model Training Beyond Human Constraints

Authors: Alfath Daryl Alhajir, Jennifer Dodgson, Joseph Lim, Truong Ma Phi, Julian Peh, Akira Rafhael Janson Pattirane, Lokesh Poovaragan

Abstract

Current large language models (LLMs) are constrained by human-derived training data and limited by a single level of abstraction that impedes definitive truth judgments. This paper introduces a novel framework in which AI models autonomously generate and validate new knowledge through direct interaction with their environment. Central to this approach is an unbounded, ungamable numeric reward - such as annexed disk space or follower count - that guides learning without requiring human benchmarks. AI agents iteratively generate strategies and executable code to maximize this metric, with successful outcomes forming the basis for self-retraining and incremental generalisation. To mitigate model collapse and the warm start problem, the framework emphasizes empirical validation over textual similarity and supports fine-tuning via GRPO. The system architecture employs modular agents for environment analysis, strategy generation, and code synthesis, enabling scalable experimentation. This work outlines a pathway toward self-improving AI systems capable of advancing beyond human-imposed constraints toward autonomous general intelligence.

Figures & tables

Explore similar work

CardsList
  1. G-Zero: Self-Play for Open-Ended Generation from Zero Data

    May 11, 2026Chengsong Huang, Haolin Liu, Tong Zheng +7Self-PlayRecursive Self-Improvement

  2. Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning

    Nov 20, 2025Peng Xia, Kaide Zeng, Jiaqi Liu +5Self-Evolving AgentsSelf-Evolution

  3. Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration

    Apr 20, 2026Qifan Zhang, Dongyang Ma, Tianqing Fang +5On-Policy Self-EvolutionSelf-Evolution