Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter--context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.
Figures & tables
Figure 1: The context–parameter feedback loop in COEVO.
Figure 2: Policy-state signals evolve throughout training. (a) Relative attention allocation across functional prompt sections changes substantially over training, indicating that the policy’s utilization of the same context is non-stationary. (b) Policy entropy captures the concurrent evolution of behavioral exploration: under a fixed detailed context, entropy progressively decreases, whereas COEVO maintains broader exploration through context adaptation. Together, these dynamics motivate using both context utilization and policy exploration to guide context evolution.
Figure 3: Policy-aware context intervention. Entropy measures behavioral exploration, while attention reflects policy-relative context utilization. Their combination determines how the strategy context is adapted to the current policy.
Exploration
Utilization
Intervention
Context update
Low
Low
Expand
The policy is behaviorally concentrated while making limited use of the existing strategy guidance. Introduce an alternative strategy or scaffold to enlarge the available behavioral space.
Low
High
Relax
The policy is concentrated while strongly relying on the current guidance. Soften overly prescriptive instructions allowing alternative reasoning paths when appropriate.
High
Low
Clarify
The policy explores broadly but makes weak use of the relevant guidance. Rewrite the selected section so that its intended strategy, priority, or activation condition is more explicit.
High
High
Consolidate
The policy remains highly exploratory despite strong contextual utilization. Merge redundant or competing guidance into a more coherent strategy scaffold to reduce fragmentation.
Table 1: Policy-aware context intervention. Exploration is measured by policy entropy, while utilization is estimated from attention to the relevant strategy section. Both signals are interpreted relative to the current policy’s calibrated reference range.
Metric
Training
Simple
Medium
Detail
Mean ↑
Worst ↑
Δ↓
Average@12
Simple
44.72
40.28
40.28
41.76
40.28
4.44
Medium
42.78
41.94
40.28
41.67
40.28
2.50
Detail
39.17
38.89
39.72
39.26
38.89
0.83
COEVO
43.06
45.28
44.17
44.17
43.06
2.22
Pass@12
Simple
66.67
70.00
66.67
67.78
66.67
3.33
Medium
63.33
70.00
66.67
66.67
63.33
6.67
Table 2: Cross-prompt robustness on 4B-Math. Models trained with different contexts are evaluated under the same Simple, Medium, and Detail inference prompts. Worst denotes the minimum performance across inference prompts, and Δ denotes the max–min gap; higher Worst and lower Δ indicate stronger robustness to inference-prompt variation.
Step
State (Ent./Util.)
Context revision
4
L / L-L-L
Generic tool advice → explicit strategy selection
29
L / L-L-L
Add problem categories and a cost–benefit decision rule
49
H / L-L-L
Strengthen guidance into an explicit decision procedure
54
L / H-H-H
“must invoke” → “consider invoking”
69
L / L-L-H
Reintroduce classification and clearer tool-use conditions
99
H / H-H-L
Consolidate guidance into a concise three-step procedure
Table 3: Representative accepted context revisions in Case I.
Large language model agents increasingly operate in dynamic environments where tool interfaces, APIs, and user requirements change after deployment. Existing self-evolution methods mainly follow two paradigms: harness-based approaches, which externalize feedback into editable memories or skills for rapid adaptation, and parameter-based approaches, which internalize experience into model parameters for deeper capability improvement. However, using either mechanism alone creates a trade-off between flexibility and performance. This paper asks how an agent can coordinate both channels to achieve robust self-evolution. We present COVE, a unified agent self-evolution framework that combines harness-based and parameter-based learning through task-aware routing, stage-aware scheduling, and knowledge optimization. Through this design, COVE treats self-evolution not as indiscriminate accumulation of experience, but as a coordinated process that matches tasks and knowledge types to appropriate learning mechanisms. Experiments across multiple task categories show that COVE outperforms single-channel evolution strategies, demonstrating more robust and efficient improvement under changing environments.
Tianyun Ji, Zhenya Huang, Jiayu Liu +3
University of Science and Technology of China Hefei, Anhui, China · City University of Hong Kong Hong Kong, China · Hefei Normal University Hefei, Anhui, China +1
Self-evolving large language models (LLMs) learn by generating their own training tasks and solutions, reducing reliance on human-curated supervision. However, in many reasoning domains, the model must also validate generated tasks and judge generated answers to obtain training signals. This creates a training-signal challenge: erroneous self-judgments become erroneous gradient updates. Existing approaches either rely on external verifiers, which limits generality, or treat noisy self-generated feedback as supervision. We propose COSE (Confidence-Orchestrated Self-Evolution), which uses the LLM's intrinsic confidence as a lightweight uncertainty signal to modulate learning. COSE introduces confidence-weighted PPO updates and confidence-prioritized replay. Across 19 held-out benchmarks and four Qwen/Llama backbones (0.6B--4B), COSE consistently improves over base models and achieves the best average performance in general reasoning and mathematics, while remaining competitive on code. Code and data are available at https://anonymous.4open.science/r/COSE_-B5C2.
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training-time optimization methods can update model parameters but lack an explicit mechanism for accumulating transferable experience. Bridging these two paradigms requires a critical intermediate stage that remains underexplored, namely \emph{experience distillation}. To address this gap, we propose \textbf{SPEE} (\textbf{S}elf-\textbf{P}rogressive \textbf{E}xperience \textbf{E}volution), a unified post-training framework that sequentially performs explicit experience evolution followed by implicit policy optimization. During explicit experience evolution, SPEE reflects on trajectories collected from multiple interactions to extract, verify, and progressively evolve transferable experience, which is subsequently internalized into the policy through privilege-guided On-Policy Self-Distillation (OPSD). During implicit policy optimization, reward-driven reinforcement learning leverages these internalized priors to explore novel solution strategies. In the experience evolution stage, a continuously evolving global experience pool consolidates knowledge from both successful and failed trajectories, filters out low-utility experience, and mitigates post-hoc rationalization induced by individual trajectories. Experiments on five mathematical reasoning benchmarks demonstrate that SPEE consistently outperforms both test-time and training-time self-evolution baselines across three model scales. The source code is available at https://github.com/rrrsj/SPEE.
Shijie Ren, Xiting Wang, Meng Li +8
Gaoling School of Artificial Intelligence, Renmin University of China · Weixin AI, Tencent Inc, China