COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement
Organizations: Peking University · Emotional Machine
Abstract
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter--context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.
Figures & tables
| Exploration | Utilization | Intervention | Context update |
|---|---|---|---|
| Low | Low | Expand | The policy is behaviorally concentrated while making limited use of the existing strategy guidance. Introduce an alternative strategy or scaffold to enlarge the available behavioral space. |
| Low | High | Relax | The policy is concentrated while strongly relying on the current guidance. Soften overly prescriptive instructions allowing alternative reasoning paths when appropriate. |
| High | Low | Clarify | The policy explores broadly but makes weak use of the relevant guidance. Rewrite the selected section so that its intended strategy, priority, or activation condition is more explicit. |
| High | High | Consolidate | The policy remains highly exploratory despite strong contextual utilization. Merge redundant or competing guidance into a more coherent strategy scaffold to reduce fragmentation. |
| Metric | Training | Simple | Medium | Detail | Mean | Worst | |
|---|---|---|---|---|---|---|---|
| Average@12 | Simple | 44.72 | 40.28 | 40.28 | 41.76 | 40.28 | 4.44 |
| Medium | 42.78 | 41.94 | 40.28 | 41.67 | 40.28 | 2.50 | |
| Detail | 39.17 | 38.89 | 39.72 | 39.26 | 38.89 | 0.83 | |
| COEVO | 43.06 | 45.28 | 44.17 | 44.17 | 43.06 | 2.22 | |
| Pass@12 | Simple | 66.67 | 70.00 | 66.67 | 67.78 | 66.67 | 3.33 |
| Medium | 63.33 | 70.00 | 66.67 | 66.67 | 63.33 | 6.67 |
| Step | State (Ent./Util.) | Context revision |
|---|---|---|
| 4 | L / L-L-L | Generic tool advice explicit strategy selection |
| 29 | L / L-L-L | Add problem categories and a cost–benefit decision rule |
| 49 | H / L-L-L | Strengthen guidance into an explicit decision procedure |
| 54 | L / H-H-H | “must invoke” “consider invoking” |
| 69 | L / L-L-H | Reintroduce classification and clearer tool-use conditions |
| 99 | H / H-H-L | Consolidate guidance into a concise three-step procedure |
| Case | Observed evolution | Implication |
|---|---|---|
| Case I | Structure stronger constraints relaxation consolidation | Appropriate guidance changes with the current policy state. |
| Case II | Detailed operational guidance more direct strategy instructions | Redundant guidance can be removed while preserving core task requirements. |
| Method | Avg. | Pass | |
|---|---|---|---|
| Attn.-only | 40.2 | 66.7 | 0.35 |
| Ent.-only | 42.0 | 70.0 | 0.47 |
| E-SPL | 41.6 | 66.7 | 0.31 |
| COEVO | 44.3 | 73.3 | 0.52 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Domain | Training data | Evaluation data |
|---|---|---|---|
| Qwen 3.5-4B | Math | DAPO-Math-17k | AIME 2025 |
| Qwen 3.5-4B-Base | Math | DAPO-Math-17k | AIME 2025 |
| Qwen 3.5-9B | Math | DAPO-Math-17k | AIME 2025 |
| Qwen 3.8-27B | Math | DAPO-Math-17k | AIME 2025 |
| Qwen 3.5-4B | Code | LCB-v6 temporal train | 100-task stratified subset of frozen 200-task test split |
| Qwen 3.5-9B | Code | LCB-v6 temporal train | 100-task stratified subset of frozen 200-task test split |
| Setting | Math | Code |
|---|---|---|
| Initialization | Qwen 3.5-4B, Qwen 3.5-4B-Base, Qwen 3.5-9B, and Qwen 3.8-27B | Qwen 3.5-4B and Qwen 3.5-9B initialized from a shared SFT checkpoint |
| RL algorithm | GRPO-style group-mean advantage without standard-deviation normalization; PPO clip | GRPO-style group-mean advantage without standard-deviation normalization; PPO clip |
| Optimizer | Adam, , | Adam, , |
| Learning rate | ||
| Batch size (task per step) | 8 (4B), 4 (9B, 27B) | 4 |
| Rollouts per task | 8 | 8 |
| Protocol | Step | Avg.@12 | Pass@12 | Format | Code calls |
|---|---|---|---|---|---|
| Simple | 0 | 2.50 | 16.67 | 2.78 | 0.03 |
| 50 | 12.78 | 30.00 | 22.78 | 0.01 | |
| 100 | 10.83 | 20.00 | 33.89 | 0.10 | |
| Medium | 0 | 1.39 | 6.67 | 1.67 | 0.01 |
| 50 | 11.11 | 23.33 | 17.50 | 0.00 | |
| 100 | 15.00 | 30.00 | 39.17 | 0.00 |
| Protocol | Step | Avg.@12 | Pass@12 |
|---|---|---|---|
| Simple | 0 | 31.94 | 53.33 |
| 25 | 43.61 | 73.33 | |
| 50 | 50.28 | 83.33 | |
| 75 | 51.67 | 90.00 | |
| 100 | 45.28 | 76.67 | |
| Medium | 0 | 49.44 | 76.67 |
| Train | Step | S | M | D | Mean | Worst | Case |
|---|---|---|---|---|---|---|---|
| Simple | 20 | 53.00 | 53.00 | 59.00 | 55.00 | 53.00 | 62.41 |
| Simple | 40 | 57.00 | 57.00 | 56.00 | 56.67 | 56.00 | 64.05 |
| Simple | 60 | 48.00 | 52.00 | 55.00 | 51.67 | 48.00 | 61.46 |
| Simple | 80 | 60.00 | 55.00 | 55.00 | 56.67 | 55.00 | 65.95 |
| Simple | 100 | 56.00 | 58.00 | 54.00 | 56.00 | 54.00 | 63.78 |
| Medium | 20 | 48.00 | 49.00 | 48.00 | 48.33 | 48.00 | 57.61 |
| Train | Step | S | M | D | Mean | Worst | Case |
|---|---|---|---|---|---|---|---|
| Simple | 20 | 56.00 | 49.00 | 54.00 | 53.00 | 49.00 | 61.47 |
| Simple | 40 | 59.00 | 61.00 | 56.00 | 58.67 | 56.00 | 64.37 |
| Simple | 60 | 57.00 | 60.00 | 60.00 | 59.00 | 57.00 | 65.53 |
| Simple | 80 | 55.00 | 61.00 | 56.00 | 57.33 | 55.00 | 61.21 |
| Simple | 100 | 60.00 | 60.00 | 58.00 | 59.33 | 58.00 | 64.27 |
| Medium | 20 | 59.00 | 59.00 | 56.00 | 58.00 | 56.00 | 64.04 |
| Model | Method | Step | Avg.@12 | Pass@12 |
|---|---|---|---|---|
| Qwen3.5-4B | Base | 0 | 17.17 | 32.67 |
| Simple | 40 | 22.22 | 36.00 | |
| Medium | 100 | 28.56 | 41.00 | |
| Detail | 100 | 26.17 | 41.00 | |
| COEVO | 100 | 30.40 | 43.33 | |
| Qwen3.5-9B | Base | 0 | 26.28 | 42.33 |
| Domain | Setting | Method | Avg.@12 | Pass@12 |
| Math | 4B | Simple | 40.28 | 66.67 |
| Medium | 40.55 | 66.67 | ||
| Detail | 39.72 | 60.00 | ||
| COEVO | 44.17 | 73.33 | ||
| Math | 4B-Base | Simple | 10.83 | 20.00 |
| Medium | 15.00 | 30.00 |