Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
Figures & tables
Figure 1 : CAMG-RL learns to use files as memory for long-horizon tasks. Each CAMG environment provides the task-acting policy with an episode-persistent workspace that it can write, revise, search, and read through ordinary filesystem actions throughout the episode. The illustrated context-boundary case shows the policy updating a bounded CONTINUATION.md before context replacement and retrieving the saved working state through an ordinary file read afterward.
Multiple domains
Task reward
End-to-end RL
Task across contexts
Shell access
Agent-memory benchmarks
LongMemEval [ 54 ]
×
×
×
×
×
MemoryAgentBench [ 14 ]
×
×
×
×
×
MemoryArena [ 10 ]
✓
×
×
✓
×
Trainable long-horizon agent environments
AgentGym-RL [ 57 ]
✓
✓
✓
×
×
Table 1 : CAMG uniquely enables end-to-end RL of filesystem memory across verifiable, cross-context tasks in multiple domains. Each environment produces its task reward by a fixed procedure from executed policy behavior: three use programmatic checks, and DeepResearch answers are scored by a frozen semantic judge shared by every evaluated method (Appendix A.1.3 ). We position CAMG among agent-memory benchmarks and trainable agent environments. ✓ = present, × = absent.
Figure 2 : A single policy learns to use files as memory while solving long-horizon tasks. Task interactions and filesystem operations form one policy trajectory over an episode-persistent workspace. At the highlighted context boundary, the policy updates CONTINUATION.md . A later read brings the saved working state back into active context. Each completed episode then carries its full trajectory record into fully asynchronous PPO, where the learner consumes queued episodes and publishes updated weights for later episodes.
Success rate, %
Method
Shop
Coding
Deep Research
Auto Research
Average
Frozen and training-free controls
Qwen3.5-4B
11.6
10.9
36.7
10.2
17.4
Mem0 [ 2 ]
15.6
14.8
37.5
10.9
19.7
Letta Code [ 37 ]
12.6
14.8
32.0
11.7
17.8
Learned-memory methods
Table 2 : CAMG-RL achieves the highest average success rate across the CAMG test set. We compare CAMG-RL with training-free memory controls and learned-memory baselines.
Figure 3 : The frozen base already prefers filesystem memory actions, and training converts this preference into complete memory chains while barely moving the weights. (a) Frozen-base action surprisal on paired store, retrieve, revise, and remove intents. Renaming the filesystem tool to an unseen name ( Aliased files ) keeps the surprisal within 0.013 nats per token of the native actions, and the hybrid control ( AgeMem tools in native form ) recovers only a fifth of the store gap. (b) Complete store–retrieve–use chains and executable operations during training. AgeMem never completes a chain. (c) Mean token-averaged KL of each trained policy from the frozen base, and relative weight change over the language-model weights.
Success rate, %
Method
SWE-bench Verified
MLE-bench Lite
Average
Qwen3.5-4B
7.6
0.0
3.8
Qwen3.5-9B
14.8
4.5
9.7
Qwen3.5-27B
17.4
4.5
11.0
Qwen3.5-35B-A3B
15.6
4.5
10.1
Qwen3.5-122B-A10B
22.0
9.1
15.6
Table 3 : CAMG-RL-4B, trained with file-based memory, matches the 35B-A3B member of its own family on both benchmarks, and CAMG-RL-9B is competitive with the frozen 122B-A10B. We compare both with frozen Qwen3.5 models from 4B to 397B-A17B on SWE-bench Verified and MLE-bench Lite. We report issue-resolution rate on SWE-bench, Any Medal rate on MLE-bench, and their unweighted mean as Average.
Figure 4 : Memory-action gradients and general memory files both drive CAMG-RL’s gains, and the evaluation-only restriction trades Shop for the other three environments. Grouped bars report the four environment success rates and their equal-weight average for full CAMG-RL, w/o memory-action gradients, retrained w/o general files, and evaluated w/o general files on the same test tasks.
Figure 5 : The trained policy carries the task state across the context replacement by writing and reading the continuation file. It records the objective, evidence, test status, and next step in .agent_memory/CONTINUATION.md before the replacement, reads the file back after it, and completes the task with hidden tests passing (return 1.00). The patch card is an excerpt. The red divider marks the replacement.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Cross-session dependency
Natural attribute chain
A purchased attribute determines the required attribute in a later category.
Latent preference
Earlier choices establish a preference omitted from later requests.
Recency override
A later preference update replaces an earlier value.
Distractor robustness
The relevant profile must be retrieved among stale and unrelated records.
Compositional recall
Two earlier relations must be composed to identify a later target.
Negative constraint
Earlier exclusions leave one admissible option in later sessions.
Appendix
Table 4: Each Shop family requires information to persist across a six-purchase episode. We list the CAMG-synthesized families. Paired variants keep later prompts and candidates fixed while changing the relevant earlier state.
Figure 6 : Long-horizon CAMG tasks run under a bounded context window. Each panel summarizes one task input and the interaction available to the policy.
Table 5: CAMG-RL uses shared training settings with per-environment action budgets.
Environment
Reward rt
Shop
rt=rt/7 when rt>0 , and rt=rt otherwise. Correct purchases give rt=1 in sessions 1–5 and rt=2 in session 6. Invalid actions and incorrect purchases give rt=−0.01
Coding
rT=1{hidden tests pass} and rt=0 for t<T . A context-boundary response whose checkpoint write fails to persist receives rt=−0.1
DeepResearch
rT=1{semantic judge accepts the answer} and rt=0 for t<T . Invalid actions, including failed checkpoint writes, receive rt=−0.02 , and the episode ends at the twelfth invalid action
AutoResearch
On a valid submission, rT=min(1,max(0,sT)+0.1) for direction-normalized private-grader score sT , where the +0.1 is a one-time first-submission milestone. Invalid submissions give rT=0 , exhausting the action budget without submitting gives rT=−0.01 , and exceeding the stated runtime limits gives rT=−1 . For t<T , rt=0 except two one-time milestones: +0.05 for the first successful code execution and +0.1 for the first recorded validation metric. Unparseable actions receive rt=−0.01
Appendix
Table 6: Task rewards are sparse and success-based across the four environments.
Figure 7 : All four environments improve during the 200-update joint run. Each panel shows one environment’s success rate on training rollouts with temperature-1 sampling and 64 episodes per update, and the bottom panel shows the equal-weight average. Lines are non-overlapping five-update bucket means, and each panel annotates the final bucket of the checkpoint evaluated in Table 2 .
Intent
Probe variant
Paired gap
Name
Remainder
Store
AgeMem tools
+0.205
−0.001
+0.207
AgeMem tools in native form
+0.166
+0.118
+0.047
shell documentation removed
+0.354
+0.120
+0.234
Retrieve
AgeMem tools
+0.398
+0.022
+0.376
AgeMem tools in native form
+0.484
+0.193
+0.291
shell documentation removed
+0.477
+0.177
+0.300
Appendix
Table 7: The paired gap is not about the tool name. Each paired gap is the difference in mean NLL per action token against the CAMG-RL version on the same probe items. The name-token and remainder contributions each weight their segment’s nats by the target’s own token count, so the two contributions sum to the paired gap. The third variant of each group deletes every shell-command documentation passage from the hybrid variant’s prompt.
Probe variant
Store
Retrieve
Revise
Remove
CAMG-RL files
0.252
0.031
0.031
0.031
Aliased files
0.207
0.009
0.009
0.009
AgeMem tools
0.053
0.531
3.092
3.327
AgeMem tools in native form
6.371
3.763
5.184
5.306
shell documentation removed
6.757
3.323
5.319
5.095
Appendix
Table 8: Deleting the documented competitor leaves the foreign names expensive. Each cell is the mean surprisal in nats per token of the tool-name tokens of that variant’s target action, located by byte offset through the tokenizer’s offset map. The first two rows call the documented filesystem tool under its native and nonce names. The last row deletes every shell-command documentation passage from the hybrid variant’s prompt, holding the appended memory-tool block, target actions, setup turns, and responses byte-identical, which leaves the six memory tools as the only documented functions in the coding environment.
Policy
Shop
Coding
DeepResearch
AutoResearch
Average
CAMG-RL
97.1
26.6
55.5
38.3
54.4
token-level estimation
94.0
18.8
47.8
39.9
50.1
Appendix
Table 9: Token-level estimation control. Success rates (%) on the 512 test tasks of Appendix C.1 . The control replaces response-level advantage estimation with per-token GAE and changes nothing else in training or evaluation.
Figure 8 : Saved memory content drives later success. Paired success-rate change from replacing each trajectory’s memory files at its first context replacement: blank (erased) and shuffled (transplanted from a length-matched same-environment task) versus intact content, in percentage points across all four environments.
Success rate, %
Method
Shop
Coding
Deep Research
Auto Research
Average
Chain
AgeMem, common training
16.7
26.6
28.1
29.7
25.3
0.0
AgeMem, original recipe
59.3
25.0
47.7
22.7
38.7
14.1
CAMG-RL
97.1
26.6
55.5
38.3
54.4
20.9
Appendix
Table 10 : The original AgeMem recipe does not close the gap. Success rate, %, for the common-training AgeMem model of Table 2 , the same six tools trained with the full original recipe, and CAMG-RL. Chain is the share of episodes that complete the store–retrieve–use chain in the final training quarter, and the common-training model never completes one.
Success rate, %
Interface
Shop
Coding
Deep Research
Auto Research
Average
Parse
Chain
CAMG-RL, shell file edits
97.1
26.6
55.5
38.3
54.4
0.61
20.9
CAMG-RL, apply_patch tool
90.6
20.3
52.3
31.3
48.6
11.7
15.6
Appendix
Table 11 : A single out-of-distribution file tool does not recover the shell interface. Success rate, %, for CAMG-RL and a variant trained under the same protocol with file modification routed through a single Codex-style apply_patch tool. Parse is the share of responses the action parser rejects, over the whole training run for the shell interface and after training for the apply_patch variant. Chain is the share of episodes that complete the store–retrieve–use chain in the final training quarter.
Method
Scale
Wall-clock (h)
CAMG-RL-4B
4B
19.6
CompactionRL
4B
22.4
AgeMem
4B
23.6
CAMG-RL-9B
9B
23.2
Appendix
Table 12: Every formal training run finishes within about one day. All four use the shared asynchronous-PPO configuration of Table 5 on the same six training GPUs.
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Samsung Research, Beijing, China