Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each accepted skill only for its originating family raises mean hidden trajectory utility from 0.713 under global memory to 0.816 and changes harmful deployments from six of eight to none. In 27 paired randomized-order streams, Scoped-ORC improves mean trajectory utility by 0.063 [0.037, 0.094] over Global-ORC, accepts 63 rather than 12 updates, and produces multiple accepted updates in 19/27 streams, with 0/63 harmful acceptances. The global control reaches 0.713, below the static agent's 0.775, because locally valid edits can interfere with unrelated families. These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.
Figures & tables
Method
Mean traj. ↑
Final checkpoint ↑
BWT ↑
Accepted
Harmful
Static
0.775 [0.734, 0.814]
0.775 [0.734, 0.814]
0.000 [0.000, 0.000]
0.00
0.00
Frozen-compute
0.775 [0.734, 0.814]
0.775 [0.734, 0.814]
0.000 [0.000, 0.000]
0.00
0.00
Latest-only
0.703 [0.678, 0.725]
0.720 [0.681, 0.755]
-0.253 [-0.300, -0.204]
12.00
6.25
Self-judge
0.736 [0.696, 0.774]
0.722 [0.641, 0.802]
-0.140 [-0.220, -0.059]
3.38
1.75
Replay
0.707 [0.676, 0.738]
0.681 [0.621, 0.745]
-0.240 [-0.297, -0.191]
7.88
3.88
ORC
0.713 [0.653, 0.761]
0.707 [0.645, 0.760]
0.000 [0.000, 0.000]
1.00
0.75
Table 1: Main ProcStream-RSI results over eight paired streams. Brackets are seed-bootstrap 95% intervals. Harmful accepts count accepted rounds whose fixed hidden checkpoint score decreased.
Contrast
Mean difference [bootstrap 95%]
p
Holm p
ORC - Replay mean trajectory
0.006 [-0.045, 0.048]
0.8438
1.0000
ORC - Latest-only mean trajectory
0.010 [-0.036, 0.051]
0.7188
1.0000
ORC - Batch-ORC final
-0.068 [-0.123, -0.004]
0.0938
0.2812
Table 2: Prespecified paired contrasts. Intervals are descriptive percentile seed-bootstrap 95% intervals. Paired sign-flip sensitivity p -values require symmetry of seed-level differences; Holm adjustment covers the three rows.
Method
HumanEval+ pass rate
GPT-OSS hidden pass rate
Static
0.914 [0.898, 0.930]
0.941 [0.907, 0.974]
Latest-only
0.918 [0.891, 0.941]
0.911 [0.886, 0.937]
Replay
0.922 [0.895, 0.945]
0.916 [0.878, 0.951]
ORC
0.895 [0.867, 0.922]
0.951 [0.926, 0.968]
Batch-ORC
0.914 [0.898, 0.930]
0.941 [0.907, 0.974]
Scoped retrieval †
—
0.960 [0.930, 0.987]
Table 3: Terminal-policy transfer. Values are means with seed-bootstrap 95% intervals. HumanEval+ uses the fixed 32-task sample; GPT-OSS uses each stream’s 18-task sealed final bank. † Family-scoped retrieval over the same archived completions.
Deployment
Mean traj. ↑
Final checkpoint ↑
Accepted
Harmful
Static
0.775 [0.734, 0.814]
0.775 [0.734, 0.814]
0
0
Global ORC
0.713 [0.653, 0.762]
0.707 [0.645, 0.760]
8
6
Scoped retrieval
0.816 [0.782, 0.852]
0.819 [0.786, 0.855]
8
0
Table 4: Same accepted skills, different retrieval scope on the eight main streams. Scoped retrieval applies each ORC rule only to its originating family and uses the initial skill elsewhere.
Deployment
Accepted
Harmful
Checkpoint t=1
Non-current Δ
Global ORC
5
3
0.761 [0.709, 0.802]
-0.047 [-0.120, 0.005]
Scoped ORC
5
0
0.802 [0.779, 0.826]
0.000 [0.000, 0.000]
Table 5: Balanced randomized-entry replication (18 streams; two per first family). Brackets show stream-bootstrap 95% intervals.
Method
Trajectory
Final checkpoint
Accepts/stream
Streams ≥2
Harmful/accepted
Global- ORC
0.785 [0.757, 0.812]
0.789 [0.758, 0.817]
0.44
2/27
6/12
Scoped- ORC
0.848 [0.825, 0.871]
0.900 [0.871, 0.927]
2.33
19/27
0/63
Table 6: Full 12-round randomized-order extension (27 paired streams). Brackets are stream-bootstrap 95% intervals; harmful updates reduce the next hidden checkpoint.