Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret
Organizations: Independent Researcher · Department of Computer Science and Technology Kean University, Union, NJ, USA
Abstract
Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. Steps stay cheap, but anything the writer drops is lost before later decisions reveal that they need it. We quantify this loss and ask whether training can reduce it. Comparing the written state with the best state of the same size written in hindsight, we split the reader's loss into a budget loss, which any state of that size must incur, and a write-time regret, which comes from the writer's choices. In TextWorld cooking games where we control how long a fact must be carried before it is needed, a 128-token state holding the facts wins nearly every game, while prompted language-model writers win at most 17%. Almost all of the loss is write-time regret, and it grows with the delay. We then train the writer from the reader's own loss. DSSR (decision-sufficient state representations) scores candidate states by how well the reader acts after the writer carries them forward, and teaches the writer to prefer the better ones. This forward-rolled score predicts game outcomes (), whereas scoring a candidate as a fixed context, as hindsight methods usually do, does not (). On a pre-registered test split opened once, training adds +7.0 [+1.9, +12.2] points of success when facts are needed soon, bringing a plain summary writer to the level of belief- and slot-based memory prompts. The gain shrinks as the delay grows and is significant only at the shortest delay. We trace this limit to credit assignment: keeping a fact now pays off only if every later rewrite keeps it too, which a per-step score cannot see.
Figures & tables
| success, | , | success, | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 | |
| full | 97.8 | 97.8 | 97.8 | 97.8 | – | – | – | – | – | – | – | – |
| oracle-b | 100 | 100 | 100 | 97.8 | 100 | 100 | 100 | 97.8 | ||||
| lastk | 0 | 0 | 0 | 0 | 0.71 | 1.80 | 1.46 | 1.46 | 2.2 | 0 | 0 | 0 |
| summary | 2.2 | 2.2 | 0 | 2.2 | 0.85 | 1.26 | 1.47 | 1.10 | 20.0 | 6.7 | 2.2 | 2.2 |
| slots | 11.1 | 0 | 0 | 0 | 0.49 | 1.55 | 1.38 | 1.69 | 15.6 | 4.4 | 2.2 | 0 |
| success (%) | (nats) | |||||||
| 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 | |
| full (no budget) | 95.6 | 96.7 | 96.7 | 94.4 | – | – | – | – |
| oracle-b (privileged) | 98.9 | 100 | 98.9 | 97.8 | – | – | – | – |
| lastk | 0 | 0 | 0 | 0 | – | – | – | – |
| summary | 3.3 | 2.2 | 1.1 | 0 | 0.79 | 1.05 | 1.35 | 1.66 |
| belief | 11.1 | 0 | 0 | 0 | 0.57 | 1.21 | 1.55 | 1.73 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| rooms | Phase-A steps | Phase-B steps | history tokens (mean) | history tokens (max) | |
|---|---|---|---|---|---|
| 6 | 0 | 15.7 | 12.0 | 2230 | 2821 |
| 6 | 40 | 55.7 | 12.0 | 6843 | 9087 |
| 9 | 0 | 25.0 | 18.9 | 4151 | 5514 |
| 9 | 40 | 64.8 | 18.7 | 9102 | 10908 |
| 12 | 0 | 33.5 | 21.1 | 4825 | 6721 |
| 12 | 40 | 73.3 | 20.8 | 9079 | 11591 |
| gate | criterion | result | verdict |
| G1a | grows with lag: , interval excluding 0, for the write-time formats | 1/3 formats on the primary GPU (plateau after ten rewrites), 3/3 on the replicate; in the measured-lag form [0.24, 1.97] | fail / pass ( form) |
| G1b | the mean gap predicts online success across cells (Spearman) | over 76 cells ( replicate) | pass |
| G1c | and on Phase-A write points | window form: , fail; recursive form: [0.39, 0.56] under the original protocol, [0.13, 0.43] under (91/200 write points without return variance) | fail (window) |
| G2a | at , : DSSR best prompted ptsand outcome pts(intervals ); not below pts | [-0.4, +3.3] over summary ; [+0.2, +3.0] over outcome ; : [+1.9, +12.2] | fail |
| G2b | at reduced by (interval ); smaller -vs- slope | as first computed: ( ), difference ; slopes per step (seed 0), difference . Corrected (A5): ( ), ; slopes per step (three seeds), difference | fail |
| G2c | beats at (interval ); the gap shrinks at | [-0.4, +3.3] at lag; [-3.3, +10.4] at | fail |
| format | quantity | ||||
|---|---|---|---|---|---|
| full | success | 97.8 | 97.8 | 97.8 | 97.8 |
| oracle-b | success, | 100 | 100 | 100 | 97.8 |
| oracle-b | , | [-0.30, -0.08] | [-0.29, -0.07] | [-0.29, -0.13] | [-0.22, -0.09] |
| lastk | success, | 0 | 0 | 0 | 0 |
| success, | 2.2 | 0 | 0 | 0 | |
| , | 0.71 [0.54, 0.92] | 1.80 [1.39, 2.22] | 1.46 [1.10, 1.83] | 1.46 [1.03, 1.92] |
| 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 | |
|---|---|---|---|---|---|---|---|---|
| oracle-b | ||||||||
| lastk | 0.63 | 1.52 | 1.48 | 1.32 | 0.77 | 1.95 | 1.65 | 1.36 |
| summary | 0.93 | 1.42 | 1.29 | 1.56 | 0.42 | 1.31 | 1.50 | 1.50 |
| slots | 0.65 | 1.72 | 1.60 | 1.53 | 0.39 | 1.92 | 1.47 | 1.82 |
| belief | 0.70 | 1.24 | 1.11 | 1.45 | 0.50 | 1.15 | 1.05 | 0.95 |
| bin | (thr. 0.3) | (thr. 1.0) | ||||
|---|---|---|---|---|---|---|
| 1–4 | 107 | 0.61 [0.38, 0.84] | 0.30 [-0.00, 0.59] | 0.10 | 0.60 | 0.56 |
| 5–9 | 15 | 1.89 [0.33, 3.78] | [-3.76, 0.32] | 2.58 | 1.35 | 2.26 |
| 10–19 | 123 | 1.97 [1.41, 2.62] | 0.01 [-0.27, 0.24] | 0.37 | 1.89 | 2.21 |
| 20–39 | 146 | 1.87 [1.41, 2.36] | 0.29 [-0.01, 0.57] | 0.32 | 1.87 | 2.12 |
| 40+ | 141 | 2.30 [1.75, 2.92] | [-0.67, 0.15] | 0.68 | 2.18 | 2.50 |
| condition | mean |
|---|---|
| baseline, | 0.69 |
| baseline, (walk before the reveal) | 0.45 |
| brief, | 0.92 |
| post, 40 steps after the reveal | 1.76 |
| post and brief | 1.13 |
| contrast | estimate [lo, hi] |
| writer | round | 10 | 20 | 40 | mean | |
|---|---|---|---|---|---|---|
| DSSR ( ) | 0 | 13.3 | 0 | 2.2 | 0 | 3.9 |
| 1 | 11.1 | 4.4 | 2.2 | 2.2 | 5.0 | |
| 2 | 4.4 | 8.9 | 4.4 | 0 | 4.4 | |
| 3 | 13.3 | 0 | 0 | 2.2 | 3.9 | |
| 0 | 6.7 | 2.2 | 0 | 0 | 2.2 | |
| 1 | 4.4 | 2.2 | 0 | 0 | 1.7 |
| writer | round | 10 | 20 | 40 | mean | |
|---|---|---|---|---|---|---|
| DSSR ( ) | 0 | 0.71 | 1.24 | 1.17 | 1.70 | 1.20 |
| 1 | 0.52 | 1.60 | 1.59 | 1.87 | 1.40 | |
| 2 | 0.47 | 1.06 | 1.17 | 1.41 | 1.02 | |
| 3 | 0.79 | 1.11 | 1.02 | 1.14 | 1.01 | |
| 0 | 0.94 | 1.49 | 1.51 | 1.36 | 1.33 | |
| 1 | 0.80 | 1.30 | 1.44 | 1.54 | 1.27 |
| candidates | group | cov. all | cov. best | cov. worst | all zero | differ | best worst | best worst | |
|---|---|---|---|---|---|---|---|---|---|
| main, round 0 | all | 2818 | 0.24 | 0.23 | 0.24 | 0.51 | 0.09 | 0.28 | 0.29 |
| 439 | 0.44 | 0.44 | 0.43 | 0.19 | 0.20 | 0.24 | 0.33 | ||
| 632 | 0.30 | 0.30 | 0.31 | 0.41 | 0.10 | 0.29 | 0.33 | ||
| 760 | 0.20 | 0.19 | 0.20 | 0.58 | 0.07 | 0.30 | 0.26 | ||
| 987 | 0.13 | 0.13 | 0.13 | 0.67 | 0.05 | 0.31 | 0.18 | ||
| 1–4 steps after reveal | 331 | 0.42 | 0.42 | 0.41 | 0.27 | 0.25 | 0.28 | 0.27 |
| success | ||||||||
|---|---|---|---|---|---|---|---|---|
| seed | 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 |
| 0 (selected) | 26.7 | 11.1 | 4.4 | 2.2 | 0.61 | 1.13 | 1.19 | 1.46 |
| 1 | 8.9 | 11.1 | 2.2 | 2.2 | 0.51 | 1.26 | 0.91 | 1.26 |
| 2 | 22.2 | 8.9 | 8.9 | 4.4 | 0.31 | 1.34 | 1.14 | 1.55 |
| mean sd | 19.3 7.6 | 10.4 1.0 | 5.2 2.8 | 3.0 1.0 | 0.48 | 1.24 | 1.08 | 1.42 |
| untrained | 2.2 | 2.2 | 0 | 2.2 | 0.85 | 1.26 | 1.47 | 1.10 |
| reader | condition | 10 | 20 | 40 | , | |
|---|---|---|---|---|---|---|
| full | 95.6 [91.1, 98.9] | 96.7 [92.2, 100] | 96.7 [92.2, 100] | 94.4 [88.9, 98.9] | – | |
| oracle-b | 98.9 [96.7, 100] | 100 | 98.9 [96.7, 100] | 97.8 [94.4, 100] | – | |
| lastk | 0 | 0 | 0 | 0 | – | |
| summary | 3.3 [0.0, 7.8] | 2.2 [0.0, 5.6] | 1.1 [0.0, 3.3] | 0 | 1.50 [1.23, 1.80] | |
| belief | 11.1 [5.6, 17.8] | 0 | 0 | 0 | 1.64 [1.36, 1.93] | |
| slots | 10.0 [4.4, 16.7] | 0 | 0 | 0 | 1.80 [1.50, 2.12] |
| 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 | 0 | 10 | 20 | 40 | |
| summary | 0 | 0 | 0 | 0 | 7.8 | 0 | 0 | 1.1 | 7.8 | 1.1 | 2.2 | 1.1 |
| belief | 2.2 | 0 | 1.1 | 1.1 | 16.7 | 1.1 | 1.1 | 1.1 | 16.7 | 6.7 | 8.9 | 2.2 |
| slots | 0 | 0 | 0 | 0 | 12.2 | 0 | 0 | 0 | 20.0 | 2.2 | 0 | 0 |
| guide | 0 | 0 | 0 | 0 | 16.7 | 2.2 | 5.6 | 2.2 | 30.0 | 16.7 | 15.6 | 15.6 |
| DSSR (A4 r2, seed 0) | 2.2 | 0 | 0 | 0 | 14.4 | 4.4 | 4.4 | 1.1 | 18.9 | 14.4 | 2.2 | 5.6 |
| condition | 0 | 1–2 | 3–5 | 6–9 | 10–14 | 15–19 | 20–29 | 30–39 | 40–59 | 60+ |
|---|---|---|---|---|---|---|---|---|---|---|
| summary , valid | 0.02 | 1.03 | 1.24 | 1.36 | 1.28 | 1.73 | 1.54 | 1.14 | 0.94 | 0.54 |
| DSSR seed 0, valid | 0.05 | 0.39 | 0.66 | 0.27 | 1.07 | 1.59 | 1.25 | 0.59 | 1.29 | 1.13 |
| DSSR seed 1, valid | 0.01 | 0.21 | 0.98 | 0.71 | 0.82 | 1.21 | 1.31 | 0.91 | 1.20 | 0.38 |
| DSSR seed 2, valid | 0.07 | 0.21 | 0.61 | 0.25 | 0.95 | 1.37 | 1.13 | 0.94 | 1.49 | 1.03 |
| guide , valid | 0.02 | 0.82 | 0.59 | 0.85 | 0.73 | 1.15 | 0.71 | 0.90 | 1.07 | 0.66 |
| summary , test | 0.68 | 0.99 | 1.17 | 1.15 | 1.04 | 1.29 | 1.10 | 1.39 | 0.68 |
| condition | 0 | 1–2 | 3–5 | 6–9 | 10–14 | 15–19 | 20–29 | 30–39 | 40–59 | 60+ |
| summary , valid | 39 | 67 | 45 | 20 | 152 | 29 | 183 | 44 | 229 | 67 |
| DSSR seed 0, valid | 26 | 58 | 50 | 43 | 152 | 54 | 182 | 47 | 245 | 65 |
| DSSR seed 1, valid | 29 | 66 | 53 | 37 | 139 | 56 | 196 | 54 | 246 | 75 |
| DSSR seed 2, valid | 27 | 57 | 51 | 45 | 153 | 56 | 182 | 64 | 230 | 83 |
| guide , valid | 33 | 59 | 44 | 32 | 125 | 46 | 187 | 59 | 260 | 50 |
| summary , test | 69 | 132 | 130 | 69 | 247 | 85 | 364 | 107 | 493 | 151 |
| split | writer | bin | games | difference [lo, hi] | relative |
|---|---|---|---|---|---|
| valid | DSSR (3 seeds) | 1–2 | 39 | [+0.20, +1.65] | % |
| 3–5 | 30 | [+0.10, +1.25] | % | ||
| 6–9 | 11 | [-0.14, +1.63] | % | ||
| 10–14 | 44 | [-0.21, +0.73] | % | ||
| 15–19 | 15 | [-1.13, +1.65] | % | ||
| 20–29 | 43 | [-0.14, +0.69] | % |
| 0 | 2 | 3 | 5 | 7 | 9 | 10 | 15 | 20 | 30 | 40 | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DSSR, 3-seed mean (%) | 19.3 | 11.9 | 9.6 | 9.6 | 8.9 | 6.7 | 10.4 | 6.7 | 5.2 | 2.2 | 3.0 |
| summary (%) | 2.2 | 2.2 | 0 | 0 | 0 | 0 | 2.2 | 0 | 0 | 0 | 2.2 |
| gain (points) | |||||||||||
| interval, lower | |||||||||||
| interval, upper |