Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Organizations: Salesforce AI Research
Abstract
Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.
Figures & tables
| Pre-boundary function set | Boundaries | Tasks | Prefix steps | Compaction index | dU@3 [95% CI] |
|---|---|---|---|---|---|
| No write | 96 | 60 | 7.9 | 1.22 | [ , ] |
| Login-only write | 368 | 123 | 15.1 | 1.89 | [ , ] |
| Substantive write | 126 | 64 | 22.1 | 3.14 | [ , ] |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Table | Shipped SHA-256 | Python 3.10 fresh | Python 3.12 fresh |
|---|---|---|---|
| boundary_wide.csv | 8534faccb3a97dec | identical | identical |
| boundary_long.csv | 34a61f99aa9d7af7 | identical | identical |
| step_level.csv | eba9b5d8075c90ea | identical | identical |
| curves.csv | 0cd6cee8d4f5dfda | differs ( 6fba4773d74415b7 ) | identical |
| marginal.csv | 2d144932ff1bd365 | differs ( 8b2adbb1326b5d98 ) | identical |
| error_subtypes.csv | 6a2969efd8f15b91 | identical | identical |
| Function | Method | Function | Method |
|---|---|---|---|
| file_system.list_directory | GET | supervisor.show_account | GET |
| file_system.read | GET | supervisor.show_account_credentials | GET |
| file_system.read_file | GET | supervisor.show_account_profile | GET |
| file_system.show_file_metadata | GET | supervisor.show_account_usernames | GET |
| phone.get_contact | GET | supervisor.show_phone_contacts | GET |
| phone.set_access_token | POST | venmo.show_api_descriptions | GET |
| Function | Boundaries |
|---|---|
| spotify.login | 200 |
| phone.login | 146 |
| venmo.login | 143 |
| file_system.login | 99 |
| simple_note.login | 73 |
| gmail.login | 4 |
| Cell | dU | dE | dR | |
|---|---|---|---|---|
| Read-only | 1 | 0.135 [ 0.071, 0.205] | 0.053 [ 0.010, 0.101] | 0.086 [ 0.026, 0.148] |
| Read-only | 3 | 0.398 [ 0.266, 0.533] | 0.130 [ 0.061, 0.212] | 0.272 [ 0.160, 0.384] |
| Read-only | 5 | 0.681 [ 0.504, 0.859] | 0.179 [ 0.107, 0.266] | 0.516 [ 0.370, 0.665] |
| Any prior write | 1 | 0.130 [ 0.097, 0.163] | 0.119 [ 0.096, 0.144] | 0.021 [ 0.008, 0.048] |
| Any prior write | 3 | 0.335 [ 0.269, 0.405] | 0.229 [ 0.188, 0.270] | 0.140 [ 0.088, 0.194] |
| Any prior write | 5 | 0.476 [ 0.384, 0.570] | 0.299 [ 0.246, 0.353] | 0.231 [ 0.157, 0.305] |
| Cell | dU | dE | dR | |
|---|---|---|---|---|
| No substantive write | 1 | 0.153 [ 0.119, 0.188] | 0.110 [ 0.087, 0.134] | 0.049 [ 0.023, 0.075] |
| No substantive write | 3 | 0.393 [ 0.326, 0.464] | 0.224 [ 0.186, 0.264] | 0.197 [ 0.145, 0.250] |
| No substantive write | 5 | 0.566 [ 0.478, 0.658] | 0.291 [ 0.243, 0.343] | 0.316 [ 0.248, 0.385] |
| Substantive write | 1 | 0.048 [ 0.004, 0.099] | 0.104 [ 0.053, 0.160] | 0.034 [ 0.096, 0.030] |
| Substantive write | 3 | 0.172 [ 0.048, 0.291] | 0.171 [ 0.077, 0.266] | 0.031 [ 0.065, 0.123] |
| Substantive write | 5 | 0.301 [ 0.122, 0.472] | 0.241 [ 0.116, 0.366] | 0.136 [ 0.013, 0.274] |
| Channel | Original label | Post hoc label | |
|---|---|---|---|
| 1 | dU | 0.005 [ 0.078, 0.066] | 0.105 [ 0.167, 0.044] |
| 1 | dE | 0.067 [ 0.013, 0.117] | 0.006 [ 0.062, 0.053] |
| 1 | dR | 0.066 [ 0.135, 0.002] | 0.083 [ 0.150, 0.012] |
| 3 | dU | 0.062 [ 0.207, 0.079] | 0.221 [ 0.359, 0.092] |
| 3 | dE | 0.099 [ 0.003, 0.182] | 0.053 [ 0.153, 0.047] |
| 3 | dR | 0.132 [ 0.254, 0.010] | 0.166 [ 0.274, 0.066] |
| Policy | Avoided | Nested 95% CI | Task-only 95% CI | Mass advantage (nested) |
|---|---|---|---|---|
| Stump | 0.108 | [0.084, 0.239] (0.077) | [0.083, 0.238] (0.078) | [ , ] |
| Depth-2 tree | 0.195 | [0.112, 0.251] (0.069) | [0.096, 0.246] (0.075) | [ , ] |
| Logistic (2 features) | 0.206 | [0.146, 0.250] (0.052) | [0.156, 0.256] (0.050) | [ , ] |
| 17 summaries (boosted) | 0.247 | [0.223, 0.288] (0.032) | [0.217, 0.277] (0.030) | [ , ] |
| Policy | (share) | ( , ; neg.) | [task 95% CI] | |
|---|---|---|---|---|
| Stump | 54 (0.092) | 22.00 (25.00, ; 8) | 18.67 | [ , ] |
| Depth-2 tree | 93 (0.158) | 49.67 (54.27, ; 12) | 32.15 | [ , ] |
| Logistic (frozen) | 96 (0.163) | 38.20 (44.00, ; 11) | 33.18 | [ , ] |
| Boosted, 17 summaries | 118 (0.200) | 59.27 (64.73, ; 12) | 40.79 | [ , ] |
| Boosted, 144 features | 118 (0.200) | 74.07 (76.80, ; 8) | 40.79 | [ , ] |
| Logistic, 144 features | 118 (0.200) | 78.53 (82.87, ; 9) | 40.79 | [ , ] |
| Original label | Post hoc label | ||||||
|---|---|---|---|---|---|---|---|
| Channel | Stump | Depth-2 | Logistic | Stump | Depth-2 | Logistic | |
| 1 | dU | 0.546 | 0.577 | 0.467 | 0.546 | 0.572 | 0.539 |
| 1 | dE | 0.490 | 0.567 | 0.503 | 0.490 | 0.567 | 0.475 |
| 1 | dR | 0.572 | 0.602 | 0.507 | 0.572 | 0.602 | 0.515 |
| 3 | dU | 0.592 | 0.619 | 0.547 | 0.592 | 0.628 | 0.560 |
| 3 | dE | 0.468 | 0.466 | 0.489 | 0.468 | 0.466 | 0.438 |
| Policy | Allowed | Retained | Avoided | Mass |
|---|---|---|---|---|
| Stump | 0 | 0.000 | 1.000 | 1.000 |
| 59 | 0.100 | 0.904 | 0.885 | |
| 104 | 0.176 | 0.837 | 0.818 | |
| 144 | 0.244 | 0.802 | 0.788 | |
| 212 | 0.359 | 0.698 | 0.682 | |
| 258 | 0.437 | 0.657 | 0.639 |
| Subset | Replicate AUROC [95% CI] | Pearson | (Spearman–Brown) [95% CI] | Sign agr. | |||
|---|---|---|---|---|---|---|---|
| 1 | all | 590 | 0.726 [0.691, 0.760] | 0.407 | 0.578 [0.514, 0.631] | 0.630 | 0.331 |
| 1 | first | 275 | 0.765 [0.725, 0.807] | 0.550 | 0.710 [0.656, 0.763] | 0.661 | 0.393 |
| 1 | later | 315 | 0.673 [0.608, 0.726] | 0.320 | 0.484 [0.372, 0.568] | 0.602 | 0.263 |
| 3 | all | 590 | 0.718 [0.686, 0.746] | 0.397 | 0.568 [0.505, 0.624] | 0.551 | 0.282 |
| 3 | first | 275 | 0.742 [0.699, 0.783] | 0.509 | 0.674 [0.615, 0.737] | 0.609 | 0.296 |
| 3 | later | 315 | 0.683 [0.634, 0.725] | 0.331 | 0.496 [0.384, 0.580] | 0.500 | 0.237 |
| Model | Features | AUROC of rep.-avg. scores [95% CI] | Per-rep. mean (min–max) | AP |
|---|---|---|---|---|
| Boosted | H (17) | 0.657 [0.608, 0.703] | 0.645 (0.623–0.659) | 0.699 |
| L2 logistic | H (17) | 0.629 [0.575, 0.680] | 0.624 (0.606–0.641) | 0.698 |
| Boosted | H+T (20) | 0.654 [0.605, 0.700] | 0.642 (0.621–0.656) | 0.698 |
| L2 logistic | H+T (20) | 0.626 [0.574, 0.676] | 0.621 (0.599–0.638) | 0.701 |
| L2 logistic | has-written + errors | 0.482 [0.429, 0.532] | 0.544 (0.526–0.565) | 0.617 |
| Boosted | has-written + errors | 0.483 [0.429, 0.535] | 0.543 (0.524–0.563) | 0.616 |
| Model | Features | Ret. | Avoided [95% CI] | Mass [95% CI] |
|---|---|---|---|---|
| Boosted | H (17) | 0.800 | 0.247 [0.217, 0.277] | 0.238 [0.201, 0.284] |
| L2 logistic | H (17) | 0.800 | 0.262 [0.233, 0.286] | 0.248 [0.211, 0.286] |
| Boosted | H+T (20) | 0.800 | 0.250 [0.219, 0.273] | 0.249 [0.208, 0.287] |
| L2 logistic | H+T (20) | 0.800 | 0.273 [0.243, 0.299] | 0.266 [0.227, 0.303] |
| L2 logistic | has-written + errors | 0.829 | 0.209 [0.180, 0.255] | 0.164 [0.134, 0.214] |
| Boosted | has-written + errors | 0.805 | 0.227 [0.183, 0.256] | 0.179 [0.134, 0.213] |
| Model | Features | count [95% CI] | mass [95% CI] |
|---|---|---|---|
| Boosted | H (17) | 0.047 [ 0.019, 0.079] | 0.038 [ 0.003, 0.086] |
| L2 logistic | H (17) | 0.062 [ 0.035, 0.088] | 0.048 [ 0.012, 0.088] |
| Boosted | H+T (20) | 0.050 [ 0.021, 0.075] | 0.049 [ 0.010, 0.088] |
| L2 logistic | H+T (20) | 0.073 [ 0.045, 0.100] | 0.066 [ 0.029, 0.105] |
| L2 logistic | has-written + errors | 0.038 [ 0.002, 0.065] | 0.007 [ 0.051, 0.026] |
| Boosted | has-written + errors | 0.032 [ 0.004, 0.067] | 0.016 [ 0.053, 0.028] |
| Learner | Repr. | Features | AUROC [95% CI] | Per-rep. mean (min–max) | AP |
|---|---|---|---|---|---|
| L2 logistic | R1 | 144 | 0.654 [0.606, 0.698] | 0.644 (0.622–0.661) | 0.712 |
| L2 logistic | R2 | 4,240 | 0.606 [0.555, 0.655] | 0.594 (0.569–0.621) | 0.672 |
| L2 logistic | R3 | 8,336 | 0.612 [0.561, 0.660] | 0.601 (0.569–0.628) | 0.680 |
| Boosted | R1 | 144 | 0.663 [0.612, 0.708] | 0.651 (0.635–0.667) | 0.709 |
| Boosted | R2 | 4,240 | 0.642 [0.592, 0.686] | 0.628 (0.607–0.648) | 0.688 |
| Boosted | R3 | 8,336 | 0.651 [0.603, 0.695] | 0.636 (0.621–0.655) | 0.695 |
| Learner | Repr. | Features | Avoided [95% CI] | count [95% CI] | mass [95% CI] |
|---|---|---|---|---|---|
| L2 logistic | R1 | 144 | 0.273 [0.243, 0.296] | [ , ] | [ , ] |
| L2 logistic | R2 | 4,240 | 0.253 [0.221, 0.276] | [ , ] | [ , ] |
| L2 logistic | R3 | 8,336 | 0.253 [0.225, 0.279] | [ , ] | [ , ] |
| Boosted | R1 | 144 | 0.267 [0.238, 0.290] | [ , ] | [ , ] |
| Boosted | R2 | 4,240 | 0.262 [0.231, 0.284] | [ , ] | [ , ] |
| Boosted | R3 | 8,336 | 0.262 [0.234, 0.288] | [ , ] | [ , ] |