Evaluating Budgeted Context Projection with Unexecuted Companion Runs
Organizations: Independent AI Researcher
Abstract
Context projection can shorten individual requests while changing whether an agent finishes within its budget. We examine how sequential evaluation obscures this trade-off when a capped first continuation prevents its companion from running. In a recorded ReVerPi source-reading campaign, 15 pairs with two final answers yield 12 historically scored successes per arm. Retaining all 27 intervention boundaries distinguishes observed failures from ten unexecuted companions and bounds projected-minus-full success between and tasks. Under the archived scoring contract, a frozen projection selector has a success difference from full context of ; outside four fitting tasks, it is across 23 boundaries. Excluding one task whose platform premise is not established by retained actor-input evidence changes the nonfitting range to across 22 boundaries, so strict inferiority is not robust to that exclusion. Among eleven historically joint-success pairs, projection uses 25% fewer aggregate logical tokens but more tokens for the median pair and 55 rather than 35 suffix requests. These retrospective results concern one adaptively assembled campaign, not population performance. The case motivates accounting that retains every boundary, preserves unknown outcomes and policy dependencies, checks task premises, and separates bounded completion from success-conditioned resource use.
Figures & tables
| Setting | Recorded value |
|---|---|
| Pi / protocol | 0.84.2 / Responses |
| Proxy-returned model | deepseek-flash |
| Effort / concurrency | low / 1 |
| Capture / suffix caps | 12 / 12 newly dispatched model requests |
| Per-phase wall limit | 3,600 seconds |
| Context / output reservation | 131,072 / 65,536 tokens |
| Runs | Inclusion rule | Use |
|---|---|---|
| 76 | Paired/capture plans | Campaign accounting and eligibility |
| 27 | Realized paired boundary | Primary bounded-success and selector ranges |
| 17 | Both bounded outcomes known | Secondary observed-expenditure policy comparison |
| 15 | Both arms return a final answer | Historical completed-pair summary |
| 11 | Both recorded answers correct | Resource use conditional on joint success |
| Recorded endpoint | Full | Projected |
|---|---|---|
| Correct | 14 | 12 |
| Known bounded failure | 6 | 12 |
| Unexecuted | 7 | 3 |
| Frame | Runs | Success-count range | Difference (pp) |
|---|---|---|---|
| All paired boundaries | 27 | ||
| Outside four fitting tasks | 23 | ||
| Outside four fitting sources | 22 |
| Reduced frame | Runs | Count range | Difference (pp) |
|---|---|---|---|
| All boundaries except pathspec-util | 26 | ||
| Outside fitting tasks, also excluding pathspec-util | 22 | ||
| Outside fitting sources, also excluding pathspec-util | 21 |
| Frame | Runs | F failures | Rule failures | F tokens | Rule tokens |
|---|---|---|---|---|---|
| Fitting | 4 | 1 | 0 | 404,703 | 298,261 |
| Outside fitting | 13 | 2 | 3 | 1,419,197 | 1,541,879 |
| All comparable | 17 | 3 | 3 | 1,823,900 | 1,840,140 |
| Category | Full | Projected |
|---|---|---|
| Uncached input | 264,132 | 353,816 |
| Cached input | 982,272 | 574,592 |
| Output | 20,632 | 21,366 |
| Total | 1,267,036 | 949,774 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Selected minus full | ||
|---|---|---|---|
| packaging-compatibility ∗ | 0 | 1 | |
| tomlkit-datetimes ∗ | 1 | 1 | |
| filelock-contention | 1 | 1 | |
| more-itertools-window | 1 | 1 | |
| pathspec-util | 1 | 0 | |
| docutils-tree | ? | 0 |
| Task | First | History bytes | |||||
|---|---|---|---|---|---|---|---|
| click-parameters ∗ | F | 62,102 | 1/1 | 3/8 | 76,016 | 103,353 | 1.360 |
| packaging-compatibility ∗ | F | 95,519 | 0/1 | 1/1 | 57,171 | 51,359 | 0.898 |
| boltons-indexedset ∗ | F | 52,266 | 1/1 | 2/6 | 73,859 | 133,797 | 1.812 |
| cerberus-errors | F | 52,968 | 0/0 | 2/5 | 48,494 | 63,500 | 1.309 |
| dotenv-stream | P | 25,110 | 1/1 | 2/5 | 26,662 | 38,865 | 1.458 |
| filelock-basics | F | 72,397 | 1/1 | 3/11 | 91,143 | 161,204 | 1.769 |
| Task | First | History bytes | Rule | ||
|---|---|---|---|---|---|
| attrs-define | P | 58,749 | F | ?/0 | 0/12 |
| docutils-nodes | F | 72,019 | F | 1/0 | 8/12 |
| docutils-tree | P | 154,876 | P | ?/0 | 0/12 |
| markdown-extensions | P | 130,201 | P | ?/0 | 0/12 |
| marshmallow-fields | P | 51,348 | F | ?/0 | 0/12 |
| pathspec-util | F | 94,493 | P | 1/0 | 3/12 |
| Task | Paired | Repeat | Choice | Outcome | Tokens | |
|---|---|---|---|---|---|---|
| filelock-basics | 72,397 | 77,200 | F F | 2 | correct | 79,050 |
| filelock-contention | 96,013 | 123,601 | P P | 10 | correct | 206,300 |
| more-itertools-recipes | 56,382 | 110,194 | F P | 2 | correct | 100,078 |
| pathspec-match | 63,576 | 97,045 | F P | 12 | capped | 244,370 |
| sqlparse-grouping | 89,881 | 64,817 | F F | 5 | correct | 135,759 |
| urllib3-headers | 62,556 | 54,818 | F F | 2 | correct | 49,185 |
| Task | ||||||||
|---|---|---|---|---|---|---|---|---|
| more-itertools-recipes | 60,479 | 352,768 | 5,117 | 15,514 | 18,944 | 1,368 | 9 | 1 |
| tomlkit-datetimes ∗ | 28,063 | 167,424 | 2,170 | 42,956 | 52,480 | 1,591 | 6 | 5 |
| filelock-contention | 24,917 | 87,552 | 2,395 | 24,380 | 38,528 | 2,557 | 3 | 4 |
| pathspec-match | 22,217 | 41,472 | 2,759 | 28,724 | 17,792 | 2,359 | 2 | 2 |
| sqlparse-grouping | 24,755 | 49,152 | 2,067 | 46,517 | 45,056 | 2,477 | 2 | 6 |
| more-itertools-window | 29,340 | 33,792 | 861 | 44,430 | 36,992 | 1,225 | 1 | 2 |
| Policy | Failures | Suffix tokens | Logical tokens | Change vs F |
|---|---|---|---|---|
| Always-full | 3 | 1,313,508 | 1,823,900 | — |
| Always-projected | 5 | 1,240,753 | 1,751,145 | |
| Frozen threshold 93,641 | 3 | 1,329,748 | 1,840,140 | |
| Hindsight threshold 95,519 | 2 | 1,276,951 | 1,787,343 |
| Summary | Historical | Strict format |
|---|---|---|
| Jointly correct pairs | 11 | 10 |
| Full logical tokens | 1,267,036 | 1,191,020 |
| Projected logical tokens | 949,774 | 846,421 |
| Ratio of sums | 0.749603 | 0.710669 |
| Geometric mean ratio | 0.883374 | 0.846091 |
| Median ratio | 1.291501 | 1.264712 |
| Source/version | Files | Boundary | Pair | Snapshot SHA prefix |
|---|---|---|---|---|
| attrs-26.1.0 | 20 | 1 | 0 | fc4d96802c3d |
| boltons-26.2.0 | 31 | 1 | 1 | b80f6739e270 |
| cachetools-7.1.4 | 9 | 0 | 0 | 1fa516c2685d |
| cerberus-1.3.8 | 7 | 1 | 1 | f4bb235519e6 |
| click-8.1.8 | 17 | 1 | 1 | 44353e20a2fd |
| colorama-0.4.6 | 14 | 0 | 0 | be58df74478d |
| Task | Non-allowlisted read intents | Aborted results | Pending results |
|---|---|---|---|
| platformdirs-paths | 1 | 1 | 0 |
| sqlparse-format | 2 | 1 | 1 |
| wheel-metadata | 3 | 1 | 2 |