cs.AISep 25, 2026

Evaluating Budgeted Context Projection with Unexecuted Companion Runs

Authors: Guangzhe Zhang

Organizations: Independent AI Researcher

Abstract

Context projection can shorten individual requests while changing whether an agent finishes within its budget. We examine how sequential evaluation obscures this trade-off when a capped first continuation prevents its companion from running. In a recorded ReVerPi source-reading campaign, 15 pairs with two final answers yield 12 historically scored successes per arm. Retaining all 27 intervention boundaries distinguishes observed failures from ten unexecuted companions and bounds projected-minus-full success between −9-9 and +1+1 tasks. Under the archived scoring contract, a frozen projection selector has a success difference from full context of [−3,0][-3,0]; outside four fitting tasks, it is [−4,−1][-4,-1] across 23 boundaries. Excluding one task whose platform premise is not established by retained actor-input evidence changes the nonfitting range to [−3,0][-3,0] across 22 boundaries, so strict inferiority is not robust to that exclusion. Among eleven historically joint-success pairs, projection uses 25% fewer aggregate logical tokens but more tokens for the median pair and 55 rather than 35 suffix requests. These retrospective results concern one adaptively assembled campaign, not population performance. The case motivates accounting that retains every boundary, preserves unknown outcomes and policy dependencies, checks task premises, and separates bounded completion from success-conditioned resource use.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 30, 2026cs.AI

Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare and differ widely in how much they matter. Standard benchmarks spend this budget uniformly: a read-only lookup is sampled as often as an irreversible payment action. We instead formulate evaluation as a sequential allocation problem. Given a fixed trial budget and a set of scenarios whose failure behavior is unknown, which scenarios should be run, and run again? We propose a risk-aware contextual Thompson Sampling policy that combines a pre-execution scenario context vector and a fixed impact score with the failure outcomes observed during evaluation, and we test it by offline replay over 70 ττ-bench airline scenarios and 824 recorded trials. Our main result is at the smallest budget: with only 50 trials (6%6\% of the corpus), the policy recovers 86%86\% of the impact-weighted failures an oracle could find, compared to 25%25\% for uniform allocation. It discovers 3.5×3.5\times more impact-weighted failures (215.4 vs. 62.2) with the same number of trials, delivers 5×5\times the discovery per dollar, and cuts the budget wasted on scenarios that never fail from 34%34\% to 2.8%2.8\%. The rest of our analysis demonstrates and qualifies this result: a budget sweep shows the advantage shrinks as the budget approaches the corpus size, and paired significance tests show that scenario context helps mainly at small budgets while posterior-based exploration helps at moderate ones. Risk-aware adaptive allocation therefore helps most exactly where evaluation budget is scarcest.
Jul 29, 2026cs.LG

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance

Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary errors to compound; longer tasks may also contain harder individual decisions or become harder as conversation history, tool outputs, and environment changes accumulate. We use trajectory-induced degradation to mean this last possibility: earlier execution makes later work harder. When the harmful accumulation is specifically the text visible to the model, it is often called context rot. In this position paper, we argue that to claim a "long-horizon failure", benchmarks must compare actual full-task success against a baseline prediction built from short, individual stages. We call the log-ratio between this prediction and actual success the horizon residual. The comparison must use the same agent configuration and specify in advance how stages, checkpoints, information, and budgets will be chosen. The residual shows that the full rollout differs from the chosen baseline; targeted experiments are still needed to explain why.
Jul 27, 2026cs.AI

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes intended reasoning from answer acquisition. Outcome evidence and exposure detection do not establish whether success depended on an acquired target; we call this missing evaluation object success provenance. AcquaBench audits it through matched CLEAN, GOLD, and SHAM value substitution on four standardized surfaces with joint qid-clustered analysis. CLEAN retains benchmark-authorized information. GOLD makes the correct target available. SHAM preserves source structure and exposure opportunity but substitutes a matched incorrect value. GOLD minus CLEAN measures the total score response to correct-target availability; GOLD minus SHAM tests whether that response tracks target correctness beyond matched source exposure. In D0, GOLD exceeds SHAM by 19.1 to 25.9 percentage points, showing that success follows the correct value. In D2, GOLD still exceeds SHAM under distributed sufficiency while coloc no longer transfers as a high-score marker, with AUROC 0.376 and 0.142. Behavioral dependence can thus persist beyond this probe's intended observation unit. In model comparison, a supported 5.0-point CLEAN score gap compresses to a raw GOLD difference of -0.6 points without establishing rank inversion. Agent benchmarks should report success together with whether the evaluated information state supported it.