Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.
Figures & tables
Figure 1: LongHarness (ours) reveals harness differences in both accuracy and cost.
Figure 2: Examples from our benchmark. Top left: Constraint Solving Search combines evidence from documents. Top right: Equivalent Program Pair Search distinguishes equivalent programs from near matches. Bottom left: Program Execution Tracing requires reading prior outputs and matching code semantics. Bottom right: Outlier Memo Detection identifies a conflict using supporting memos.
LongBenchV2
HELMET
LongProc
Oolong
LongHarness
Diverse, adaptive retrieval demand
∼
∼
×
×
✓
Semantically confusable evidence
×
✓
∼
∼
✓
Extensive reasoning steps
×
×
✓
✓
✓
Step-dependent retrieval
×
×
✓
×
✓
Multiple strategies with different costs
×
∼
×
∼
✓
Wide range of SoTA system accuracy
×
×
×
×
✓
Table 1: Comparison of long-context benchmarks along properties relevant to evaluating LM harnesses. ✓ : supported; ∼ : partially supported; × : not supported. Diverse, adaptive retrieval demand means choosing an appropriate way to access context, such as lexical search, semantic retrieval, or direct reading, as information needs change during reasoning. Step-dependent retrieval means that intermediate findings determine what must be retrieved next. The final two rows indicate whether model–harness configurations are widely separated in accuracy and cost.
Model
Harness
Constraint
Memos
Program
Pairs
Macro avg.
GPT-5.6-sol
Direct
58% ($0.690)
44% ($1.17)
94% ($0.648)
44% ($0.780)
60.0% ($0.822)
RLM
44% ($0.659)
92% ($7.23)
76% ($1.47)
2% ($7.86)
53.5% ($4.30)
OpenCode
56% ($1.72)
10% ($3.11)
96% ($1.37)
0% ($5.41)
40.5% ($2.90)
mini-swe-agent
42% ($0.320)
100% ($0.583)
88% ($1.02)
42% ($1.34)
68.0% ($0.816)
ReAct
32% ($1.40)
6% ($2.77)
82% ($0.866)
52% ($2.68)
43.0% ($1.93)
GLM-5.3
Direct
0% ($0.728)
0% ($1.09)
38% ($0.785)
0% ($1.19)
9.5% ($0.948)
Table 2: Exact accuracy and estimated cost across four suites. Each cell reports accuracy (%) followed by USD per-instance cost in parentheses; macro averages the four task accuracies and per-task mean costs. Appendix A gives token counts and Appendix C gives the API rates.
Figure 3: Macro-average accuracy and estimated cost across the four tasks (Table 2 ). Color denotes model; shape denotes harness. Callouts and larger markers mark each model’s best accuracy. ReAct appears where available.
Figure 4: Harness effects relative to direct inference on OOLONG-Synth, LongBench-v2, and our suites. External benchmarks use GPT-5.6-sol. The reference is zero accuracy change and 1× cost. Green denotes higher accuracy at lower cost; pink denotes lower accuracy at higher cost.
Harness
OOLONG-Synth
LongBench-v2
Direct
64% ($0.686)
62% ($0.532)
RLM
78% ($3.70)
62% ($0.641)
OpenCode
74% ($2.08)
60% ($0.463)
mini-swe-agent
86% ($0.217)
68% ($0.190)
ReAct
70% ($0.751)
68% ($0.354)
Table 3: GPT-5.6-sol on existing long-context benchmarks. Each cell reports accuracy and average USD per instance in parentheses.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Private specification
Controlled distractors
Gold derivation
Constraint Solving Search
Typed fact world and predicate support sets
People satisfying all but one condition
Set intersection and fact-to-passage map
Outlier Memo Detection
Typed relation graph and composed relations
Endpoint substitutions and near-twin claims
Graph-composition proofs
Program Execution Tracing
Computation graph, typed inputs, and function contracts
Same-output functions and alternate routes
Execution and contract probes
Equivalent Program Pair Search
Executable program families and test signatures
Behavior-changing syntax-tree mutations
Differential execution
Appendix
Table 5: Underlying structures used to construct the four tasks. Public contexts expose only the rendered items, not their roles or provenance.
Model
Input
Cache read
Cache write
Output
GPT-5.6-sol
4.00
0.400
5.00
20.00
GPT-5.6-sol (long)
8.00
0.800
10.00
30.00
GLM-5.3
4.86
0.972
4.86
12.15
Qwen3.8-27B
2.48
–
–
7.46
Gemini 3.8 Flash
0.75
0.075
–
3.75
Kimi-K2.6
5.15
1.03
5.15
12.81
Appendix
Table 6: API rates in USD per million tokens. A dash indicates that we do not apply a separate rate for that token category.