Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.
Figures & tables
Figure 1: LongHarness (ours) reveals harness differences in both accuracy and cost.
Figure 2: Examples from our benchmark. Top left: Constraint Solving Search combines evidence from documents. Top right: Equivalent Program Pair Search distinguishes equivalent programs from near matches. Bottom left: Program Execution Tracing requires reading prior outputs and matching code semantics. Bottom right: Outlier Memo Detection identifies a conflict using supporting memos.
LongBenchV2
HELMET
LongProc
Oolong
LongHarness
Diverse, adaptive retrieval demand
∼
∼
×
×
✓
Semantically confusable evidence
×
✓
∼
∼
✓
Extensive reasoning steps
×
×
✓
✓
✓
Step-dependent retrieval
×
×
✓
×
✓
Multiple strategies with different costs
×
∼
×
∼
✓
Wide range of SoTA system accuracy
×
×
×
×
✓
Table 1: Comparison of long-context benchmarks along properties relevant to evaluating LM harnesses. ✓ : supported; ∼ : partially supported; × : not supported. Diverse, adaptive retrieval demand means choosing an appropriate way to access context, such as lexical search, semantic retrieval, or direct reading, as information needs change during reasoning. Step-dependent retrieval means that intermediate findings determine what must be retrieved next. The final two rows indicate whether model–harness configurations are widely separated in accuracy and cost.
Model
Harness
Constraint
Memos
Program
Pairs
Macro avg.
GPT-5.6-sol
Direct
58% ($0.690)
44% ($1.17)
94% ($0.648)
44% ($0.780)
60.0% ($0.822)
RLM
44% ($0.659)
92% ($7.23)
76% ($1.47)
2% ($7.86)
53.5% ($4.30)
OpenCode
56% ($1.72)
10% ($3.11)
96% ($1.37)
0% ($5.41)
40.5% ($2.90)
mini-swe-agent
42% ($0.320)
100% ($0.583)
88% ($1.02)
42% ($1.34)
68.0% ($0.816)
ReAct
32% ($1.40)
6% ($2.77)
82% ($0.866)
52% ($2.68)
43.0% ($1.93)
GLM-5.3
Direct
0% ($0.728)
0% ($1.09)
38% ($0.785)
0% ($1.19)
9.5% ($0.948)
Table 2: Exact accuracy and estimated cost across four suites. Each cell reports accuracy (%) followed by USD per-instance cost in parentheses; macro averages the four task accuracies and per-task mean costs. Appendix A gives token counts and Appendix C gives the API rates.
Figure 3: Macro-average accuracy and estimated cost across the four tasks (Table 2 ). Color denotes model; shape denotes harness. Callouts and larger markers mark each model’s best accuracy. ReAct appears where available.
Figure 4: Harness effects relative to direct inference on OOLONG-Synth, LongBench-v2, and our suites. External benchmarks use GPT-5.6-sol. The reference is zero accuracy change and 1× cost. Green denotes higher accuracy at lower cost; pink denotes lower accuracy at higher cost.
Harness
OOLONG-Synth
LongBench-v2
Direct
64% ($0.686)
62% ($0.532)
RLM
78% ($3.70)
62% ($0.641)
OpenCode
74% ($2.08)
60% ($0.463)
mini-swe-agent
86% ($0.217)
68% ($0.190)
ReAct
70% ($0.751)
68% ($0.354)
Table 3: GPT-5.6-sol on existing long-context benchmarks. Each cell reports accuracy and average USD per instance in parentheses.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Private specification
Controlled distractors
Gold derivation
Constraint Solving Search
Typed fact world and predicate support sets
People satisfying all but one condition
Set intersection and fact-to-passage map
Outlier Memo Detection
Typed relation graph and composed relations
Endpoint substitutions and near-twin claims
Graph-composition proofs
Program Execution Tracing
Computation graph, typed inputs, and function contracts
Same-output functions and alternate routes
Execution and contract probes
Equivalent Program Pair Search
Executable program families and test signatures
Behavior-changing syntax-tree mutations
Differential execution
Appendix
Table 5: Underlying structures used to construct the four tasks. Public contexts expose only the rendered items, not their roles or provenance.
Model
Input
Cache read
Cache write
Output
GPT-5.6-sol
4.00
0.400
5.00
20.00
GPT-5.6-sol (long)
8.00
0.800
10.00
30.00
GLM-5.3
4.86
0.972
4.86
12.15
Qwen3.8-27B
2.48
–
–
7.46
Gemini 3.8 Flash
0.75
0.075
–
3.75
Kimi-K2.6
5.15
1.03
5.15
12.81
Appendix
Table 6: API rates in USD per million tokens. A dash indicates that we do not apply a separate rate for that token category.
Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. However, existing long-context evaluations - from Needle-in-a-Haystack (NIAH) tests to more recent multi-hop reasoning and summarization tasks - predominantly measure average-case performance, and many are either saturated or lack robustness. Notably absent is a systematic way to probe how models perform as we scale up the difficulty of tasks along various axes. We address this gap by proposing PredicateLongBench, a benchmark that stress-tests long-context reasoning by asking models to identify the longest contiguous subsequence of words in a long input that satisfies given predicates/constraints (e.g., lexicographic ordering), drawn from a broader predicate class. The central innovation of our benchmark is the identification and systematic exploration of multiple different axes of difficulty which test multiple aspects of long context understanding. We provide two complementary generation pipelines - a fully synthetic setup using random word-like strings, and a real-world setup that samples words from natural documents while preserving their distributional properties. We find that frontier models struggle to perform well as we scale up the difficulty of tasks along our axes, demonstrating the utility of our benchmark in understanding the limitations of current long-context capabilities. Furthermore, the tasks in PredicateLongBench, though challenging, are conceptually simple and do not require LLM-based generations or judges.
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent's seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task family, masking two failure modes: performance can collapse as length grows, and strong retrieval need not transfer to downstream use. We present ATLAS, a benchmarking framework that redefines long-context evaluation as length-dependent capability profiling. ATLAS contributes three methodological principles:(i) a layered taxonomy separating foundational operations from application workloads so failures can be attributed, (ii) length-aware AUC scoring that integrates score-length curves over a fixed 8K-1M grid, replacing single-point metrics with full degradation profiles, and (iii) ATLAScore, a harmonic-mean aggregate over taxonomy categories that penalizes imbalanced profiles, with end-to-end uncertainty propagation from subset scores through the nonlinear final aggregate. We instantiate the framework across eight capability dimensions with nine auditable components and 6,438 instances, and evaluate 26 models. Gemini-3.1-Pro-Preview leads at 128K, Claude-Opus-4.6 leads at 1M. Rankings reshuffle substantially between ATLASscore@8K-128K and ATLASscore@8K-1M: 7 models move by at least two ranks, and the two taxonomy layers share only 61% of cross-model variance, with individual rank gaps up to 12 positions. These results support reporting long-context quality by capability and length, not by a single headline score.