Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than others? In this work, we define a notion of Corpus Task Complexity (CTC) that characterizes tasks by how their difficulty grows with corpus size; for instance, a retrieval query only requires a single linear pass over a corpus, while finding contradictions requires checking a quadratically growing set of claim pairs. Observing that prior work has largely only studied tasks whose difficulty grows linearly with corpus size, which we call low CTC tasks, we introduce 10 new tasks belonging to a class of high CTC whose difficulty grows quadratically or more in corpus size. We find that high-CTC tasks not only grow much more challenging on average at longer contexts for LCLMs, they reverse many modeling conclusions drawn solely from low-CTC evaluations. For instance, efficient block-sparse and hybrid attention approaches consistently match full attention performance on low-CTC tasks, but degrade much more on high-CTC tasks. Large-corpus high-CTC reasoning thus remains an open challenge as full attention is too costly to scale, motivating future research on these tasks. We release our code, data, and 22-task suite (CTC-Bench), to facilitate future research in this area.
Figures & tables
Figure 1: Performance of in-distribution trained Qwen3.5-4B models on low- and high-CTC tasks. (Left) Performance degrades faster on high-CTC tasks than on low-CTC tasks as corpus size increases. (Right) Block-sparse attention causes almost no performance degradation relative to full attention on low-CTC tasks, but substantial degradation on high-CTC tasks.
Figure 2: Circles represent documents, with shaded circles for the gold documents. As corpus size N=∣C∣ grows from 3 documents to 100 documents, a contradiction search task may in worst-case require many more oracle operations (4950) than factoid retrieval (100).
O(N) : Low Complexity Tasks
Higher-complexity tasks
Dataset
Description
Dataset
Description
NIAH-contra
Find the contradicting claim
Contradiction
Find all contradicting claim pairs
SciFact
Retrieve evidence for a scientific claim
X-Absence
Find unmatched docs across 2 shuffled, near-identical corpora
FiQA
Retrieve relevant financial opinions
QDmatch (HPQA)
Match questions to two documents
MS MARCO
Retrieve relevant web passages
QDmatch (NQ)
Match questions to documents
OBLIQ
Retrieve passages for subjective queries
QDmatch (FiQA)
Match financial queries to documents
Table 1: Overview of the 22 CTC-Bench Tasks. Linear-time tasks are shown on the left; tasks requiring higher-order comparisons are shown on the right (details, examples in Appendix H )
Figure 3: Scaling of full vs. block-sparse attention (mask mix) We show block-sparse vs full attention performance at different contexts on our full task suite. While both methods perform near-identically on low-CTC, high-CTC tasks correspond to growing gaps between them on larger corpora.
Figure 4: Hybrid vs Full Models We plot CTC-Bench -10 results for a full-attention (solid) version of OLMo-3-7B vs OLMo-3-Hybrid (dot-dashed). On low-CTC tasks both perform similarly, but on high-CTC tasks we observe consistent performance gaps.
Figure 5: (Left) Gain of curriculum mask-mixing over pure block-sparse training (no mixing) , averaged across 2k–32k contexts on CTC-Bench -10. Mask-mixing consistently improves block-sparse, especially on High-CTC tasks and tasks more dispersed across contexts (such as OOLONG). (Right) Length generalization (full attention). For 64k and 128k context eval sets (trained only to 32k), we plot percentage of 32k score. High-CTC tasks are harder for length generalization.
Figure 6: (Left) Model scale/family Percentage of full-attention performance lost to block-sparse attention at 4k, 8k and 16k context, for three Qwen3.5 scales and two other model families, on a low-CTC (HotpotQA, OT(N) ) and high-CTC (Contradiction, OT(N2) ) task. The gap is minimal on low-CTC and large and growing on high-CTC; within the Qwen3.5 family (shades of purple) smaller models degrade faster. (Right) Oracle Difficulty For 4 task pairs with matched oracle operations (block-sparse attention), ratio of OT(N) to OT(N2) task performance on the same corpus (pairs named in the legend). Tasks with harder base operations (darker) see the ratio grow faster.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Example
O(N) : One-pass tasks
NQ
“who sold out jesus for 30 pieces of silver” → [1]
HotpotQA (bridge)
“What company did Rex Maughan aquire?” → [8, 9]
NIAH-contra
“ … a hypertonic solution of 14.4% has an ir reversible ciliostatic effect” → [20] , the one claim saying reversible
BEIR SciFact
“0-dimensional biomaterials show inductive properties.” → [1]
BEIR FiQA
“Where should I park my rainy-day / emergency fund?” → [2, 8, 10, 12, 19]
Appendix
Table 2: One real example per task. All content is verbatim from the evaluation sets, elided with … where a passage is too long to show; document IDs are 1-indexed exactly as the model sees them. Appendix H shows each example in full, with the task instruction, the surrounding corpus and the distractor types.
Documents per example ( N )
Task
Source corpus
CTC
Metric
2k
4k
8k
16k
32k
Chars/doc
Eval
NQ
Wikipedia 100w (DPR)
OT(N)
gold-ID F1
11
23
48
104
208
636
500
HotpotQA (bridge)
Wikipedia 100w (DPR)
OT(N)
gold-ID F1
17
36
72
144
296
432
500
NIAH-contra
PubMed claims
OT(N)
gold-ID F1
40
86
180
365
740
60
500
BEIR SciFact
SciFact abstracts
OT(N)
gold-ID F1
5
10
21
43
88
1517
300
BEIR FiQA
FiQA-2018 posts
OT(N)
gold-ID F1
8
19
40
82
166
835
500
Appendix
Table 3: Corpus statistics for every task in the evaluation suite Column headings 2k–32k are approximate token targets per task; the entries are the median number of documents actually present in an example of that rung, measured from the evaluation files the reported results were graded on. Chars/doc is the median document length in characters at the deepest available rung.
Figure 7: Full model performance on several tasks and Qwen3.5 model scales. Some high-CTC tasks have minimum model scales necessary to learn effectively.
Figure 8: Full vs block-sparse attention with a few different model families.
Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investigate whether LLMs, trained on human-preferred corpora that embed attentional biases, exhibit a similar limitation: \emph{failing to attend to subtle yet important contextual cues under explicit task instructions}. To evaluate this, we introduce the task of \textbf{explicit-implicit reasoning} and present \textbf{MixRea}, a benchmark of 2,246 multiple-choice questions across 9 reasoning types with varying distributions of explicit and implicit information. Evaluation of 21 advanced LLMs shows that even the best-performing reasoning model (Gemini 2.5 Pro) achieves only 42.8% consistency, revealing widespread inattentional blindness. To mitigate this, we propose \textbf{Potential Relation Completion Prompting (PRCP)}, a prompting method that improves reasoning by recovering overlooked causal relations. Further analysis shows that this limitation persists across diverse multi-source reasoning tasks, highlighting the need for more cognitively aligned models.
Yuanqing Cai, Ziyi Huang, Minhao Liu +3
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evaluations have recently been released, these evaluations tend to rely on retrieval from one or more sections of the context, which allows nearly all of the context tokens to be disregarded as noise. This represents only one type of task that might be performed with long context. We introduce Oolong, a benchmark of long-context reasoning tasks that require analyzing individual chunks of text on an atomic level, and then aggregating these analyses to answer distributional questions. Oolong is separated into two task sets: Oolong-synth, a set of naturalistic synthetic tasks, where we can easily ablate components of the reasoning problem; and Oolong-real, a downstream setting which requires reasoning over real-world conversational data. Oolong requires models to reason over large quantities of examples, to perform both classification and counting in-context, and to reason over temporal and user relations. Even frontier models struggle on Oolong, with GPT-5, Claude-Sonnet-4, and Gemini-2.5-Pro all achieving less than 50% accuracy on both splits at 128K. We release the data and evaluation harness for Oolong to enable further development of models that can reason over large quantities of text.
Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely exists outside of curated benchmarks built on structured sources. To construct it directly from plain, unannotated text, existing methods ask a single teacher model to jointly discover an evidence path through a document and verbalize it as a question-answer pair. However, these methods degrade sharply when documents are structured around repetitive templates and densely cross-referencing clauses, conditions that characterize most real-world specialized corpora. In this work, we decouple the two operations: reasoning paths are enumerated offline over a graph of contextual keyword centroids, and the teacher is invoked only to verbalize pre-validated paths. The graph enforces five geometric admissibility constraints, for which we provide Gram-matrix arguments establishing that local similarity bounds alone admit endpoint drift up to ∼91∘, and that an upper similarity bound is necessary to exit dense embedding cliques formed by boilerplate text. A matched-size ablation isolates the mechanism: at equal training scale, constrained and unconstrained chains yield indistinguishable downstream performance, and the gain at full scale comes from a 4.4× expansion of the usable corpus rather than from higher per-chain quality -- reframing the role of graph constraints, in this setting, as raising teacher synthesizability rather than improving chain content. Fine-tuning Qwen3-32B on 80K examples constructed from the CUAD legal contract corpus improves closed-book Token F1 from 21.66% to 38.58%. We have released our codes at https://github.com/hkgai-official/GCSCS.
Pengyu Chen, Yonggang Zhang, Mingming Chen +3
The Hong Kong University of Science and Technology · Hong Kong Generative AI Research and Development Center · Hong Kong Baptist University