No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow
Organizations: UC Berkeley · Carnegie Mellon University · Allen Institute for AI
Abstract
Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than others? In this work, we define a notion of Corpus Task Complexity (CTC) that characterizes tasks by how their difficulty grows with corpus size; for instance, a retrieval query only requires a single linear pass over a corpus, while finding contradictions requires checking a quadratically growing set of claim pairs. Observing that prior work has largely only studied tasks whose difficulty grows linearly with corpus size, which we call low CTC tasks, we introduce 10 new tasks belonging to a class of high CTC whose difficulty grows quadratically or more in corpus size. We find that high-CTC tasks not only grow much more challenging on average at longer contexts for LCLMs, they reverse many modeling conclusions drawn solely from low-CTC evaluations. For instance, efficient block-sparse and hybrid attention approaches consistently match full attention performance on low-CTC tasks, but degrade much more on high-CTC tasks. Large-corpus high-CTC reasoning thus remains an open challenge as full attention is too costly to scale, motivating future research on these tasks. We release our code, data, and 22-task suite (CTC-Bench), to facilitate future research in this area.
Figures & tables
| : Low Complexity Tasks | Higher-complexity tasks | ||
|---|---|---|---|
| Dataset | Description | Dataset | Description |
| NIAH-contra | Find the contradicting claim | Contradiction | Find all contradicting claim pairs |
| SciFact | Retrieve evidence for a scientific claim | X-Absence | Find unmatched docs across 2 shuffled, near-identical corpora |
| FiQA | Retrieve relevant financial opinions | QDmatch (HPQA) | Match questions to two documents |
| MS MARCO | Retrieve relevant web passages | QDmatch (NQ) | Match questions to documents |
| OBLIQ | Retrieve passages for subjective queries | QDmatch (FiQA) | Match financial queries to documents |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Example |
|---|---|
| : One-pass tasks | |
| NQ | “who sold out jesus for 30 pieces of silver” [1] |
| HotpotQA (bridge) | “What company did Rex Maughan aquire?” [8, 9] |
| NIAH-contra | “ a hypertonic solution of 14.4% has an ir reversible ciliostatic effect” [20] , the one claim saying reversible |
| BEIR SciFact | “0-dimensional biomaterials show inductive properties.” [1] |
| BEIR FiQA | “Where should I park my rainy-day / emergency fund?” [2, 8, 10, 12, 19] |
| Documents per example ( ) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | Source corpus | CTC | Metric | 2k | 4k | 8k | 16k | 32k | Chars/doc | Eval |
| NQ | Wikipedia 100w (DPR) | gold-ID F1 | 11 | 23 | 48 | 104 | 208 | 636 | 500 | |
| HotpotQA (bridge) | Wikipedia 100w (DPR) | gold-ID F1 | 17 | 36 | 72 | 144 | 296 | 432 | 500 | |
| NIAH-contra | PubMed claims | gold-ID F1 | 40 | 86 | 180 | 365 | 740 | 60 | 500 | |
| BEIR SciFact | SciFact abstracts | gold-ID F1 | 5 | 10 | 21 | 43 | 88 | 1517 | 300 | |
| BEIR FiQA | FiQA-2018 posts | gold-ID F1 | 8 | 19 | 40 | 82 | 166 | 835 | 500 | |