Solving complex tasks in domains such as science, medicine, law, and finance often requires assembling interdependent information scattered across vast, heterogeneous sources far beyond model context limits. Existing approaches tackle this challenge by organizing information into more manageable representations over which models can reason, such as graphs, textual memories, and retrieval collections. These representations dictate what downstream reasoning is possible and, ultimately, whether it succeeds; yet their design and construction remain largely ad hoc. In this work, drawing on the cognitive theory of relevance realization, we propose concrete principles for designing AI systems that construct effective representations of very large contexts. We analyze existing approaches and show how their successes and failures map onto their alignment with these principles, and introduce R3Con, a harness designed to operationalize the principles more systematically. We evaluate R3Con against nine state-of-the-art baselines on two recent benchmarks of reasoning over large document corpora. On these benchmarks, R3Con substantially outperforms the strongest baseline, by 20 and 8.4 percentage points. It also enables smaller models to outperform much larger ones: R3Con with 4B and 9B models outperforms all evaluated 35B baselines, while R3Con with a 35B-A3B model outperforms Claude Code with Claude-Sonnet-5 at 3.7× lower cost. Our results show that context representations following our principled approach can reduce reliance on model scale, pointing toward a future of AI systems with frontier-level performance powered by smaller models. Our code is available at https://github.com/michaeltheologitis/r3con
Figures & tables
Figure 1: R3Con with a 3B -active model outperforms Claude Code equipped with Anthropic’s frontier Sonnet models on CorpusQA ( Lu et al., 2026 ) , while costing roughly 3.7× to 3.9× less.
Figure 2: Running example based on Bram Stoker’s Dracula ( 1897 ) . Full details in § B .
Figure 3: An investigation of the events aboard the ship Demeter. Figuring out that the entire crew of nine died, and that Dracula was behind it, requires connecting evidence across partial and incomplete accounts.
Figure 4: R3Con answers a task over a large context by first constructing an effective representation using our design principles (§ 3.2 ). It extracts the relevant context R (Phase 1) and then uses R to further structure the context (Phase 2). Together, the extracted relevant context and structured context form the representation, which is then used by a coding agent to produce the answer (Phase 3).
Figure 5: Evolution of the relevance snippets for the ship’s log ( D37 ) and The Dailygraph newspaper ( D12 ) across two rounds of extracting relevance. For example, the top-left and top-right cells show R37(1) and R37(2) , respectively. ‡ Document ordering is arbitrary. The full snippets are shown in Fig. 12 .
Table 1: Accuracy (% ↑ ) across real-world domains requiring reasoning over large document corpora. R3Con achieves the highest accuracy across all domains, outperforming the strongest baseline by +19.88 pp on CorpusQA and +8.42 pp on Loong overall (SEs in Tab. 10 ). Accuracy is micro-averaged.
Figure 7
Variant
Loong ( ↑ )
CorpusQA ( ↑ )
R3Con (default)
37.43%
70.64%
[0.6pt/2pt] w/o extracting relevance
23.96% (-13.47)
52.58% (-18.1)
w/o structuring
32.97% (-4.46)
63.22% (-7.42)
Table 3: Accuracy in the ablation of the two phases of R3Con (see Fig. 4 ).
Table 4: Taxonomy of baselines according to the principles (§ 3.2 ) satisfied by their representation-construction phase. See § C.2 for detailed analysis. = limited support; N/A = not applicable.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Document type
# Docs.
Description
Journal / diary
5
Personal journals and diaries (e.g., Jonathan Harker’s journal)
Letter
20
Letters between characters (e.g., Mina to Lucy)
Telegram
8
Telegrams between characters (e.g., Van Helsing to Seward)
Memorandum
4
Written accounts of events (e.g., Lucy’s final memorandum)
Newspapers
4
Newspaper cuttings (e.g., The Pall Mall Gazette)
Note
2
Short written notes (e.g., Van Helsing’s contingency note)
Appendix
Table 5: The reconstructed document collection in Bram Stoker’s Dracula ( 1897 ) .
Model
Parametric knowledge
With corpus in context
Answer
Correct
Answer
Correct
Qwen3.5-35B-A3B ( 2026 )
31
✗
7
✗
Appendix
Table 6: Answers to the Dracula question using parametric knowledge alone and with the full novel provided in context. The ground-truth answer is 13, but also requires the correct people named. Full model responses are shown in Fig. 8 .
Model
Temp.
Top- p
Top- k
Min- p
Pres. pen.
Rep. pen.
Qwen3.5-4B
1.0
0.95
20
0.0
1.5
1.0
Qwen3.5-9B
1.0
0.95
20
0.0
1.5
1.0
Qwen3.5-35B-A3B
1.0
0.95
20
0.0
1.5
1.0
Appendix
Table 7: Sampling parameters used for each model, following the recommendations from the official HuggingFace model cards.
Method
CorpusQA
Loong
Qwen3.5-4B
StructRAG ( 2025b )
29.79%
16.84%
CodeAgent ( 2025 )
10.92%
19.44%
R3Con (ours)
49.24%
30.23%
Qwen3.5-9B
StructRAG ( 2025b )
40.43%
20.30%
Appendix
Table 8: Accuracy ( ↑ ) across benchmarks and model sizes of the strongest baselines.
Figure 9: Distribution of the representation components (i.e., final relevant context and structuring) for the two benchmarks with the default Qwen3.5-35B-A3B model.
Figure 12: Evolution of the relevance snippets for the log of the Demeter ( D37 ) and The Dailygraph newspaper ( D12 ) across two rounds of extracting relevance. For example, the top-left and top-right cells show R37(1) and R37(2) , respectively. ‡ Document ordering is arbitrary.
Downstream Reasoning
CorpusQA
Loong
LLM
65.43%
40.17%
Coding Agent (default)
70.64%
37.43%
Appendix
Table 9: Accuracy ( ↑ ) with different downstream reasoning methods over the representation constructed by R3Con .
Figure 13: Effect of the number of relevance-extraction rounds N on accuracy and cost.
Method
CorpusQA ( Lu et al., 2026 )
Loong ( Wang et al., 2024a )
Edu.
Finance
Real-estate
Overall
Paper
Finance
Legal
Overall
RAPTOR ( 2024 )
8.00 ± 3.13
11.11 ± 2.47
17.95 ± 4.35
12.06 ± 1.84
0.00 ± 0.00
22.17 ± 1.71
5.71 ± 1.24
12.12 ± 0.92
ReadAgent ( 2024 )
42.68 ± 5.46
37.42 ± 3.79
31.25 ± 5.18
37.23 ± 2.68
15.15 ± 2.08
35.33 ± 2.03
14.93 ± 1.95
24.49 ± 1.25
MemAgent ( 2026 )
7.32 ± 2.88
7.83 ± 2.09
8.64 ± 3.12
7.90 ± 1.49
0.00 ± 0.00
26.80 ± 1.81
4.00 ± 1.05
13.63 ± 0.96
HippoRAG2 ( 2025 )
7.32 ± 2.88
7.83 ± 2.09
20.99 ± 4.52
10.94 ± 1.72
0.32 ± 0.32
17.42 ± 1.55
6.29 ± 1.30
10.06 ± 0.85
LinearRAG ( 2026 )
8.54 ± 3.09
6.63 ± 1.93
16.05 ± 4.08
9.42 ± 1.61
0.61 ± 0.43
9.38 ± 1.19
5.43 ± 1.21
6.04 ± 0.67
Appendix
Table 10: Accuracy (% ↑ ) across real-world domains requiring reasoning over large document corpora. R3Con achieves the highest accuracy across all domains, outperforming the strongest baseline by +19.88 pp on CorpusQA and +8.42 pp on Loong overall. Overall accuracy is micro-averaged. We report standard errors (SE) as ± values.
Method
Spot.
Comp.
Chain
Clust.
Overall
RAPTOR ( 2024 )
51.31 ± 3.62
16.25 ± 2.38
2.04 ± 0.82
1.54 ± 0.54
12.12 ± 0.92
ReadAgent ( 2024 )
60.11 ± 3.67
33.77 ± 3.13
18.49 ± 2.27
10.70 ± 1.40
24.49 ± 1.25
MemAgent ( 2026 )
42.13 ± 3.52
35.00 ± 3.08
0.64 ± 0.45
0.95 ± 0.42
13.63 ± 0.96
HippoRAG2 ( 2025 )
46.70 ± 3.55
13.33 ± 2.19
0.33 ± 0.33
0.38 ± 0.27
10.06 ± 0.85
LinearRAG ( 2026 )
32.49 ± 3.34
3.75 ± 1.23
0.64 ± 0.45
0.38 ± 0.27
6.04 ± 0.67
StructRAG ( 2025b )
44.16 ± 3.54
30.83 ± 2.98
24.76 ± 2.45
12.19 ± 1.43
23.72 ± 1.19
Appendix
Table 11: Accuracy (% ↑ ) across Loong’s reasoning facets. We show standard errors as ± values.
Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning differences underlying this gap remain underexplored. Across benchmarks in mathematics, physics, chemistry, and programming, we observe stable performance gaps: averaged over datasets, Qwen3-32B outperforms Qwen3-8B by 6.43%, while GPT-OSS-120B exceeds GPT-OSS-20B by 7.38%. To study the reasoning differences behind these gains, we develop AdvCluster, an automated framework that identifies questions where the larger model shows a stable advantage, extracts fine-grained advantage descriptions from paired reasoning traces produced by larger and smaller models, and organizes them through semantic clustering with quantitative evaluation and selection guided by a reviewer model. Our analysis yields a systematic taxonomy of larger model reasoning advantages, spanning both common advantages that recur across domains and specialized advantages associated with particular domains. Across these patterns, a recurring theme is Constraint-Guided Reasoning: larger models are better at identifying explicit and implicit constraints, organizing them into structured reasoning, and using them to rule out infeasible paths and verify intermediate steps.
Guan-Yi Lin, Hen-Hsen Huang
National Chengchi University Taipei, Taiwan · Academia Sinica Taipei, Taiwan
Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured Thoughts, a framework that organizes reasoning into alternating <try> and <outcome> blocks: <try> captures exploratory scratch work, while <outcome> contains the distilled conclusion of that step. We construct a dataset of structured thoughts by segmenting reasoning traces into <try> blocks and prompting an LLM to summarize each step into its corresponding <outcome>. Fine-tuning pretrained foundation models on this reformatted data produces models that adopt the structured reasoning style, leading to performance gains of up to 8.08% on reasoning benchmarks compared to standard SFT. The explicit structure also enables context pruning: after each <try>/<outcome> pair, the <try> can be pruned, allowing the model to retain conclusions without keeping the full scratch work in the context. A proof-of-concept pruning implementation achieves an average of 85% memory / context savings with an 8.67% performance drop across mathematical tasks.
Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents. In all these applications, memory (the ability to retain, access, and reason over information accumulated over long contexts and multiple interactions) plays a crucial role in determining the reliability of any agent. We introduce RECON (Reasoning over Extended Contexts with Obfuscated Narratives), a benchmark for evaluating compositional reasoning over long contexts. RECON spans 24 case files across three domains (criminal, medical, and financial), each ranging from 50k to 100k tokens, and tests agents on six memory intensive tasks: reconstructing multi-hop evidence chains, propagating cascading invalidations, resolving source conflicts, counterfactual reasoning, satisfying temporal constraints, and temporal fact retrieval. Recent memory benchmarks evaluate whether agents can retrieve scattered facts or detect if a fact has changed whereas RECON evaluates what happens after the change, whether agents can trace which downstream conclusions are affected, which survive through independent support, and how alternative timelines would have unfolded. Our evaluation reveals substantial limitations across current architectures: even the strongest non-Oracle system reaches only 22.4% Accuracy, with retrieval and reasoning each surfacing as challenges.
Mihir Shriniwas Arya
Department of Computer Science and Engineering RV College of Engineering