Solving complex tasks in domains such as science, medicine, law, and finance often requires assembling interdependent information scattered across vast, heterogeneous sources far beyond model context limits. Existing approaches tackle this challenge by organizing information into more manageable representations over which models can reason, such as graphs, textual memories, and retrieval collections. These representations dictate what downstream reasoning is possible and, ultimately, whether it succeeds; yet their design and construction remain largely ad hoc. In this work, drawing on the cognitive theory of relevance realization, we propose concrete principles for designing AI systems that construct effective representations of very large contexts. We analyze existing approaches and show how their successes and failures map onto their alignment with these principles, and introduce R3Con, a harness designed to operationalize the principles more systematically. We evaluate R3Con against nine state-of-the-art baselines on two recent benchmarks of reasoning over large document corpora. On these benchmarks, R3Con substantially outperforms the strongest baseline, by 20 and 8.4 percentage points. It also enables smaller models to outperform much larger ones: R3Con with 4B and 9B models outperforms all evaluated 35B baselines, while R3Con with a 35B-A3B model outperforms Claude Code with Claude-Sonnet-5 at 3.7× lower cost. Our results show that context representations following our principled approach can reduce reliance on model scale, pointing toward a future of AI systems with frontier-level performance powered by smaller models. Our code is available at https://github.com/michaeltheologitis/r3con
Figures & tables
Figure 1: R3Con with a 3B -active model outperforms Claude Code equipped with Anthropic’s frontier Sonnet models on CorpusQA ( Lu et al., 2026 ) , while costing roughly 3.7× to 3.9× less.
Figure 2: Running example based on Bram Stoker’s Dracula ( 1897 ) . Full details in § B .
Figure 3: An investigation of the events aboard the ship Demeter. Figuring out that the entire crew of nine died, and that Dracula was behind it, requires connecting evidence across partial and incomplete accounts.
Figure 4: R3Con answers a task over a large context by first constructing an effective representation using our design principles (§ 3.2 ). It extracts the relevant context R (Phase 1) and then uses R to further structure the context (Phase 2). Together, the extracted relevant context and structured context form the representation, which is then used by a coding agent to produce the answer (Phase 3).
Figure 5: Evolution of the relevance snippets for the ship’s log ( D37 ) and The Dailygraph newspaper ( D12 ) across two rounds of extracting relevance. For example, the top-left and top-right cells show R37(1) and R37(2) , respectively. ‡ Document ordering is arbitrary. The full snippets are shown in Fig. 12 .
Table 1: Accuracy (% ↑ ) across real-world domains requiring reasoning over large document corpora. R3Con achieves the highest accuracy across all domains, outperforming the strongest baseline by +19.88 pp on CorpusQA and +8.42 pp on Loong overall (SEs in Tab. 10 ). Accuracy is micro-averaged.
Figure 7
Variant
Loong ( ↑ )
CorpusQA ( ↑ )
R3Con (default)
37.43%
70.64%
[0.6pt/2pt] w/o extracting relevance
23.96% (-13.47)
52.58% (-18.1)
w/o structuring
32.97% (-4.46)
63.22% (-7.42)
Table 3: Accuracy in the ablation of the two phases of R3Con (see Fig. 4 ).
Table 4: Taxonomy of baselines according to the principles (§ 3.2 ) satisfied by their representation-construction phase. See § C.2 for detailed analysis. = limited support; N/A = not applicable.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Document type
# Docs.
Description
Journal / diary
5
Personal journals and diaries (e.g., Jonathan Harker’s journal)
Letter
20
Letters between characters (e.g., Mina to Lucy)
Telegram
8
Telegrams between characters (e.g., Van Helsing to Seward)
Memorandum
4
Written accounts of events (e.g., Lucy’s final memorandum)
Newspapers
4
Newspaper cuttings (e.g., The Pall Mall Gazette)
Note
2
Short written notes (e.g., Van Helsing’s contingency note)
Appendix
Table 5: The reconstructed document collection in Bram Stoker’s Dracula ( 1897 ) .
Model
Parametric knowledge
With corpus in context
Answer
Correct
Answer
Correct
Qwen3.5-35B-A3B ( 2026 )
31
✗
7
✗
Appendix
Table 6: Answers to the Dracula question using parametric knowledge alone and with the full novel provided in context. The ground-truth answer is 13, but also requires the correct people named. Full model responses are shown in Fig. 8 .
Model
Temp.
Top- p
Top- k
Min- p
Pres. pen.
Rep. pen.
Qwen3.5-4B
1.0
0.95
20
0.0
1.5
1.0
Qwen3.5-9B
1.0
0.95
20
0.0
1.5
1.0
Qwen3.5-35B-A3B
1.0
0.95
20
0.0
1.5
1.0
Appendix
Table 7: Sampling parameters used for each model, following the recommendations from the official HuggingFace model cards.
Method
CorpusQA
Loong
Qwen3.5-4B
StructRAG ( 2025b )
29.79%
16.84%
CodeAgent ( 2025 )
10.92%
19.44%
R3Con (ours)
49.24%
30.23%
Qwen3.5-9B
StructRAG ( 2025b )
40.43%
20.30%
Appendix
Table 8: Accuracy ( ↑ ) across benchmarks and model sizes of the strongest baselines.
Figure 9: Distribution of the representation components (i.e., final relevant context and structuring) for the two benchmarks with the default Qwen3.5-35B-A3B model.
Figure 12: Evolution of the relevance snippets for the log of the Demeter ( D37 ) and The Dailygraph newspaper ( D12 ) across two rounds of extracting relevance. For example, the top-left and top-right cells show R37(1) and R37(2) , respectively. ‡ Document ordering is arbitrary.
Downstream Reasoning
CorpusQA
Loong
LLM
65.43%
40.17%
Coding Agent (default)
70.64%
37.43%
Appendix
Table 9: Accuracy ( ↑ ) with different downstream reasoning methods over the representation constructed by R3Con .
Figure 13: Effect of the number of relevance-extraction rounds N on accuracy and cost.
Method
CorpusQA ( Lu et al., 2026 )
Loong ( Wang et al., 2024a )
Edu.
Finance
Real-estate
Overall
Paper
Finance
Legal
Overall
RAPTOR ( 2024 )
8.00 ± 3.13
11.11 ± 2.47
17.95 ± 4.35
12.06 ± 1.84
0.00 ± 0.00
22.17 ± 1.71
5.71 ± 1.24
12.12 ± 0.92
ReadAgent ( 2024 )
42.68 ± 5.46
37.42 ± 3.79
31.25 ± 5.18
37.23 ± 2.68
15.15 ± 2.08
35.33 ± 2.03
14.93 ± 1.95
24.49 ± 1.25
MemAgent ( 2026 )
7.32 ± 2.88
7.83 ± 2.09
8.64 ± 3.12
7.90 ± 1.49
0.00 ± 0.00
26.80 ± 1.81
4.00 ± 1.05
13.63 ± 0.96
HippoRAG2 ( 2025 )
7.32 ± 2.88
7.83 ± 2.09
20.99 ± 4.52
10.94 ± 1.72
0.32 ± 0.32
17.42 ± 1.55
6.29 ± 1.30
10.06 ± 0.85
LinearRAG ( 2026 )
8.54 ± 3.09
6.63 ± 1.93
16.05 ± 4.08
9.42 ± 1.61
0.61 ± 0.43
9.38 ± 1.19
5.43 ± 1.21
6.04 ± 0.67
Appendix
Table 10: Accuracy (% ↑ ) across real-world domains requiring reasoning over large document corpora. R3Con achieves the highest accuracy across all domains, outperforming the strongest baseline by +19.88 pp on CorpusQA and +8.42 pp on Loong overall. Overall accuracy is micro-averaged. We report standard errors (SE) as ± values.
Method
Spot.
Comp.
Chain
Clust.
Overall
RAPTOR ( 2024 )
51.31 ± 3.62
16.25 ± 2.38
2.04 ± 0.82
1.54 ± 0.54
12.12 ± 0.92
ReadAgent ( 2024 )
60.11 ± 3.67
33.77 ± 3.13
18.49 ± 2.27
10.70 ± 1.40
24.49 ± 1.25
MemAgent ( 2026 )
42.13 ± 3.52
35.00 ± 3.08
0.64 ± 0.45
0.95 ± 0.42
13.63 ± 0.96
HippoRAG2 ( 2025 )
46.70 ± 3.55
13.33 ± 2.19
0.33 ± 0.33
0.38 ± 0.27
10.06 ± 0.85
LinearRAG ( 2026 )
32.49 ± 3.34
3.75 ± 1.23
0.64 ± 0.45
0.38 ± 0.27
6.04 ± 0.67
StructRAG ( 2025b )
44.16 ± 3.54
30.83 ± 2.98
24.76 ± 2.45
12.19 ± 1.43
23.72 ± 1.19
Appendix
Table 11: Accuracy (% ↑ ) across Loong’s reasoning facets. We show standard errors as ± values.