Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-internal mechanism can select evidence across this range. We introduce UNifying REtrieval And Long-Context with a Single Model (UNREAL), a model-native evidence selection framework to span corpus retrieval and long-context inference. UNREAL encodes chunks and derives retrieval queries directly from the frozen LLM's internal representations. It adds fewer than 500K trainable parameters and leaves the backbone unchanged. On a 3B-token, 21M-chunk Wikipedia index, all four dense and hybrid UNREAL backbones outperform state-of-the-art retriever-reranker systems. The best model raises recall from 49.1% to 73.2% on HotpotQA, from 31.7% to 60.1% on 2WikiMultiHopQA, and from 8.8% to 14.4% on MuSiQue. Applied to long-context tasks, the same selection mechanism removes distractors before generation, raising NoLiMa accuracy from 1.0% to 24.83% at its maximum context length of 128K tokens, and LV-Eval's F1 score from 49.97% to 54.66% at 256K. UNREAL also reduces FLOPs and time-to-first-token relative to full-context inference from roughly 32K tokens onward, with larger gains as context grows. Together, these results establish model-internal evidence selection as a common foundation for corpus retrieval and evidence-sparse long-context inference.
Figures & tables
Figure 1: Exact match on question-answering as context length grows from 8K to 100M tokens. Full context evaluation is shown only until computationally infeasible. For retrieval-based methods, relevant context is first extracted from the input and then fed to the model to generate the answer. Unlike baselines that rely on external models, UNREAL leverages the frozen LLM’s internal representations to retrieve relevant context, maintaining the highest accuracy across long context lengths.
Figure 2: UNREAL trains retrieval tokens ρ and summation weights α (fire indicates training), while keeping the LLM frozen, enhancing its ability to identify relevant information in a context. Given a long context, these trained variables enable UNREAL to select relevant chunks and pass them to the same LLM for retrieval-augmented generation. Numbers indicate the order of operations.
Figure 3: UNREAL scales to full-corpus Wikipedia retrieval. Complete-evidence recall@10 on 21M Wiki-2018 chunks across eight datasets; an example is correct only when all annotated evidence chunks are retrieved. Across four architectures, UNREAL consistently outperforms baselines, especially on multi-hop datasets. Bars show means with 95% BCa confidence intervals.
HotpotQA
2Wiki
MuSiQue
Method
EM
F1
EM
F1
EM
F1
No context
28.5
37.7
29.9
34.6
6.8
16.6
Oracle
64.3
78.9
57.5
66.8
47.5
58.2
BM25
39.1
49.7
29.1
35.5
6.4
14.6
BGE-large-en-v1.5
39.6
51.2
28.5
35.6
8.6
17.6
Qwen3-Emb-0.6B
37.9
48.5
28.2
35.5
10.1
18.1
Table 1: Exact match (EM) and macro-averaged F1 on three multi-hop QA datasets, generated by Nemotron-3.5-Lightning under each retrieval method’s top-5 retrieved context. UNREAL achieves the highest EM and F1 scores across all three datasets while relying on a single LLM for both retrieval and generation.
Figure 4: Left: NoLiMa accuracy across context lengths (4K–128K tokens) using Nemotron-3-Nano. Right: LV-Eval F1 across context lengths (16K–256K words, averaged over LooGLE-SD and MultiFieldQA-en) using Nemotron-3.5-Lightning. In both cases, full-context performance degrades as context grows, whereas UNREAL achieves the highest performance.
Figure 5: Substring exact-match on the RAG subset of the HELMET long context benchmark, with Nemotron-3.5-Lightning-30B-A3B and Qwen3.5-35B-A3B as generators. UNREAL outperforms full-context inference and other retrieval methods for both generators.
Figure 6: Recall and accuracy on the NoLiMa benchmark for Nemotron-3-Nano. Left: The effect of the number of selected chunks n on UNREAL’s accuracy and recall. Right: Embeddings from the frozen LLM outperform random retrieval, motivating UNREAL, which is trained to enhance this capability.
Figure 7: Time-to-first-token speedup of UNREAL over full-context inference. Solid curves use directly measured full-context runtimes; dashed curves use FLOP-scaled projections beyond the last measured context length.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Complete-evidence recall@20 on 21M Wiki-2018 chunks across 8 datasets; an example is correct only when all annotated evidence chunks are retrieved. Bars show means with 95% BCa confidence intervals.
Figure 9: Average recall@10 on 21M Wiki-2018 chunks across 8 datasets. Bars show means with 95% BCa confidence intervals.
Figure 10: Average recall@20 on 21M Wiki-2018 chunks across 8 datasets. Bars show means with 95% BCa confidence intervals.
HotpotQA
2Wiki
MuSiQue
Method
EM
F1
EM
F1
EM
F1
No context
25.0
34.6
27.7
33.3
6.9
14.8
Oracle
60.8
74.9
53.6
61.7
42.6
54.6
BM25
36.8
47.3
28.1
34.4
7.0
15.2
BGE-large-en-v1.5
36.9
47.3
27.6
33.7
9.7
18.1
Qwen3-Emb-0.6B
34.9
46.0
28.5
35.0
7.7
15.4
Appendix
Table 2: Exact match (EM) and macro-averaged F1 on three QA datasets, generated by Qwen3.5-35B-A3B under each retrieval method’s top-5 retrieved context. “Oracle” denotes the upper-bound result, obtained by answering the question based only on the oracle chunks.
Figure 11: Distractor identity matters. With the gold chunk always present, retrieved distractors cause more interference than the same number of random distractors, and the gap grows with the budget.
Figure 12: LOFT generation. (a) All selective budgets outperform full-context reading. (b) Performance by dataset at n=10 against gold-only context. We use each dataset’s designated LOFT metric.
HotpotQA
MuSiQue
SQuAD v2
Design choice
Avg R@10
All R@10
Avg R@10
All R@10
Avg R@10
Baseline
81.28
72.93
32.70
13.79
54.43
R=64→16
77.93 (-3.35)
67.83 (-5.10)
31.91 (-0.79)
13.04 (-0.75)
53.82 (-0.61)
G=4→1
77.91 (-3.37)
68.37 (-4.56)
29.69 (-3.01)
11.49 (-2.30)
52.94 (-1.49)
ML → SL
73.02 (-8.26)
61.01 (-11.92)
26.20 (-6.50)
9.86 (-3.93)
46.93 (-7.50)
HN =500→100
79.40 (-1.88)
70.59 (-2.34)
28.77 (-3.93)
11.58 (-2.21)
52.86 (-1.57)
Appendix
Table 3: Recipe ablation for UNREAL-Nemotron-3.5-Lightning-30B-A3B, single-knob departures from the main method. R is the number of retrieval tokens; G is the number of groups they are mean-pooled into on the query side; ML (multi-layer) vs. SL (single-layer) is whether the read-out combines all 6 global-attention blocks or only the block used for the corpus readout; HN is the number of hard negatives per query; ∣C0(x)∣ is the number of BM25-retrieved context chunks prepended to the query; P is the number of learnable soft prompt tokens used at the beginning of the sequence. Avg. R@10 is recall averaged over each query’s oracle chunks; All R@10 requires every oracle chunk to be retrieved.
HotpotQA
MuSiQue
SQuAD v2
Corpus layer
Avg R@10
All R@10
Avg R@10
All R@10
Avg R@10
Layer 8
79.68
70.25
31.74
12.33
53.66
Layer 18
78.98
69.42
31.65
12.45
54.05
Layer 30
80.19
71.39
32.40
13.33
54.85
Layer 38
80.62
71.99
31.94
12.29
54.01
Layer 42
81.28
72.93
32.70
13.79
54.43
Appendix
Table 4: Ablation on the layer used to create the corpus chunk embeddings, for UNREAL-Nemotron-3.5-Lightning-30B-A3B.
HotpotQA
MuSiQue
2Wiki
SQuAD v2
Lp
Avg R@10
All R@10
Avg R@10
All R@10
Avg R@10
All R@10
Avg R@10
1
78.20
69.11
30.28
12.45
67.43
55.21
48.52
3
79.34
70.58
31.15
12.70
69.51
57.11
51.61
7
81.28
72.93
32.70
13.79
71.52
58.90
54.43
11
81.54
73.40
34.34
14.67
72.38
59.68
55.42
Appendix
Table 5: Effect of the number of pooled vectors per corpus chunk Lp on retrieval recall@10, for UNREAL-Nemotron-3.5-Lightning-30B-A3B. Avg. R@10 is the recall averaged over each query’s oracle chunks; All R@10 requires every oracle chunk of a query to be retrieved. Recall increases monotonically with Lp , but at the cost of higher memory and compute during evaluation.
N∗ (tokens)
Backbone
Θ
La
κ
ℓc/L
λ
n=5
n=10
Muse-Glimmer-30B
25.2B
13
4.73×105
47/52
0.096
13,049
17,437
Nemotron-3.5-Lightning
0 2.87B
0 6
1.17×105
42/52
0.201
0 6,321
0 8,471
Qwen3.5-35B-A3B
0 2.44B
10
5.95×104
27/40
0.323
0 4,117
0 5,569
Appendix
Table 6: Approximate break-even context length N∗ (tokens; equation 10 ), with T=141 , Tq=Tg=T , and ∣xret∣=910 . Θ counts active non-embedding matrix parameters.
Figure 13: Measured time-to-first-token, baseline vs. UNREAL, for all four backbones (random weights/inputs, served with vLLM on an NVIDIA H100, bf16). Solid markers are directly measured; the dashed baseline segment is a FLOPs-scaled projection past the point where a single full-context forward pass exceeds our measurement budget or the available KV cache. Dotted vertical connectors mark the speedup (baseline time / UNREAL time) at a few representative context lengths, read directly off the time gap between the two curves.
Figure 14: Measured prefill throughput (tokens/second), baseline vs. UNREAL, for all four backbones. Same data and conventions as Fig. 13 .
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization. In this work, we propose Recursive Evidence Replay as LLM Harness for Long-Context Reasoning (RECONTEXT), a training-free inference method for improving long-context reasoning. RECONTEXT uses model-internal relevance signals to construct a query-conditioned evidence pool and replays it before final generation while preserving the full original context. This recursive selection process separates evidence organization from answer generation without training, external memory, or context pruning. We also provide a theoretical analysis based on associative memory, which characterizes the context as a memory store, the question as a retrieval cue, attention as cue-trace association, and replay as trace reactivation. Experiments on eight long-context datasets with 128K context length show that RECONTEXT consistently improves evidence utilization across Qwen3-4B, Qwen3-8B, and Llama3-8B, achieving the best average rank on all three backbones. Code is available at https://github.com/Yanjun-Zhao/ReContext.
Retrieval augmented generation (RAG) depends critically on the quality and granularity of retrieved evidence. Large retrieval units preserve context but often introduce irrelevant content, which can dilute answer bearing evidence and worsen long context utilization. Fine-grained units are more compact, but they may be difficult to retrieve reliably because short chunks can lack semantic, lexical, or bridging cues needed to match the query. We propose Uncertainty-aware Multi-Granularity RAG (UMG-RAG), a training-free hybrid retrieval framework that treats chunk granularity as query-specific reliability estimation. Instead of training a new retriever or modifying the generator, UMG-RAG uses existing dense and sparse retrievers as complementary experts across multiple chunk granularities. For each query, it converts each expert-granularity score list into an evidence distribution, estimates reliability from distribution entropy, and fuses candidates according to query-specific semantic, lexical, and granularity confidence. We further introduce UMGP-RAG, a parent promotion variant that uses fine-grained hits to locate relevant evidence while returning broader non-redundant parent chunks for local coherence. Experiments on question answering benchmarks show that uncertainty-aware fusion and parent promotion improve generation quality while maintaining a lightweight, plug-and-play retrieval pipeline.
Hoin Jung, Xiaoqian Wang
Elmore Family School of Electrical and Computer Engineering Purdue University West Lafayette, IN 47907
Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with retrieval-augmented generation (RAG) remains fundamentally misaligned. Current RAG systems optimize for providing context before reasoning begins, while reasoning models require evidence injection during multi-step inference chains. We introduce ReaLM-Retrieve, a reasoning-aware retrieval framework that addresses this mismatch through three key innovations: (1) a step-level uncertainty detector that identifies knowledge gaps at reasoning-step granularity rather than token or sentence level; (2) a retrieval intervention policy that learns when external evidence maximally benefits ongoing reasoning; and (3) an efficiency-optimized integration mechanism that reduces per-retrieval overhead by 3.2x compared to naive integration. Experiments on MuSiQue, HotpotQA, and 2WikiMultiHopQA demonstrate that ReaLM-Retrieve achieves on average 10.1% absolute improvement in answer F1 over standard RAG (range: 9.0-11.8% across the three benchmarks) while reducing retrieval calls by 47% compared to fixed-interval approaches like IRCoT (all improvements significant at p<0.01, paired bootstrap). On the challenging MuSiQue benchmark requiring 2-4 hop reasoning, our method achieves 71.2% F1 with an average of only 1.8 retrieval calls per question. Analysis shows that ReaLM-Retrieve also improves retrieval quality itself, achieving 81.3% Recall@5 with consistently higher precision and MRR than fixed-interval baselines on supporting evidence, establishing new state-of-the-art efficiency-accuracy trade-offs for reasoning-intensive retrieval tasks.
Dongxin Guo, Jikun Wu, Siu Ming Yiu
The University of Hong Kong Hong Kong, China · Brain Investing Limited Hong Kong, China · Stellaris AI Limited Hong Kong, China