Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.
Figures & tables
Figure 1: Overview of RECAST. At each round, RouterLM uses the task, compact source profile, and evidence and feedback from earlier rounds to either select and formulate a primitive or synthesized operation or pass the evidence to AnswerLM when it considers the evidence sufficient.
Figure 2: RouterLM action space and output structure. Together, the outputs represent action selection, task-specific formulation, and justification.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
6.7
1.0
23.0
41.7
69.7
1.0
23.8
Fixed Retrieval
–
36.0
27.3
72.3
79.3
64.0
55.3
55.7
Direct program generation
Direct Code
Qwen3.5-9B
80.7
1.0
26.0
64.0
54.7
3.3
38.3
Direct Code
Gemini 3.5 Flash
88.0
5.0
46.0
78.3
48.0
5.0
45.1
Table 1: In-domain success rates (%) of RECAST across six benchmark families. The listed LM directly generates the program for Direct Code, guides evidence retrieval for IRCoT and Interact-RAG, and serves as RouterLM for RECAST. Bold denotes the best result in each column.
Variant
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Evidence-operation space
RECAST Train-Free, Primitive Only
25.3
61.0
63.7
80.7
67.7
42.7
56.8
RECAST Train-Free, Synthesis Only
85.3
45.0
54.3
88.7
63.0
23.0
59.9
RECAST Train-Free, Full
86.0
57.0
60.7
83.0
68.3
29.3
64.1
GRPO reward and data selection
RECAST SFT+GRPO, Accuracy Only
82.0
58.0
65.0
87.3
66.0
55.0
68.9
Table 2: Success rates (%) of RECAST ablations across six benchmark families. The variants evaluate the evidence-operation space, training stages, GRPO reward, and GRPO data selection. Bold denotes the best result in each column.
Method
2Wiki
TAT-QA
WTQ
Avg.
Direct
72.0
1.0
10.0
27.7
Direct Code
45.0
38.0
59.0
47.3
Fixed Retrieval
58.0
66.0
52.0
58.7
IRCoT
69.0
73.0
51.0
64.3
Interact-RAG
77.0
57.0
48.0
60.7
RECAST SFT+GRPO
83.0
81.0
74.0
79.3
Table 3: Zero-shot success rate (%) of RECAST across three unseen benchmarks.
CompilerLM
In-domain avg.
Held-out avg.
Qwen3.5-9B
67.2
73.3
Gemini 3.5 Flash Lite
73.2
79.3
Gemini 3.5 Flash
75.6
79.3
Table 4: RECAST in-domain and held-out success rates (%) with alternative CompilerLM s.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Questions
Raw traj.
SFT rows
GRPO rows
DataBench
885
2,080
665
97
FinQA
410
1,109
349
101
HiTab
935
2,381
567
96
HotpotQA
785
1,674
660
96
LaMP
735
1,814
519
95
MultiHiertt
910
2,505
608
123
Appendix
Table 5: Composition of the trajectory collection and the final SFT and GRPO training sets.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
1.2
0.0
1.0
3.8
1.5
0.0
0.4
Fixed Retrieval
–
1.0
1.2
0.6
1.5
1.7
1.5
0.3
Direct program generation
Direct Code
Qwen3.5-9B
0.6
1.0
1.0
0.0
1.5
0.6
0.5
Direct Code
Gemini 3.5 Flash
1.7
2.0
1.7
0.6
1.0
0.0
0.4
Appendix
Table 6: Standard deviations of in-domain success rates (%) across three evaluations. The corresponding mean results are reported in Table 1 .
Benchmark
Primitive calls
Synthesis calls
Synthesis usage (%)
Routing rounds
DataBench
0.04
1.12
98.0
2.16
FinQA
3.00
0.22
21.0
4.22
HiTab
2.60
0.07
6.0
3.67
HotpotQA
1.46
0.00
0.0
2.43
LaMP
2.89
0.04
4.0
3.93
MultiHiertt
2.11
0.40
35.0
3.50
Appendix
Table 7: Operation usage and routing rounds of trained RECAST. Statistics are computed from one evaluation run, with 100 questions per benchmark. Calls and rounds are averaged per question; synthesis usage denotes the percentage of questions with at least one synthesis call.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
7.0
1.0
21.0
39.0
73.0
1.0
23.7
Fixed Retrieval
–
35.0
26.0
68.0
79.0
68.0
57.0
55.5
Direct program generation
Direct Code
Qwen3.5-9B
70.0
3.0
21.0
62.0
57.0
0.0
35.5
Direct Code
Gemini 3.5 Flash
77.0
7.0
41.0
78.0
47.0
5.0
42.5
Appendix
Table 8: In-domain success rates (%) judged by GPT-4.1. Same runs as Table 1 (one replicate per row), re-judged with gpt-4.1-2025-04-14 using the identical judge prompt. Bold denotes the best result in each column.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
7.0
1.0
19.0
38.0
68.0
1.0
22.3
Fixed Retrieval
–
36.0
26.0
66.0
77.0
62.0
56.0
53.8
Direct program generation
Direct Code
Qwen3.5-9B
79.0
2.0
23.0
64.0
56.0
2.0
37.7
Direct Code
Gemini 3.5 Flash
87.0
6.0
42.0
78.0
45.0
5.0
43.8
Appendix
Table 9: In-domain success rates (%) judged by DeepSeek-V4-Flash. Same runs as Table 1 (one replicate per row), re-judged with deepseek-ai/DeepSeek-V4-Flash-0731 using the identical judge prompt and decoding at temperature 0. Bold denotes the best result in each column.
Method
LM
Gemini
GPT-4.1
Agree (%)
κ
Direct
–
23.7
23.7
98.7
0.96
Fixed Retrieval
–
56.0
55.5
96.8
0.94
Direct Code
Qwen3.5-9B
37.7
35.5
95.2
0.90
Direct Code
Gemini 3.5 Flash
44.7
42.5
96.5
0.93
IRCoT
Qwen3.5-9B
57.0
56.7
97.0
0.94
IRCoT
Gemini 3.5 Flash
58.2
57.0
95.8
0.91
Appendix
Table 10: Agreement between the Gemini judge and GPT-4.1. Average success rate (%) under each judge on the same 600 predictions per row, raw agreement, and Cohen’s κ .
Method
LM
Gemini
DeepSeek-V4-Flash
Agree (%)
κ
Direct
–
23.7
22.3
98.0
0.94
Fixed Retrieval
–
56.0
53.8
96.5
0.93
Direct Code
Qwen3.5-9B
37.7
37.7
97.0
0.94
Direct Code
Gemini 3.5 Flash
44.7
43.8
98.8
0.98
IRCoT
Qwen3.5-9B
57.0
54.8
97.2
0.94
IRCoT
Gemini 3.5 Flash
58.2
55.5
95.3
0.90
Appendix
Table 11: Agreement between the Gemini judge and DeepSeek-V4-Flash. Average success rate (%) under each judge on the same 600 predictions per row, raw agreement, and Cohen’s κ .
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
6.7
0.7
8.7
37.5
53.9
0.0
17.9
Fixed Retrieval
–
38.6
11.8
47.1
72.6
56.8
32.9
43.3
Direct program generation
Direct Code
Qwen3.5-9B
69.5
0.0
20.3
61.2
40.0
0.0
31.8
Direct Code
Gemini 3.5 Flash
60.9
1.7
31.2
73.3
43.0
1.7
35.3
Appendix
Table 12: In-domain token-level F1 (%) across six benchmark families. Computed on the same runs as Table 8 with SQuAD-style normalization. Bold denotes the best result in each column.
Method
LM
Success (%)
Qwen (k)
Gemini (k)
Total (k)
Retrieval baselines
Direct
–
23.8
0.0
0.6
0.6
Fixed Retrieval
–
55.7
0.0
1.6
1.6
Direct program generation
Direct Code
Qwen3.5-9B
38.3
4.7
0.0
4.7
Direct Code
Gemini 3.5 Flash
45.1
0.0
4.8
4.8
Appendix
Table 13: Average task success rates (%) and inference token usage per question, in thousands (k). Token counts include input and output tokens across all inference calls and exclude the evaluation judge. The LM column follows Table 1 . Totals are computed before rounding.
Large language models (LLMs) are widely used in retrieval-augmented generation (RAG) to incorporate external knowledge at inference time. However, when retrieved contexts are noisy, incomplete, or heterogeneous, a single generation process often struggles to reconcile evidence effectively. We propose \textbf{MASS-RAG}, a multi-agent synthesis approach to retrieval-augmented generation that structures evidence processing into multiple role-specialized agents. MASS-RAG applies distinct agents for evidence summarization, evidence extraction, and reasoning over retrieved documents, and combines their outputs through a dedicated synthesis stage to produce the final answer. This design exposes multiple intermediate evidence views, allowing the model to compare and integrate complementary information before answer generation. Experiments on four benchmarks show that MASS-RAG consistently improves performance over strong RAG baselines, particularly in settings where relevant evidence is distributed across retrieved contexts.
Xingchen Xiao, Heyan Huang, Runheng Liu +1
School of Computer Science and Technology, Beijing Institute of Technology · Department of Mathematical Sciences, Tsinghua University
Large reasoning models such as DeepSeek-R1 and OpenAI o1 generate extended chains of thought spanning thousands of tokens, yet their integration with retrieval-augmented generation (RAG) remains fundamentally misaligned. Current RAG systems optimize for providing context before reasoning begins, while reasoning models require evidence injection during multi-step inference chains. We introduce ReaLM-Retrieve, a reasoning-aware retrieval framework that addresses this mismatch through three key innovations: (1) a step-level uncertainty detector that identifies knowledge gaps at reasoning-step granularity rather than token or sentence level; (2) a retrieval intervention policy that learns when external evidence maximally benefits ongoing reasoning; and (3) an efficiency-optimized integration mechanism that reduces per-retrieval overhead by 3.2x compared to naive integration. Experiments on MuSiQue, HotpotQA, and 2WikiMultiHopQA demonstrate that ReaLM-Retrieve achieves on average 10.1% absolute improvement in answer F1 over standard RAG (range: 9.0-11.8% across the three benchmarks) while reducing retrieval calls by 47% compared to fixed-interval approaches like IRCoT (all improvements significant at p<0.01, paired bootstrap). On the challenging MuSiQue benchmark requiring 2-4 hop reasoning, our method achieves 71.2% F1 with an average of only 1.8 retrieval calls per question. Analysis shows that ReaLM-Retrieve also improves retrieval quality itself, achieving 81.3% Recall@5 with consistently higher precision and MRR than fixed-interval baselines on supporting evidence, establishing new state-of-the-art efficiency-accuracy trade-offs for reasoning-intensive retrieval tasks.
Dongxin Guo, Jikun Wu, Siu Ming Yiu
The University of Hong Kong Hong Kong, China · Brain Investing Limited Hong Kong, China · Stellaris AI Limited Hong Kong, China
Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the benefit of shorter prompts. We revisit RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views. We first show that modern compressors have unstable gains over simple truncation and can add substantial inference-time latency. We then propose Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agnostic score store. REVA maps token-level attention to readable word units, aggregates importance across repeated document accesses, and renders budget-specific plain-text views that preserve document order and the standard RAG interface. Across four representative benchmarks and modern LLMs, REVA improves generation quality by 1.0--5.8 points over existing advances, while reducing compression overhead by a factor of 5.3 to 15.6, adding less than 40 ms of latency.
Tuan Nguyen, Qiran Hu, Banruo Liu +3
VinUni-Illinois Smart Health Center, VinUniversity, Vietnam · University of Illinois Urbana-Champaign, USA