Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.
Figures & tables
Figure 1: Overview of RECAST. At each round, RouterLM uses the task, compact source profile, and evidence and feedback from earlier rounds to either select and formulate a primitive or synthesized operation or pass the evidence to AnswerLM when it considers the evidence sufficient.
Figure 2: RouterLM action space and output structure. Together, the outputs represent action selection, task-specific formulation, and justification.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
6.7
1.0
23.0
41.7
69.7
1.0
23.8
Fixed Retrieval
–
36.0
27.3
72.3
79.3
64.0
55.3
55.7
Direct program generation
Direct Code
Qwen3.5-9B
80.7
1.0
26.0
64.0
54.7
3.3
38.3
Direct Code
Gemini 3.5 Flash
88.0
5.0
46.0
78.3
48.0
5.0
45.1
Table 1: In-domain success rates (%) of RECAST across six benchmark families. The listed LM directly generates the program for Direct Code, guides evidence retrieval for IRCoT and Interact-RAG, and serves as RouterLM for RECAST. Bold denotes the best result in each column.
Variant
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Evidence-operation space
RECAST Train-Free, Primitive Only
25.3
61.0
63.7
80.7
67.7
42.7
56.8
RECAST Train-Free, Synthesis Only
85.3
45.0
54.3
88.7
63.0
23.0
59.9
RECAST Train-Free, Full
86.0
57.0
60.7
83.0
68.3
29.3
64.1
GRPO reward and data selection
RECAST SFT+GRPO, Accuracy Only
82.0
58.0
65.0
87.3
66.0
55.0
68.9
Table 2: Success rates (%) of RECAST ablations across six benchmark families. The variants evaluate the evidence-operation space, training stages, GRPO reward, and GRPO data selection. Bold denotes the best result in each column.
Method
2Wiki
TAT-QA
WTQ
Avg.
Direct
72.0
1.0
10.0
27.7
Direct Code
45.0
38.0
59.0
47.3
Fixed Retrieval
58.0
66.0
52.0
58.7
IRCoT
69.0
73.0
51.0
64.3
Interact-RAG
77.0
57.0
48.0
60.7
RECAST SFT+GRPO
83.0
81.0
74.0
79.3
Table 3: Zero-shot success rate (%) of RECAST across three unseen benchmarks.
CompilerLM
In-domain avg.
Held-out avg.
Qwen3.5-9B
67.2
73.3
Gemini 3.5 Flash Lite
73.2
79.3
Gemini 3.5 Flash
75.6
79.3
Table 4: RECAST in-domain and held-out success rates (%) with alternative CompilerLM s.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Questions
Raw traj.
SFT rows
GRPO rows
DataBench
885
2,080
665
97
FinQA
410
1,109
349
101
HiTab
935
2,381
567
96
HotpotQA
785
1,674
660
96
LaMP
735
1,814
519
95
MultiHiertt
910
2,505
608
123
Appendix
Table 5: Composition of the trajectory collection and the final SFT and GRPO training sets.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
1.2
0.0
1.0
3.8
1.5
0.0
0.4
Fixed Retrieval
–
1.0
1.2
0.6
1.5
1.7
1.5
0.3
Direct program generation
Direct Code
Qwen3.5-9B
0.6
1.0
1.0
0.0
1.5
0.6
0.5
Direct Code
Gemini 3.5 Flash
1.7
2.0
1.7
0.6
1.0
0.0
0.4
Appendix
Table 6: Standard deviations of in-domain success rates (%) across three evaluations. The corresponding mean results are reported in Table 1 .
Benchmark
Primitive calls
Synthesis calls
Synthesis usage (%)
Routing rounds
DataBench
0.04
1.12
98.0
2.16
FinQA
3.00
0.22
21.0
4.22
HiTab
2.60
0.07
6.0
3.67
HotpotQA
1.46
0.00
0.0
2.43
LaMP
2.89
0.04
4.0
3.93
MultiHiertt
2.11
0.40
35.0
3.50
Appendix
Table 7: Operation usage and routing rounds of trained RECAST. Statistics are computed from one evaluation run, with 100 questions per benchmark. Calls and rounds are averaged per question; synthesis usage denotes the percentage of questions with at least one synthesis call.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
7.0
1.0
21.0
39.0
73.0
1.0
23.7
Fixed Retrieval
–
35.0
26.0
68.0
79.0
68.0
57.0
55.5
Direct program generation
Direct Code
Qwen3.5-9B
70.0
3.0
21.0
62.0
57.0
0.0
35.5
Direct Code
Gemini 3.5 Flash
77.0
7.0
41.0
78.0
47.0
5.0
42.5
Appendix
Table 8: In-domain success rates (%) judged by GPT-4.1. Same runs as Table 1 (one replicate per row), re-judged with gpt-4.1-2025-04-14 using the identical judge prompt. Bold denotes the best result in each column.
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
7.0
1.0
19.0
38.0
68.0
1.0
22.3
Fixed Retrieval
–
36.0
26.0
66.0
77.0
62.0
56.0
53.8
Direct program generation
Direct Code
Qwen3.5-9B
79.0
2.0
23.0
64.0
56.0
2.0
37.7
Direct Code
Gemini 3.5 Flash
87.0
6.0
42.0
78.0
45.0
5.0
43.8
Appendix
Table 9: In-domain success rates (%) judged by DeepSeek-V4-Flash. Same runs as Table 1 (one replicate per row), re-judged with deepseek-ai/DeepSeek-V4-Flash-0731 using the identical judge prompt and decoding at temperature 0. Bold denotes the best result in each column.
Method
LM
Gemini
GPT-4.1
Agree (%)
κ
Direct
–
23.7
23.7
98.7
0.96
Fixed Retrieval
–
56.0
55.5
96.8
0.94
Direct Code
Qwen3.5-9B
37.7
35.5
95.2
0.90
Direct Code
Gemini 3.5 Flash
44.7
42.5
96.5
0.93
IRCoT
Qwen3.5-9B
57.0
56.7
97.0
0.94
IRCoT
Gemini 3.5 Flash
58.2
57.0
95.8
0.91
Appendix
Table 10: Agreement between the Gemini judge and GPT-4.1. Average success rate (%) under each judge on the same 600 predictions per row, raw agreement, and Cohen’s κ .
Method
LM
Gemini
DeepSeek-V4-Flash
Agree (%)
κ
Direct
–
23.7
22.3
98.0
0.94
Fixed Retrieval
–
56.0
53.8
96.5
0.93
Direct Code
Qwen3.5-9B
37.7
37.7
97.0
0.94
Direct Code
Gemini 3.5 Flash
44.7
43.8
98.8
0.98
IRCoT
Qwen3.5-9B
57.0
54.8
97.2
0.94
IRCoT
Gemini 3.5 Flash
58.2
55.5
95.3
0.90
Appendix
Table 11: Agreement between the Gemini judge and DeepSeek-V4-Flash. Average success rate (%) under each judge on the same 600 predictions per row, raw agreement, and Cohen’s κ .
Method
LM
DataBench
FinQA
HiTab
HotpotQA
LaMP
MultiHiertt
Avg.
Retrieval baselines
Direct
–
6.7
0.7
8.7
37.5
53.9
0.0
17.9
Fixed Retrieval
–
38.6
11.8
47.1
72.6
56.8
32.9
43.3
Direct program generation
Direct Code
Qwen3.5-9B
69.5
0.0
20.3
61.2
40.0
0.0
31.8
Direct Code
Gemini 3.5 Flash
60.9
1.7
31.2
73.3
43.0
1.7
35.3
Appendix
Table 12: In-domain token-level F1 (%) across six benchmark families. Computed on the same runs as Table 8 with SQuAD-style normalization. Bold denotes the best result in each column.
Method
LM
Success (%)
Qwen (k)
Gemini (k)
Total (k)
Retrieval baselines
Direct
–
23.8
0.0
0.6
0.6
Fixed Retrieval
–
55.7
0.0
1.6
1.6
Direct program generation
Direct Code
Qwen3.5-9B
38.3
4.7
0.0
4.7
Direct Code
Gemini 3.5 Flash
45.1
0.0
4.8
4.8
Appendix
Table 13: Average task success rates (%) and inference token usage per question, in thousands (k). Token counts include input and output tokens across all inference calls and exclude the evaluation judge. The LM column follows Table 1 . Totals are computed before rounding.