TRACE: Trajectory Selection for Parallel Scaling of Search Agents
Organizations: New York University Shanghai · New York University
Abstract
Parallel search may generate a correct answer that final-answer voting fails to select. We formulate this consolidation stage as trajectory selection and introduce TRACE (Trajectory Ranking with Aggregated Cross-Rollout Evidence), a lightweight learned selector that ranks completed trajectories using the search evidence behind their answers. TRACE preserves individual query and evidence occurrences, connects rollouts through shared content or document identity, and propagates information across these relations. Each candidate answer then reads the updated states of its own trajectory, preserving retrieval provenance while incorporating evidence from related rollouts. Trained with answer-level supervision over frozen text embeddings, TRACE returns an existing answer without additional search or autoregressive aggregation. One selector per search setting transfers across rollout policies and agent backbones without agent-specific fine-tuning, improving over voting across six WebQA policies and six long-horizon dataset-backbone combinations at . On Qwen2.5-14B Base/SFT WebQA pools, TRACE achieves 45.2/49.2% EM, compared with 43.9/48.0% for the strongest Qwen3-32B generative aggregators. On long-horizon FRAMES, GAIA, and BrowseComp, it reaches 78.6% average accuracy, exceeding majority voting by 3.1 percentage points. On Base WebQA pools, TRACE with only 8 rollouts comes within 0.4 points of majority voting over 64. TRACE also achieves at least higher processing throughput than SolAgg, SummAgg, and AggAgent across all seven WebQA benchmarks. These results show that reusing cross-rollout search evidence provides an effective and efficient alternative to heavyweight generative aggregation for parallel search. Code is available at https://github.com/Jaasssoooonnnnn/TRACE.
Figures & tables
| NQ | HotpotQA | TriviaQA | PopQA | 2Wiki | MuSiQue | Bamboogle | Overall | |||||||||
| Method | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT |
| Single rollout | 29.4 | 41.4 | 27.6 | 36.8 | 49.6 | 64.2 | 35.2 | 38.6 | 24.6 | 31.4 | 10.4 | 18.4 | 31.2 | 45.6 | 29.5 | 38.8 |
| Majority Voting | 44.6 | 46.6 | 40.4 | 46.8 | 63.8 | 72.2 | 46.0 | 46.2 | 40.8 | 46.2 | 18.4 | 22.4 | 48.8 | 59.2 | 42.6 | 47.2 |
| Weighted Voting | 44.4 | 46.6 | 40.6 | 47.2 | 64.8 | 72.6 | 46.0 | 46.2 | 41.6 | 45.6 | 19.0 | 22.4 | 48.8 | 60.0 | 43.0 | 47.3 |
| Fewest Tools | 36.0 | 45.0 | 30.2 | 42.6 | 58.2 | 71.2 | 39.4 | 43.4 | 26.0 | 41.0 | 9.6 | 20.2 | 32.8 | 57.6 | 33.2 | 44.4 |
| SolAgg | 43.2 | 43.4 | 42.6 | 49.0 | 65.0 | 72.4 | 45.0 | 46.4 | 45.8 | 48.6 | 19.6 | 25.0 | 52.8 | 60.8 | 43.9 | 48.0 |
| BrowseComp-Plus | FRAMES | GAIA | |||||
| Method | OR | OSS | OR | OSS | OR | OSS | Avg. |
| Single rollout | 38.3 | 45.2 | 73.1 | 84.9 | 55.4 | 61.7 | 59.7 |
| Majority Voting | 64.5 | 66.0 | 88.4 | 90.4 | 66.0 | 77.7 | 75.5 |
| Weighted Voting | 64.5 | 70.2 | 88.4 | 90.6 | 67.0 | 79.6 | 76.7 |
| Fewest Tools | 64.6 | 72.7 | 86.0 | 89.0 | 51.5 | 76.7 | 73.4 |
| SolAgg | 67.0 | 71.4 | 88.6 | 89.0 | 66.0 | 77.7 | 76.6 |
| Configuration | WebQA | Long-horizon avg. |
|---|---|---|
| Full TRACE | 45.2 | 78.6 |
| w/o cross-rollout communication | 43.9 | 72.7 |
| w/o GNN | 43.7 | 75.5 |
| Fixed-query readout | 44.7 | 75.6 |
| Method | GPU-hours | Relative cost |
|---|---|---|
| TRACE | 0.271 | |
| SolAgg | 3.28 | |
| SummAgg | 32.96 | |
| AggAgent | 69.93 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Matching rule | Groups | Occurrence pairs | Graphs (%) | ||
|---|---|---|---|---|---|---|
| Mean | Mean | Median | P90 | |||
| WebQA | Identical chunk | 7.6 | 243.6 | 228 | 430 | 99.9 |
| Browsing | Identical evidence | 15.4 | 462.0 | 314 | 1,080 | 99.0 |
| Browsing | Same document | 6.5 | 1,797.6 | 1,142.5 | 4,224 | 99.0 |
| Configuration | WebQA | Long-horizon avg. |
|---|---|---|
| Full TRACE | 45.2 | 78.6 |
| Occurrence representation | ||
| Merge identical evidence occurrences | 44.1 | 77.1 |
| Merge identical answers | 44.9 | 76.3 |
| Merge both | 44.4 | 76.4 |
| Answer readout and training | ||
| WebQA | Long-horizon avg. | ||||
|---|---|---|---|---|---|
| EM | F1 | TRACE | Majority | Pass@ | |
| 1 | 28.5 | 36.4 | 59.7 | 59.7 | 59.7 |
| 2 | 35.2 | 43.9 | 67.8 | 65.5 | 72.2 |
| 4 | 40.2 | 48.8 | 71.5 | 69.7 | 80.1 |
| 8 | 43.3 | 51.8 | 75.3 | 73.8 | 86.1 |
| 16 | 45.2 | 53.4 | 78.6 | 75.5 | 90.9 |
| Policy | 1 | 2 | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|---|---|
| Q2.5-7B Base | 24.9 | 35.3 | 44.8 | 52.6 | 58.6 | 63.4 | 67.4 |
| Q2.5-14B | 30.5 | 40.1 | 48.0 | 54.3 | 59.3 | 63.6 | 67.3 |
| SearchR1 RL | 43.7 | 46.5 | 48.8 | 50.6 | 52.1 | 53.3 | 54.2 |
| Q2.5-32B | 30.5 | 39.7 | 47.6 | 54.2 | 59.4 | 63.6 | 67.4 |
| Q3-32B | 36.1 | 43.8 | 50.0 | 54.7 | 58.5 | 61.7 | 64.2 |
| SFT (7B) | 36.3 | 43.4 | 49.4 | 54.4 | 58.5 | 62.1 | 65.3 |
| Qwen2.5-7B rollouts | Qwen2.5-14B rollouts | |||||
| Method | Base | SFT | RL | Base | SFT | RL |
| Single rollout | 24.6 | 36.6 | 44.0 | 29.5 | 38.8 | 46.9 |
| Majority Voting | 38.1 | 46.0 | 45.1 | 42.6 | 47.2 | 47.6 |
| Weighted Voting | 38.5 | 46.1 | 45.3 | 43.0 | 47.3 | 47.7 |
| Fewest Tools | 27.2 | 42.0 | 45.3 | 33.2 | 44.4 | 47.1 |
| SolAgg | 42.4 | 48.1 | 46.5 | 43.9 | 48.0 | 48.5 |
| NQ | HotpotQA | TriviaQA | PopQA | 2Wiki | MuSiQue | Bamboogle | Overall | |||||||||
| Method | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT | Base | SFT |
| Single rollout | 29.4 | 41.4 | 27.6 | 36.8 | 49.6 | 64.2 | 35.2 | 38.6 | 24.6 | 31.4 | 10.4 | 18.4 | 31.2 | 45.6 | 29.5 | 38.8 |
| Majority Voting | 44.6 | 46.6 | 40.4 | 46.8 | 63.8 | 72.2 | 46.0 | 46.2 | 40.8 | 46.2 | 18.4 | 22.4 | 48.8 | 59.2 | 42.6 | 47.2 |
| Weighted Voting | 44.4 | 46.6 | 40.6 | 47.2 | 64.8 | 72.6 | 46.0 | 46.2 | 41.6 | 45.6 | 19.0 | 22.4 | 48.8 | 60.0 | 43.0 | 47.3 |
| Fewest Tools | 36.0 | 45.0 | 30.2 | 42.6 | 58.2 | 71.2 | 39.4 | 43.4 | 26.0 | 41.0 | 9.6 | 20.2 | 32.8 | 57.6 | 33.2 | 44.4 |
| SolAgg | 43.2 | 43.4 | 42.6 | 49.0 | 65.0 | 72.4 | 45.0 | 46.4 | 45.8 | 48.6 | 19.6 | 25.0 | 52.8 | 60.8 | 43.9 | 48.0 |
| Stage | Seconds |
|---|---|
| Read and parse candidate texts | 23.6 |
| Fresh text encoding | 905.2 |
| Assemble tensors and construct graphs | 38.2 |
| Score and extract answers | 7.2 |
| Total | 974.2 |