We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator can be swapped without retraining. T-Search is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic search tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks with gold evidence annotations, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts, outperforming larger open models. We release the model, harness, live demo, and three benchmarks, including TRuST, the first native-Russian hard-search benchmark.
Figures & tables
Figure 1: T-Search and its playground. The policy searches a selected backend, carries explicit memory between rounds, and returns ranked evidence. Optional parallel rollouts are fused with RRF. The playground displays the search trace, while a separate downstream model generates the answer. Fixed-index evaluation uses corpus backends; live web search is available in the demo.
Figure 2: A run in the playground on BrowseComp-Plus question #1217 over the Qwen3-Embedding-8B ( Zhang et al., 2025 ) index. Top: the question and the benchmark picker, with gold and evidence documents marked for badging. Below, the trace around a round boundary: a search after the 75% lock is rejected and flagged with the rule it broke; the save closes round one with eight chunks, the sources shown badged as gold documents; round two opens with a fresh context and its first query with the sources it returned.
BrowseComp-Plus
SealQA
SynthComp
Model
Harness
En
Ru
En
Ru
En
Ru
TRuST
Avg.
T-Search ( N=3 )
T-Search
72.65
62.93
66.08
61.98
58.52
58.00
49.12
61.33
T-Search ( N=1 )
T-Search
65.35
55.95
61.16
57.72
54.52
53.13
43.92
55.96
GLM-5.1
T-Search
64.32
58.18
55.49
53.21
51.69
51.71
43.11
53.96
GLM-5.2 1
T-Search
63.01
52.54
55.30
54.69
52.29
49.37
37.07
52.04
Kimi-K2.6 2
T-Search
60.71
49.76
56.86
52.46
48.25
47.06
42.39
51.07
Table 1: Retrieval results (%). Top and middle blocks report final-set Recall@10; the bottom block reports trajectory recall. Bold marks the highest final-set Recall@10 in each column. N=3 denotes three rollouts fused with RRF. ∗ WebShaper uses the WebDancer inference recipe.
Figure 3: Test-time scaling. Each curve sweeps the round cap r∈{1,3,5} for one model in the T-Search harness, left to right; the T-Search N=3 curve additionally fuses three parallel rollouts with RRF. Latency is per-query wall-clock time of the complete agent loop on 35 questions; recall is the mean final-set Recall@10 across benchmarks.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Synthetic task construction and revision. A taxonomy specification and paraphrased corpus facts guide incremental composition. Original support chunks remain available for validation. The composer model also acts as the search agent in the difficulty check. Actionable feedback returns to the composer for revision and rechecking; retained tasks provide questions, reference answers, annotated support, and fixed retrieval indices.
Figure 5: The training recipe. English and Russian are trained separately at each stage: per-language SFT is combined with DARE ( Yu et al., 2024 ) , per-language GSPO reinforcement learning with SLERP ( Shoemake, 1985 ; Goddard et al., 2024 ) .
Figure 6: RL training diagnostics for the two GSPO runs (English, Russian), trained separately before the SLERP merge. The reward panel plots the finalization recall of Eq. 1 , the only training signal; it rises over training, as does held-out accuracy; the number of search calls per episode grows; the gradient norm decays and the KL to the reference grows only modestly under the 10−3 penalty; and the rollout/train perplexity ratio converges toward 1, indicating that the vLLM rollout and Megatron training passes stay consistent without router replay. Held-out accuracy is best-of-4 answer accuracy on training-distribution questions held out from both SFT and RL, evaluated in the fixed round-based configuration; it is not a split of any evaluation benchmark.
BrowseComp-Plus
SealQA
SynthComp
Backend
En
Ru
En
Ru
En
Ru
TRuST
Avg.
Qwen3-8B + LLM reranking
75.04
66.24
64.82
59.95
62.71
62.39
48.93
62.87
Qwen3-8B (default)
65.35
55.95
61.16
57.72
54.52
53.13
43.92
55.96
Jina v5 text-small
60.52
51.41
62.46
55.60
56.37
54.56
39.31
54.32
BM25
39.49
31.97
55.33
50.00
66.87
65.15
49.18
51.14
Qwen3-0.6B
51.70
43.71
56.80
49.53
54.72
52.27
36.95
49.38
Appendix
Table 2: Reported Recall@10 (%) under backend substitution, with the same T-Search agent and N=1 . Qwen3-8B and Qwen3-0.6B denote Qwen3-Embedding models; Jina denotes jina-embeddings-v5-text-small-retrieval ( Akram et al., 2026 ) .
Benchmark
Base
SFT
+RL
BrowseComp-Plus (En)
43.7
54.0
65.4
BrowseComp-Plus (Ru)
38.6
47.7
56.0
SealQA (En)
46.1
51.5
61.2
SealQA (Ru)
43.3
50.8
57.7
SynthComp-En
41.8
50.0
54.5
SynthComp-Ru
43.9
48.9
53.1
Appendix
Table 3: Final-set Recall@10 (%). Base is Qwen3.6-35B-A3B, SFT is the model after supervised fine-tuning, and +RL is the released model after SFT followed by reinforcement learning. Avg. is the unweighted mean across the benchmark rows. Scores and averages retain the one-decimal precision of the published stage comparison. The Base and +RL columns report the same results as the base model and T-Search ( N=1 ) in Table 1 , rounded to one decimal place.
search_corpus
Semantic search over the corpus. Returns up to top_k items as {chunk_id, snippet, score}. Parallel calls allowed — issue multiple search_corpus tool calls in one assistant turn to cover independent angles. Dedup: within the current round, chunks already shown this round are filtered. Chunks in your saved set are filtered (already in context at top). Chunks seen in PREVIOUS rounds but not saved CAN resurface — re-consider them under the new angle.
Field
Type
Required
Description
query
string
Yes
Search query string.
top_k
integer
No
Number of results to return. Default: 5.
Appendix
Table 4: Released schema for search_corpus . Tool type: function ; parameters: object .
save_and_advance
End the current round and start a new one with fresh context. Only saved chunks survive the transition. Requires ≥MIN_SEARCHES_BEFORE_SAVE searches in the current round and ≥ 1 chunk saved. DISABLED in the last round. chunk_ids: saved come from this round’s search results or the currently-saved set (re-passing updates the reason); drop from the current saved set; no overlap between saved_chunks and drop_from_saved. You MUST declare progress via covered_concepts and unresolved_concepts — at least one of the two must be non-empty, and they must not overlap.
Field
Type
Required
Description
saved_chunks
array of object
Yes
Chunks to append to / update in the saved set.
saved_chunks[].chunk_id
string
Yes
saved_chunks[].reason
string
Yes
≥MIN_REASON_LEN chars. Specific fact or entity this chunk contributes.
drop_from_saved
array of string
No
Optional. chunk_ids to remove from the saved set. Use [] if dropping nothing.
Appendix
Table 5: Released schema for save_and_advance . Tool type: function ; parameters: object .
finalize_ranking
Submit final ranked list of most relevant documents for the user query and end the session. Empty ranking rejected. chunk_ids must come from your current context (saved + this round’s seen). You MUST declare covered_concepts and unresolved_concepts — harness will reject finalize if unresolved ≥ 50% of (covered + unresolved), except in the last round (where save_and_advance is disabled). Retry allowed on reject.
Field
Type
Required
Description
ranking
array of object
Yes
Ordered most → least relevant.
ranking[].chunk_id
string
Yes
ranking[].reason
string
Yes
≥MIN_REASON_LEN chars. Specific fact or entity this chunk contributes.
covered_concepts
array of string
Yes
Query atoms supported by evidence in the final ranking. Short labels.
Appendix
Table 6: Released schema for finalize_ranking . Tool type: function ; parameters: object .