T-Search: An Open Agentic Retriever and Playground for Hard Multi-Step Search
Organizations: T-Tech
Abstract
We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator can be swapped without retraining. T-Search is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic search tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks with gold evidence annotations, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts, outperforming larger open models. We release the model, harness, live demo, and three benchmarks, including TRuST, the first native-Russian hard-search benchmark.
Figures & tables
| BrowseComp-Plus | SealQA | SynthComp | |||||||
| Model | Harness | En | Ru | En | Ru | En | Ru | TRuST | Avg. |
| T-Search ( ) | T-Search | 72.65 | 62.93 | 66.08 | 61.98 | 58.52 | 58.00 | 49.12 | 61.33 |
| T-Search ( ) | T-Search | 65.35 | 55.95 | 61.16 | 57.72 | 54.52 | 53.13 | 43.92 | 55.96 |
| GLM-5.1 | T-Search | 64.32 | 58.18 | 55.49 | 53.21 | 51.69 | 51.71 | 43.11 | 53.96 |
| GLM-5.2 1 | T-Search | 63.01 | 52.54 | 55.30 | 54.69 | 52.29 | 49.37 | 37.07 | 52.04 |
| Kimi-K2.6 2 | T-Search | 60.71 | 49.76 | 56.86 | 52.46 | 48.25 | 47.06 | 42.39 | 51.07 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| BrowseComp-Plus | SealQA | SynthComp | ||||||
|---|---|---|---|---|---|---|---|---|
| Backend | En | Ru | En | Ru | En | Ru | TRuST | Avg. |
| Qwen3-8B + LLM reranking | 75.04 | 66.24 | 64.82 | 59.95 | 62.71 | 62.39 | 48.93 | 62.87 |
| Qwen3-8B (default) | 65.35 | 55.95 | 61.16 | 57.72 | 54.52 | 53.13 | 43.92 | 55.96 |
| Jina v5 text-small | 60.52 | 51.41 | 62.46 | 55.60 | 56.37 | 54.56 | 39.31 | 54.32 |
| BM25 | 39.49 | 31.97 | 55.33 | 50.00 | 66.87 | 65.15 | 49.18 | 51.14 |
| Qwen3-0.6B | 51.70 | 43.71 | 56.80 | 49.53 | 54.72 | 52.27 | 36.95 | 49.38 |
| Benchmark | Base | SFT | +RL |
|---|---|---|---|
| BrowseComp-Plus (En) | 43.7 | 54.0 | 65.4 |
| BrowseComp-Plus (Ru) | 38.6 | 47.7 | 56.0 |
| SealQA (En) | 46.1 | 51.5 | 61.2 |
| SealQA (Ru) | 43.3 | 50.8 | 57.7 |
| SynthComp-En | 41.8 | 50.0 | 54.5 |
| SynthComp-Ru | 43.9 | 48.9 | 53.1 |
| search_corpus | |||
| Semantic search over the corpus. Returns up to top_k items as {chunk_id, snippet, score}. Parallel calls allowed — issue multiple search_corpus tool calls in one assistant turn to cover independent angles. Dedup: within the current round, chunks already shown this round are filtered. Chunks in your saved set are filtered (already in context at top). Chunks seen in PREVIOUS rounds but not saved CAN resurface — re-consider them under the new angle. | |||
| Field | Type | Required | Description |
| query | string | Yes | Search query string. |
| top_k | integer | No | Number of results to return. Default: 5. |
| save_and_advance | |||
| End the current round and start a new one with fresh context. Only saved chunks survive the transition. Requires MIN_SEARCHES_BEFORE_SAVE searches in the current round and 1 chunk saved. DISABLED in the last round. chunk_ids: saved come from this round’s search results or the currently-saved set (re-passing updates the reason); drop from the current saved set; no overlap between saved_chunks and drop_from_saved. You MUST declare progress via covered_concepts and unresolved_concepts — at least one of the two must be non-empty, and they must not overlap. | |||
| Field | Type | Required | Description |
| saved_chunks | array of object | Yes | Chunks to append to / update in the saved set. |
| saved_chunks[].chunk_id | string | Yes | |
| saved_chunks[].reason | string | Yes | MIN_REASON_LEN chars. Specific fact or entity this chunk contributes. |
| drop_from_saved | array of string | No | Optional. chunk_ids to remove from the saved set. Use [] if dropping nothing. |
| finalize_ranking | |||
| Submit final ranked list of most relevant documents for the user query and end the session. Empty ranking rejected. chunk_ids must come from your current context (saved + this round’s seen). You MUST declare covered_concepts and unresolved_concepts — harness will reject finalize if unresolved 50% of (covered + unresolved), except in the last round (where save_and_advance is disabled). Retry allowed on reject. | |||
| Field | Type | Required | Description |
| ranking | array of object | Yes | Ordered most least relevant. |
| ranking[].chunk_id | string | Yes | |
| ranking[].reason | string | Yes | MIN_REASON_LEN chars. Specific fact or entity this chunk contributes. |
| covered_concepts | array of string | Yes | Query atoms supported by evidence in the final ranking. Short labels. |