Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
Figures & tables
Figure 1: Search decision performance and end-to-end agentic search. (a) On SearchDecision-Bench, SearchJev achieves 5.2–5.3 × faster decisions than same-size Qwen3.5. The x-axis uses Qwen3.5-4B as the 1 × reference. (b) On BrowseComp-Plus, dual-system agents with SearchJev -0.8B/4B reach 46.0%/54.0% answer accuracy, compared with 45.0% for System-2-only and 46.0% for Jev 1.13. Relative to System-2-only, they achieve 3.7–4.7 × faster search with 3.3–4.8 × fewer output tokens.
Decision Type
Decision Task Output Type Task Example
Routing
Intent classification Choice Query: { query }. Which intent does it express? Options: { intent labels }. Retrieval need Score Query: { query }. How much information must be retrieved? Levels: none / one fact / several facts / multi-source research. Query clarity Choice Query: { query }. What information most needs to be added? Options: none / time / location / object / aspect / { other }.
Rewriting
Query equivalence Noul Queries: { query 1 } and { query 2 }. Do they express the same intent? Options: True / False. Rewrite equivalence Noul Dialogue: { history }. Utterance: { utterance }. Rewrite: { rewrite }. Does the rewrite keep the utterance’s intent? Options: True / False.
Relevance
Relevance judgment Noul / Score Query: { query }. Document: { document }. Is it relevant? Options: True / False, or { relevance grades }. Document usefulness Noul Sub-query: { current sub-query }. Document: { document }. Does the document answer the sub-query? Options: True / False. Document quality Score Query: { query }. Document: { site, date, content }. How authoritative is the source? Levels: untrusted / user-generated / professional / authoritative.
Sufficiency
Evidence sufficiency Noul Question: { original question }. Evidence: { retrieved evidence }. Does the evidence cover all facts needed to answer? Options: True / False.
Navigation
Next-action selection Choice Search state: { state }. What should the agent do next? Options: search / read page / answer. Next sub-query selection Choice Question: { original question }. Solved steps: { history }. Which sub-query comes next? Options: { four candidate queries }. Query-fit assessment Noul Next unsolved step: { step }. Candidate query: { query }. Does the query target this step? Options: True / False.
Verification
Claim verification Choice Claim: { claim }. Evidence: { evidence }. Is the claim supported? Options: supports / refutes / insufficient. Evidence detection Noul Claim: { claim }. Sentence: { sentence }. Is the sentence evidence for the claim? Options: True / False.
Table 1: Representative decision tasks for agentic search in SearchDecision-Bench. The illustrative examples span six decision types and three output types: Choice (categorical selection), Score (ordinal scoring), and Noul (binary judgment). Braces denote placeholders.
Decision Type
Training
ID Test
OOD Test
Routing
77,627
26,427
10,000
Rewriting
74,318
32,746
10,000
Relevance
153,990
17,581
5,976
Sufficiency
14,652
1,036
0
Navigation
56,981
10,213
0
Verification
47,830
10,334
20,631
Table 2: Statistics of SearchDecision-Bench. Counts denote state-schema pairs.
Figure 2: SearchJev decision model. A search state, the question and output type of task f , and its legal options form the decision input xf . Choice uses letter labels; Score and Noul show example values. The model normalizes option-label logits into a decision distribution without autoregressive output generation.
Figure 3: Architecture of the Dual-System Search Agent. System 2 handles planning, query generation, and final answer synthesis, while SearchJev (System 1) makes search decisions specified by the schema at each step. Confident decisions are accepted; uncertain ones are delegated to System 2. The search controller uses the resolved decisions to select actions and update the search state with retrieved evidence.
Relevance
Sufficiency
Routing
Navigation
Rewriting
Verification
Latency
Model
NDCG ↑
ECE ↓
Acc ↑
ECE ↓
Acc ↑
ECE ↓
Acc ↑
ECE ↓
Acc ↑
ECE ↓
Acc ↑
ECE ↓
ms ↓
Jev 1.13 (closed)
98.5
7.6
79.8
5.6
73.5
5.5
72.1
1.5
79.5
2.8
79.8
1.6
265.4
Qwen3.5-0.8B AR (JSON)
85.7
15.7
49.9
24.2
24.7
24.0
42.0
19.7
40.6
32.1
56.8
9.5
144.2
Qwen3.5-4B AR (JSON)
96.7
4.2
57.7
15.0
53.3
6.7
37.8
26.3
66.1
5.5
55.1
4.7
195.2
SearchJev -0.8B
98.6
5.8
89.4
5.1
83.2
5.0
56.8
4.6
84.0
2.3
91.7
9.3
27.6
SearchJev -4B
99.0
6.5
96.0
5.6
86.6
5.7
57.2
8.4
85.9
2.0
93.2
8.8
37.1
Table 3: Main results on the in-distribution test sets of SearchDecision-Bench, covering decision quality, calibration, and latency. NDCG@10, accuracy, and ECE are scaled by 100; latency is the median per-decision time in milliseconds at batch size 1. Bold highlights the best performance, underlined indicates the second-best.
Agent
Acc ↑
System-2 Calls ↓
Output Tokens (k) ↓
Time (s) ↓
System-2 only (Qwen3.5-27B)
45.0
90.3
73.7
4080
Dual-system Search Agent (Jev 1.13)
46.0
76.0
18.3
1018
Dual-system Search Agent ( SearchJev -0.8B)
46.0
62.1
15.5
868
Dual-system Search Agent ( SearchJev -4B)
54.0
84.6
22.1
1102
Table 4: End-to-end agentic search results on BrowseComp-Plus, covering answer quality and efficiency with System 2 fixed at 27B. Dual-system agent sizes refer to SearchJev ; the Jev 1.13 agent uses Jev in its place. Acc is answer accuracy (%). Calls and output tokens are for System 2 and are reported per question; k denotes thousands. Time is the median active time per question, including System-2 calls, decision-model inference, and retrieval. Bold highlights the best performance, underlined indicates the second-best.
Routing
Rewriting
Relevance
Verification
Latency
Model
Acc ↑
ECE ↓
Acc ↑
ECE ↓
NDCG ↑
ECE ↓
Acc ↑
ECE ↓
ms ↓
Jev 1.13 (closed)
62.3
20.0
97.7
5.6
98.7
3.4
84.5
3.7
273.3
Qwen3.5-0.8B AR (JSON)
29.2
13.4
47.3
38.4
82.2
32.0
53.9
22.6
153.2
Qwen3.5-4B AR (JSON)
43.7
30.3
95.8
24.1
93.8
10.6
78.0
3.7
195.8
SearchJev -0.8B
41.1
15.5
99.6
28.1
95.5
31.4
55.7
23.6
27.7
SearchJev -4B
55.8
3.0
99.6
26.9
99.0
11.4
65.0
15.4
37.1
Table 5: Results on the out-of-distribution test sets of SearchDecision-Bench, covering decision quality, calibration, and latency across four decision families. NDCG@10, accuracy, and ECE are scaled by 100; latency is the median per-decision time in milliseconds at batch size 1. NDCG@10 is computed over the candidates of each FRAMES question. Bold highlights the best performance, underlined indicates the second-best. Ties receive the same marking.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
0.8B
4B
Tuning
LoRA
LoRA
LoRA rank / α
16 / 32
16 / 32
Learning rate
5×10−5
5×10−5
Batch size
132
132
Steps (1 epoch)
5,176
5,176
Max input length
2,048
2,048
Appendix
Table 6: Training hyperparameters of SearchJev . All models use AdamW ( Loshchilov and Hutter, 2019 ) (weight decay 0.01, gradient clipping 1.0), a cosine schedule ( Loshchilov and Hutter, 2017 ) with 5% warm-up, and bf16 precision, and minimize the cross-entropy to the soft labels plus λB times the Brier score; LoRA ( Hu et al., 2022 ) covers the attention, Gated DeltaNet ( Yang et al., 2025b ) , and MLP projections. wf applies to LLM-assigned labels, demonstrations, and constructed rewrites (1 otherwise). T∗ : temperature fitted per output type on 5,000 validation decisions.
AR (JSON)
SearchJev
Family
Dataset
n
0.8B
4B
0.8B
4B
Relevance
DuReader-Retrieval
642
71.8
92.7
93.6
96.6
QBQTC
3,681
34.8
30.0
70.5
72.8
Query-document quality
10,000
29.2
56.6
77.4
83.1
T2Ranking (reranking)
3,370
54.2
72.1
81.2
82.9
T2Ranking (retrieval)
5,868
71.5
88.1
91.2
91.7
Appendix
Table 7: Accuracy (%) on each in-distribution test set; n : test decisions. Table 3 instead uses NDCG@10 for the aggregate relevance result, so its relevance value is not an average of the relevance rows reported here.
Agent
System-2 decision calls (per question) ↓
Reasoning tokens (thousands per question) ↓
System-2 time (seconds per question) ↓
Decision time (seconds per search step) ↓
System-2 only (27B)
60.4
69.9
4079
95.7
+ SearchJev -0.8B
0
11.9
693
1.8
+ SearchJev -4B
0
16.1
759
3.0
Appendix
Table 8: Detailed resource use on BrowseComp-Plus, complementing the end-to-end results in Table 4 . System-2 decision calls are explicit calls dedicated to short decisions and are a subset of the total System-2 calls in Table 4 ; reasoning tokens are a subset of System-2 output tokens. System-2 time and decision time are medians. Bold highlights the best performance, underlined indicates the second-best.
Routing
Rewriting
Relevance
Sufficiency
Navigation
Verification
General JEV benchmarks
System
When2Call
B77
CLINC
MASSIVE
Typed
PAWS
MS MARCO
SQuAD2
OJ-OOD
MNLI
BoolQ
RAGTruth
JB-Hard
JB-All
Kev-T4
SemIf
Bev
Qwen3.5-4B backbone
JevK5 v0.3
73.2
65.2
70.0
73.8
–
–
–
–
–
–
–
71.2
78.4
87.9
–
–
66.3
Plumb-4B
–
–
–
–
–
–
–
–
–
–
–
–
80.2
89.6
–
–
–
decider-4b v2
–
–
–
–
–
–
–
–
–
–
–
–
67.6
84.0
–
–
–
Decision 4B v1.2
–
–
–
–
–
–
–
–
–
–
–
–
78.4
88.3
–
–
–
Appendix
Table 9: Comparison with System-1 peers from the JevBench leaderboard on the datasets they report, with the closed Jev 1.13 API included as an external reference. Results are grouped by decision family. Accuracy × 100; RAGTruth: F1 of the hallucinated class. The best, second-best, and third-best results on each dataset are in bold, underlined, and italic; ties share a style. Peer results are self-reported, except for Jev 1.13, which we measure through its API on the same items as SearchJev .
License
Datasets and models
Apache-2.0
T2Ranking ( Xie et al., 2023 ) , DuReader-Retrieval ( Qiu et al., 2022 ) and CBLUE ( Zhang et al., 2022 ) (repository licenses), CFEVER ( Lin et al., 2024 ) , FRAMES ( Krishna et al., 2025 ) , XYZ-Aquila SFT, CValues ( Xu et al., 2023 ) (repository license; data for research use), typed-decisions; Qwen3.5-0.8B/4B/27B ( Qwen Team, 2026 ) and Qwen3-32B ( Yang et al., 2025a )
MIT
RAGTruth ( Niu et al., 2024 ) , TrendFact ( Zhang et al., 2026 ) , BrowseComp-Plus ( Chen et al., 2026 ) , JevBench, Chinese Multi-Emotion Dialogue, SemIf-OpenJev (code)
CC BY 4.0
MuSiQue ( Trivedi et al., 2022 ) , BANKING77 ( Casanueva et al., 2020 ) , MASSIVE ( FitzGerald et al., 2023 ) , HelpSteer2 ( Wang et al., 2024 ) , HelpSteer3 ( Wang et al., 2025a ) , When2Call ( Ross et al., 2025 ) ; Search Arena ( Miroyan et al., 2026 ) prompts (model outputs follow the providers’ terms)
CC BY 3.0
CLINC150 ( Larson et al., 2019 )
CC BY-SA 3.0
QReCC ( Anantha et al., 2021 ) , VitaminC ( Schuster et al., 2021 ) (code: MIT), BoolQ ( Clark et al., 2019 )
CC BY-SA 4.0
SQuAD 2.0 ( Rajpurkar et al., 2018 )
Appendix
Table 10: Licenses declared by the official releases of the datasets and models we use.
We present T-Search, an open-weight agentic retriever for hard multi-step search. Given a question and a search tool over a fixed corpus, it runs a bounded multi-round search and returns a ranked list of evidence chunks with short justifications, leaving answer generation to a downstream model, so backend and generator can be swapped without retraining. T-Search is built on Qwen3.6-35B-A3B and trained on adversarially filtered synthetic search tasks with round-sliced supervised fine-tuning followed by GSPO on a recall reward. Averaged over seven English and Russian benchmarks with gold evidence annotations, it reaches 56.0 Recall@10 with one rollout, 14.4 points above its base, and 61.3 with three fused rollouts, outperforming larger open models. We release the model, harness, live demo, and three benchmarks, including TRuST, the first native-Russian hard-search benchmark.
Olga Tsymboi, Ramil Latypov, Aleksandr Medvedev +5
Probability-only models, which TypeSafe calls System One models, return calibrated probabilities for fixed choices in milliseconds and generate no text. We study one such model, Jev, through two tasks that require decisions under tight constraints. In bullet chess, a bot that places Jev's judgment inside Stockfish search alongside an opening book and endgame tablebases climbs above a 2200 Lichess bullet rating against other bots. Live model calls are too slow for search, so we distill pairwise judgments into a compact evaluator that runs at every position. We then ask how best to spend a fixed labeling budget when an LLM, Qwen3-32B, is available as a second teacher. In chess, averaging both judges' labels beats spending the whole budget on Qwen alone by 9.6 Elo (95% interval 4.3 to 14.9), and the gain replicates on fresh openings; a second answer from the same judge is no substitute, and Jev is the strongest partner for Qwen among the models tested. In passage reranking, Jev's labels alone train a reranker that scores as high as Qwen's, from 21 minutes of API calls instead of 5.1 GPU-hours, and adding Qwen gains at most a few thousandths in ranking quality. Search supplies the lookahead, distillation makes the judgment cheap enough to use at every position, and an LLM partner pays off in chess.
Mohamad Yazan Sadoun, Sarah Sharif, Yaser Mike Banad
INQUIRE Lab, School of Electrical and Computer Engineering University of Oklahoma, Norman, OK, USA
LLM agents often use generative models for bounded decisions, raising the question of when these decisions can be handled more efficiently without reducing task success. We study REFLEX, an agent architecture that uses Jev as a fast, typed decision layer and calls a strong LLM when confidence is low, or generation is required. On a frozen 100-task benchmark, REFLEX achieves 95% success with 72.7% fewer strong-model calls than a strong-only agent, with reductions persisting across three fallback families. Controlled interventions show that reliability depends on action-set size and near-valid alternatives near authorization boundaries. External BFCL and τ-style evaluations reveal limited advantages over a cheap generative cascade when ordinary routing is already highly accurate. These findings identify when selective control with Jev can reduce computation and where its benefits are limited.