SearchJev: A Fast and Calibrated System-1 Model for Search Agents
Organizations: Huawei Technologies Co., Ltd. · Leiden University
Abstract
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
Figures & tables
| Decision Type | Decision Task Output Type Task Example |
|---|---|
| Routing | Intent classification Choice Query: { query }. Which intent does it express? Options: { intent labels }. Retrieval need Score Query: { query }. How much information must be retrieved? Levels: none / one fact / several facts / multi-source research. Query clarity Choice Query: { query }. What information most needs to be added? Options: none / time / location / object / aspect / { other }. |
| Rewriting | Query equivalence Noul Queries: { query 1 } and { query 2 }. Do they express the same intent? Options: True / False. Rewrite equivalence Noul Dialogue: { history }. Utterance: { utterance }. Rewrite: { rewrite }. Does the rewrite keep the utterance’s intent? Options: True / False. |
| Relevance | Relevance judgment Noul / Score Query: { query }. Document: { document }. Is it relevant? Options: True / False, or { relevance grades }. Document usefulness Noul Sub-query: { current sub-query }. Document: { document }. Does the document answer the sub-query? Options: True / False. Document quality Score Query: { query }. Document: { site, date, content }. How authoritative is the source? Levels: untrusted / user-generated / professional / authoritative. |
| Sufficiency | Evidence sufficiency Noul Question: { original question }. Evidence: { retrieved evidence }. Does the evidence cover all facts needed to answer? Options: True / False. |
| Navigation | Next-action selection Choice Search state: { state }. What should the agent do next? Options: search / read page / answer. Next sub-query selection Choice Question: { original question }. Solved steps: { history }. Which sub-query comes next? Options: { four candidate queries }. Query-fit assessment Noul Next unsolved step: { step }. Candidate query: { query }. Does the query target this step? Options: True / False. |
| Verification | Claim verification Choice Claim: { claim }. Evidence: { evidence }. Is the claim supported? Options: supports / refutes / insufficient. Evidence detection Noul Claim: { claim }. Sentence: { sentence }. Is the sentence evidence for the claim? Options: True / False. |
| Decision Type | Training | ID Test | OOD Test |
|---|---|---|---|
| Routing | 77,627 | 26,427 | 10,000 |
| Rewriting | 74,318 | 32,746 | 10,000 |
| Relevance | 153,990 | 17,581 | 5,976 |
| Sufficiency | 14,652 | 1,036 | 0 |
| Navigation | 56,981 | 10,213 | 0 |
| Verification | 47,830 | 10,334 | 20,631 |
| Relevance | Sufficiency | Routing | Navigation | Rewriting | Verification | Latency | |||||||
| Model | NDCG | ECE | Acc | ECE | Acc | ECE | Acc | ECE | Acc | ECE | Acc | ECE | ms |
| Jev 1.13 (closed) | 98.5 | 7.6 | 79.8 | 5.6 | 73.5 | 5.5 | 72.1 | 1.5 | 79.5 | 2.8 | 79.8 | 1.6 | 265.4 |
| Qwen3.5-0.8B AR (JSON) | 85.7 | 15.7 | 49.9 | 24.2 | 24.7 | 24.0 | 42.0 | 19.7 | 40.6 | 32.1 | 56.8 | 9.5 | 144.2 |
| Qwen3.5-4B AR (JSON) | 96.7 | 4.2 | 57.7 | 15.0 | 53.3 | 6.7 | 37.8 | 26.3 | 66.1 | 5.5 | 55.1 | 4.7 | 195.2 |
| SearchJev -0.8B | 98.6 | 5.8 | 89.4 | 5.1 | 83.2 | 5.0 | 56.8 | 4.6 | 84.0 | 2.3 | 91.7 | 9.3 | 27.6 |
| SearchJev -4B | 99.0 | 6.5 | 96.0 | 5.6 | 86.6 | 5.7 | 57.2 | 8.4 | 85.9 | 2.0 | 93.2 | 8.8 | 37.1 |
| Agent | Acc | System-2 Calls | Output Tokens (k) | Time (s) |
|---|---|---|---|---|
| System-2 only (Qwen3.5-27B) | 45.0 | 90.3 | 73.7 | 4080 |
| Dual-system Search Agent (Jev 1.13) | 46.0 | 76.0 | 18.3 | 1018 |
| Dual-system Search Agent ( SearchJev -0.8B) | 46.0 | 62.1 | 15.5 | 868 |
| Dual-system Search Agent ( SearchJev -4B) | 54.0 | 84.6 | 22.1 | 1102 |
| Routing | Rewriting | Relevance | Verification | Latency | |||||
| Model | Acc | ECE | Acc | ECE | NDCG | ECE | Acc | ECE | ms |
| Jev 1.13 (closed) | 62.3 | 20.0 | 97.7 | 5.6 | 98.7 | 3.4 | 84.5 | 3.7 | 273.3 |
| Qwen3.5-0.8B AR (JSON) | 29.2 | 13.4 | 47.3 | 38.4 | 82.2 | 32.0 | 53.9 | 22.6 | 153.2 |
| Qwen3.5-4B AR (JSON) | 43.7 | 30.3 | 95.8 | 24.1 | 93.8 | 10.6 | 78.0 | 3.7 | 195.8 |
| SearchJev -0.8B | 41.1 | 15.5 | 99.6 | 28.1 | 95.5 | 31.4 | 55.7 | 23.6 | 27.7 |
| SearchJev -4B | 55.8 | 3.0 | 99.6 | 26.9 | 99.0 | 11.4 | 65.0 | 15.4 | 37.1 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| 0.8B | 4B | |
| Tuning | LoRA | LoRA |
| LoRA rank / | 16 / 32 | 16 / 32 |
| Learning rate | ||
| Batch size | 132 | 132 |
| Steps (1 epoch) | 5,176 | 5,176 |
| Max input length | 2,048 | 2,048 |
| AR (JSON) | SearchJev | |||||
| Family | Dataset | 0.8B | 4B | 0.8B | 4B | |
| Relevance | DuReader-Retrieval | 642 | 71.8 | 92.7 | 93.6 | 96.6 |
| QBQTC | 3,681 | 34.8 | 30.0 | 70.5 | 72.8 | |
| Query-document quality | 10,000 | 29.2 | 56.6 | 77.4 | 83.1 | |
| T2Ranking (reranking) | 3,370 | 54.2 | 72.1 | 81.2 | 82.9 | |
| T2Ranking (retrieval) | 5,868 | 71.5 | 88.1 | 91.2 | 91.7 | |
| Agent | System-2 decision calls (per question) | Reasoning tokens (thousands per question) | System-2 time (seconds per question) | Decision time (seconds per search step) |
|---|---|---|---|---|
| System-2 only (27B) | 60.4 | 69.9 | 4079 | 95.7 |
| + SearchJev -0.8B | 0 | 11.9 | 693 | 1.8 |
| + SearchJev -4B | 0 | 16.1 | 759 | 3.0 |
| Routing | Rewriting | Relevance | Sufficiency | Navigation | Verification | General JEV benchmarks | |||||||||||
| System | When2Call | B77 | CLINC | MASSIVE | Typed | PAWS | MS MARCO | SQuAD2 | OJ-OOD | MNLI | BoolQ | RAGTruth | JB-Hard | JB-All | Kev-T4 | SemIf | Bev |
| Qwen3.5-4B backbone | |||||||||||||||||
| JevK5 v0.3 | 73.2 | 65.2 | 70.0 | 73.8 | – | – | – | – | – | – | – | 71.2 | 78.4 | 87.9 | – | – | 66.3 |
| Plumb-4B | – | – | – | – | – | – | – | – | – | – | – | – | 80.2 | 89.6 | – | – | – |
| decider-4b v2 | – | – | – | – | – | – | – | – | – | – | – | – | 67.6 | 84.0 | – | – | – |
| Decision 4B v1.2 | – | – | – | – | – | – | – | – | – | – | – | – | 78.4 | 88.3 | – | – | – |
| License | Datasets and models |
|---|---|
| Apache-2.0 | T2Ranking ( Xie et al., 2023 ) , DuReader-Retrieval ( Qiu et al., 2022 ) and CBLUE ( Zhang et al., 2022 ) (repository licenses), CFEVER ( Lin et al., 2024 ) , FRAMES ( Krishna et al., 2025 ) , XYZ-Aquila SFT, CValues ( Xu et al., 2023 ) (repository license; data for research use), typed-decisions; Qwen3.5-0.8B/4B/27B ( Qwen Team, 2026 ) and Qwen3-32B ( Yang et al., 2025a ) |
| MIT | RAGTruth ( Niu et al., 2024 ) , TrendFact ( Zhang et al., 2026 ) , BrowseComp-Plus ( Chen et al., 2026 ) , JevBench, Chinese Multi-Emotion Dialogue, SemIf-OpenJev (code) |
| CC BY 4.0 | MuSiQue ( Trivedi et al., 2022 ) , BANKING77 ( Casanueva et al., 2020 ) , MASSIVE ( FitzGerald et al., 2023 ) , HelpSteer2 ( Wang et al., 2024 ) , HelpSteer3 ( Wang et al., 2025a ) , When2Call ( Ross et al., 2025 ) ; Search Arena ( Miroyan et al., 2026 ) prompts (model outputs follow the providers’ terms) |
| CC BY 3.0 | CLINC150 ( Larson et al., 2019 ) |
| CC BY-SA 3.0 | QReCC ( Anantha et al., 2021 ) , VitaminC ( Schuster et al., 2021 ) (code: MIT), BoolQ ( Clark et al., 2019 ) |
| CC BY-SA 4.0 | SQuAD 2.0 ( Rajpurkar et al., 2018 ) |