Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.
Table 1 : Feedback mechanisms for web-agent RL and test-time scaling. Dense = step-level feedback beyond binary outcome; URL = URL-state grounding; Cert. = certified trust; J-free = judge-free test time; TTS = test-time scaling; Xfer = reusable bank transfer.
Figure 1 : CLIFT concept. Vanilla GRPO broadcasts one trajectory-scalar reward; CLIFT scores each visited URL with verification questions and inherits the URL score to actions taken there. Full pipeline: Figure 2 .
Figure 2 : CLIFT overview. Training: rollouts are scored by a comparative judge and by a self-verifier that queries a polarity-tagged question bank. The CCC certifies question weights by URL tier, yielding a signed verifier score that is asymmetrically added to the judge reward and then converted into URL-stratified token advantages. Test-time scaling: the trained agent emits a greedy rollout and diverse retries; the frozen certified bank produces per-URL evidence for each pairwise comparison, and a conservative repeated-vote rule decides whether to keep the incumbent or swap.
Table 3 : WAI leaderboard slice. Success rates on the public 9-app comparison set; values are percentages without percent signs. The CLIFT +CTS column reports the trained Gemma-4 backbone with K=3 Conformal Trajectory Selection (one selected trajectory per task). Overall is the task-weighted success rate over the 9 apps and matches the CLIFT +CTS row of Table 6 .
Figure 3 : WAI training reward. CLIFT separates from vanilla GRPO under the same judge and browser-use stack.
Method
C
S
R
Avg.
SGV ( Andrade et al., 2026 )
52.0
57.0
33.0
50.2
WALT ( Prabhu et al., 2025 )
64.1
53.4
39.0
52.9
Gemma-4 base
41.5
28.5
24.3
30.9
Gemma-4 + CLIFT
44.9
30.7
31.0
34.4
Gemma-4 + CLIFT+CTS (K=2)
45.3
32.6
31.4
35.6
GPT-5.5 + CLIFT
60.7
43.1
47.1
48.5
Table 4 : VWA success rates. C/S/R are Classifieds, Shopping, and Reddit; Avg is weighted by env size. For CTS rows, K denotes rollout budget and one selected trajectory is emitted per task.
Origin
Construction
∣Q∣
Certified
WAI bank
keep/rewrite
36
33 ( 91.7% )
VWA bank
keep/rewrite
48
31 ( 64.6% )
OM2W expansion
live-web checks
48
36 ( 75.0% )
Final bank
132 qs
132
100 (75.8%)
Table 5 : Transferred bank on OM2W. Questions are reused from WAI/VWA or expanded, then recalibrated on OM2W rollouts; the final certified bank has 75 positive-lift and 25 negative-lift questions with mean lift 0.34 .
Figure 5 : OM2W transfer gains. Realised SR; K is rollout budget.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Origin
Stage
Questions
Certified
Mean lift
WAI bank
kept as-is
7
7
0.46
WAI bank
rewritten
29
26
0.46
VWA bank
kept as-is
10
4
0.31
VWA bank
rewritten
38
27
0.28
OM2W expansion
live-web checks
48
36
0.30
Total
–
132
100
0.34
Appendix
Table 7 : OM2W bank construction and recalibration. Questions come from WAI/VWA banks plus an OM2W-specific failure-mode expansion.
System
K
Canonical SR
Gain
Sonnet SR
Gemma-4 base
1
40.0
–
46.8
Gemma-4 + ours
2
45.0
+5.0
partial positive
GPT-5.5 low base
1
43.7
–
47.3
GPT-5.5 high base
1
49.7
–
52.0
GPT-5.5 high + ours
2
57.3
+7.7
55.7
GPT-5.5 high + ours
4
61.0
+11.3
61.7
Appendix
Table 8 : OM2W transfer performance details. K is the rollout budget and every row emits one selected trajectory. Canonical judge is WebJudge [ Xue et al., 2025 ] + o4-mini at threshold 4; Sonnet is a cross-check.
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolving, including progress verification, grounded reflection, and failure recovery. DeepSearch-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine-tuning to train stronger agents. Without distillation from more capable models, DeepSearch-World-9B achieves competitive performance compared with open-source agents, reaching 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, showing that verifiable environments enable scalable self-evolution for long-horizon web agents. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self-improving deep search agents.
As reinforcement learning continues to scale the training of large language model-based agents, reliably verifying agent behaviors in complex environments has become increasingly challenging. Existing approaches rely on rule-based verifiers or LLM-as-a-Judge models, which struggle to generalize beyond narrow domains. Agent-as-a-Judge addresses this limitation by actively interacting with environments and tools to acquire verifiable evidence, yet its capabilities remain underexplored. We introduce a benchmark AJ-Bench to systematically evaluate Agent-as-a-Judge across three domains-search, data systems, and graphical user interfaces-comprising 155 tasks and 516 annotated trajectories. The benchmark comprehensively assesses judge agents' abilities in information acquisition, state verification, and process verification. Experiments demonstrate consistent performance gains over LLM-as-a-Judge baselines, while also revealing substantial open challenges in agent-based verification. Our data and code are available at https://aj-bench.github.io/.
Wentao Shi, Yu Wang, Yuyang Zhao +8
University of Science and Technology of China · National University of Singapore · Meituan
Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a promising way to improve these agents, but current approaches can fail, because correct answers are often sparse and score-based selection depends on model calibration. We propose FineVerify, a fine-grained self-verification framework that decomposes each question into checkable sub-questions, verifies sampled candidates against each sub-question, and selects the candidate with the highest aggregated score. This per-check structure turns selection into simpler local judgments and produces scores under the same explicit criteria. Across four agentic search benchmarks and two models, FineVerify consistently outperforms standard scaling baselines. With only four sampled trajectories, it improves GPT-5-mini by 8.2 accuracy points and Gemini-3-flash by 5.6% on average. With 12 samples, FineVerify enables GPT-5-mini to surpass frontier GPT-5 on BrowseComp-Plus. Beyond accuracy, FineVerify produces interpretable verification traces that help audit benchmark errors, suggesting broader applications for inspecting agentic search systems. Code and data are available at https://github.com/XuZhao0/fineverify