cs.CLOct 5, 2026

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Authors: Yifan Zhang, Yutong Dai, Viraj Prabhu, Zhiyuan Hu, Ran Xu, Zeyuan Chen

Organizations: Salesforce AI Research

Abstract

Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

    Jul 8, 2026Xinyu Geng, Xuanhua He, Sixiang Chen +7Search AgentsSelf-Improving Agents

  2. AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation

    Apr 20, 2026Wentao Shi, Yu Wang, Yuyang Zhao +8Llm-As-A-Judge

  3. FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

    May 30, 2026James Xu Zhao, Hui Chen, Bryan Hooi +1Agentic SearchAgentic Benchmarks