CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Organizations: Salesforce AI Research
Abstract
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.
Figures & tables
| Approach | Dense | URL | Cert. | J-free | TTS | Xfer |
| Outcome RLVR / GRPO (trajectory outcome) ( Shao et al., 2024 ) | ✓ | |||||
| Stratified GRPO (structural strata) ( Zhu et al., 2025 ) | ✓ | ✓ | ||||
| Rich-feedback self-distillation (teacher feedback) ( Hübotter et al., 2026 ) | ✓ | ✓ | ✓ | |||
| Process / verifier reward models (learned PRM) ( Lightman et al., 2024 ) | ✓ | ✓ | ||||
| Post-hoc search / BoN (sample then select) ( Snell et al., 2025 ) | ✓ | ✓ | ||||
| Conformal routing / judging (uncertainty gate) ( Badshah et al., 2026 ) | ✓ |
| Benchmark | Training role | Eval role |
|---|---|---|
| WAI | train Gemma-4 | in-domain + board |
| VWA | train bank | transfer to GPT-5.5 |
| OM2W | none | zero-shot transfer |
| Gemini 3 | Qwen 3.5+ | Kimi | Gemma-4 | Gemma-4 + CLIFT +CTS | |
|---|---|---|---|---|---|
| App (n) | Flash + BU | K2.5 | (base, ours) | ( , ours) | |
| Elation Clinical (120) | 82.0 | 54.0 | 50.0 | 69.2 | 91.7 |
| Elation Prescription (120) | 81.0 | 42.0 | 23.0 | 70.0 | 91.7 |
| GitLab Plan & Track (140) | 64.0 | 37.0 | 39.0 | 59.3 | 69.3 |
| Gmail (60) | 75.0 | 57.0 | 70.0 | 60.8 | 78.3 |
| Handshake (200) | 50.0 | 50.0 | 50.0 | 47.0 | 56.0 |
| Method | C | S | R | Avg. |
|---|---|---|---|---|
| SGV ( Andrade et al., 2026 ) | ||||
| WALT ( Prabhu et al., 2025 ) | ||||
| Gemma-4 base | ||||
| Gemma-4 + CLIFT | ||||
| Gemma-4 + CLIFT+CTS (K=2) | ||||
| GPT-5.5 + CLIFT |
| Origin | Construction | Certified | |
|---|---|---|---|
| WAI bank | keep/rewrite | ( ) | |
| VWA bank | keep/rewrite | ( ) | |
| OM2W expansion | live-web checks | ( ) | |
| Final bank | 132 qs | 132 | 100 (75.8%) |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Origin | Stage | Questions | Certified | Mean lift |
|---|---|---|---|---|
| WAI bank | kept as-is | 7 | 7 | 0.46 |
| WAI bank | rewritten | 29 | 26 | 0.46 |
| VWA bank | kept as-is | 10 | 4 | 0.31 |
| VWA bank | rewritten | 38 | 27 | 0.28 |
| OM2W expansion | live-web checks | 48 | 36 | 0.30 |
| Total | – | 132 | 100 | 0.34 |
| System | Canonical SR | Gain | Sonnet SR | |
|---|---|---|---|---|
| Gemma-4 base | 1 | 40.0 | – | 46.8 |
| Gemma-4 + ours | 2 | 45.0 | +5.0 | partial positive |
| GPT-5.5 low base | 1 | 43.7 | – | 47.3 |
| GPT-5.5 high base | 1 | 49.7 | – | 52.0 |
| GPT-5.5 high + ours | 2 | 57.3 | +7.7 | 55.7 |
| GPT-5.5 high + ours | 4 | 61.0 | +11.3 | 61.7 |