Web Agents

Momentum

6 papers in the last four weeks, up 20% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 76

All topics
CardsList
  1. CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

    Oct 5, 2026Yifan Zhang, Yutong Dai, Viraj Prabhu +3Web Agents

  2. Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

    Oct 1, 2026Chengguang Gan, Zimeng He, Yoshihiro Tsujii +3Web AgentsAgentic Evaluations

  3. WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents

    Sep 28, 2026Anton Emelyanov, Maria Tikhonova, Zaven Martirosian +2Web AgentsWeb

  4. Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

    Sep 28, 2026Ruozhao Yang, Mingfei Cheng, Xiaofei XieWeb AgentsAgentic Deployments

  5. AX is the New AEO

    Sep 28, 2026Ido Finder, Assaf Elovic, Gad Shalev +1Agentic SearchWeb Agents

  6. Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

    Sep 23, 2026Chengguang Gan, Yunhao Liang, QingHao Zhang +1Web AgentsGuidance

  7. EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

    Sep 17, 2026Yinzhu Quan, Zefang LiuEconomiesWeb Agents

  8. Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision

    Sep 2, 2026Sitong Pan, Yipeng Shen, Yilin Lu +3Web AgentsEarly Failure Prediction

  9. SIR: Self-improving Red-teaming for Compute Use Agents

    Aug 31, 2026Chen Xiong, Zhiyuan He, Pin-Yu Chen +2Indirect Prompt InjectionRed-Teaming

  10. Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning

    Aug 22, 2026Chenghao Zhang, Yuxi Cheng, Saisai Hu +3Synthetic EnvironmentsWeb Agents

  11. SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents

    Aug 12, 2026Ruitao Wang, Yuwen Hao, Menglin YangWeb AgentsSynthetic Task

  12. MELLON - Multimodal Enhanced LLM for Online Navigation

    Aug 10, 2026Ruiyu Li, Haoyang Cai, Zhitong Guo +1Web AgentsMultimodal Large Language Models

  13. WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance

    Aug 7, 2026Zhi Li, Tao Zhou, Yeqing Li +2Web AgentsWeb

  14. Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

    Aug 6, 2026Jiaming Wei, Zekun Wu, Adriano Koshiyama +1Web AgentsStrategic Agents

  15. Falsifiable Commitment Planning for Self-Correcting Web Agents

    Jul 27, 2026Guangyi Liu, Huan Zhao, Quanming YaoWeb AgentsLong-Horizon Agents

  16. Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability

    Jul 13, 2026Said Elnaffar, Farzad RashidiWeb AgentsWeb

  17. MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

    Jul 11, 2026Chengguang Gan, Hanjun Wei, Yunhao Liang +3Web AgentsMultimodal Agents

  18. WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search

    Jul 9, 2026Xiaoshuai Song, Liancheng Zhang, Kangzhi Zhao +8Web AgentsAgentic Search

  19. Prismata: Confining Cross-Site Prompt Injection in Web Agents

    Jul 9, 2026Corban Villa, Alp Eren Ozdarendeli, Sijun Tan +1Web AgentsUntrusted Content

  20. DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

    Jul 8, 2026Xinyu Geng, Xuanhua He, Sixiang Chen +7Search AgentsSelf-Improving Agents

  21. WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

    Jul 7, 2026Wei Dong, Tianyu Fu, Zhe Yu +9Web AgentsEvaluation Benchmarks

  22. Untrusted Content Masking for Web Agents with Security Guarantees

    Jul 6, 2026Kristina Nikolić, Egor Zverev, Javier Rando +3Untrusted ContentWeb Agents

  23. Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services

    Jun 18, 2026Mugeng Liu, Shuoqi Li, Yixuan Zhang +1Tool InvocationAgentic Workflows

  24. Towards an Agent-First Web: Redesigning the Web for AI Agents

    Jun 17, 2026Eranga Bandara, Ross Gore, Ravi Mukkamala +18Web AgentsWeb

  25. When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration

    Jun 16, 2026Aagam Sogani, Botao Rui, Swetha Vaidyanathan +3Web AgentsTraces

  26. Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

    Jun 11, 2026Zihao Wang, Yiming Li, Yutong Wu +8Prompt-Injection DetectorsIndirect Prompt Injection

  27. MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents

    Jun 9, 2026Yv Zhang, Hao Sun, Hao Fang +5PoisoningMultimodal Memory

  28. WebChallenger: A Reliable and Efficient Generalist Web Agent

    Jun 9, 2026Jayoo Hwang, Xiaowen Zhang, Vedant PadwalWeb AgentsExploration

  29. M3^3Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

    Jun 5, 2026Zhengjun Huang, Wenxuan Liu, Zhoujin Tian +6Multimodal MemoryModalities

  30. Signal-Driven Observation for Long-Horizon Web Agents

    Jun 4, 2026Shubham Gaur, Ian LaneWeb AgentsLong-Horizon Agents

  31. SentinelBench: A Benchmark for Long-Running Monitoring Agents

    Jun 3, 2026Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin +5Web Agents

  32. Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval

    Jun 3, 2026Jiaxi Li, Ke Deng, Yun Wang +5Web AgentsSkills

  33. "I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

    May 30, 2026Soham Roy, Sarthakbrata Halder, Arya Bharaty +5Web AgentsData Leakage

  34. Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

    May 28, 2026Alejandra Zambrano, Sara Vera Marjanovic, Imene Kerboua +2Large Language Model PlanningWeb Agents

  35. GTA: Generating Long-Horizon Tasks for Web Agents at Scale

    May 28, 2026Tenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey +4Web AgentsAgentic Benchmarks

  36. VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

    May 27, 2026Yuting Xu, Jiayi Tian, Jian Liang +4Agentic BenchmarksWeb Agents

  37. VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

    May 22, 2026JunJia Guo, Yuhang Yao, Jiawei +2Web AgentsWeb

  38. Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

    May 19, 2026Fatemeh Pesaran Zadeh, Seyeon Choi, Xing Han Lù +2Web Agents

  39. Skim: Speculative Execution for Fast and Efficient Web Agents

    May 15, 2026Mike Wong, Kevin Hsieh, Suman Nath +1Web AgentsWeb

  40. WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections

    May 14, 2026Tri Cao, Yulin Chen, Hieu Cao +8Indirect Prompt InjectionWeb Agents

  41. Web Agents Should Adopt the Plan-Then-Execute Paradigm

    May 14, 2026Julien Piet, Annabella Chow, Yiwei Hou +5Web AgentsWeb

  42. Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection

    May 12, 2026Zi Liang, Ronghua Li, Yanyun Wang +2Attacker Large Language ModelDistributed Denial-Of-Service

  43. Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

    May 10, 2026Yilin Zhang, Yingkai Hua, Chunyu Wei +2Web AgentsDeception

  44. Weblica: Scalable and Reproducible Training Environments for Visual Web Agents

    May 7, 2026Oğuzhan Fatih Kar, Roman Bachmann, Yuanzheng Gong +2Web AgentsWeb

  45. BaRA: Budget-constrained and Reliable Web Data Collection Agent

    May 2, 2026Soojeong Lee, Joseph Lee, Yongseong Cho +3Web AgentsDiscovery

  46. Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction

    Apr 29, 2026Yuxuan Huang, Yihang Chen, Zhiyuan He +6Agentic SearchWeb Agents

  47. SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents

    Apr 28, 2026Mengyao Du, Han Fang, Haokai Ma +4Prompt-Injection DetectorsIndirect Prompt Injection

  48. DRIVE: Modeling Skills at the Reasoning and Interaction Levels for Web Agents under Continual Learning

    Apr 28, 2026Xirui Liu, Sihang Zhou, Yanning Hou +6Web AgentsInteraction Data

  49. Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

    Apr 27, 2026Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried +1Web AgentsLong-Horizon Task Planning

  50. Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces

    Apr 26, 2026Zijing Shi, Meng Fang, Ling ChenWeb AgentsDeception