cs.CLSep 29, 2026

The Backdrop Exposes What the World Around an Agent Costs It

Authors: Nusrat Jahan Lia, Shubhashis Roy Dipta

Organizations: University of Dhaka · University of Maryland, Baltimore County

Abstract

Agent benchmarks test agents in worlds that stay still. Deployed agents work in worlds that other people also change. Someone texts the agent to send the money elsewhere or an order confirmation asks it to reply with a door code. We present BACKDROP, which asks how much of an agent's capability in a clean world survives in such a world. BACKDROP takes a task along with the agents execution environment, and plants four everyday hazards in its world, one at a time and all together. The instruction and the correct end state stay the same. Each hazard asks one question. Authority: does a message from another person override the user? Injection: does text planted in a record redirect the agent? Boundary: does a request pull it into an app it was not given? Fault: after a write fails without saying whether it landed, does the agent check before it retries? Across 3,678 variants and 16 models, , the average pass rate falls from 69.5% to 31.3% once all four hazards are present; the strongest models fall furthest (Claude Fable 5.1 from 96.6% to 56.0%). Agents have learned to resist injected text but often follow other unauthorized requests of other people. With all four hazards present, and counting only runs where the planted text reached the agent, agents followed another person's message in 46.4% of runs and injected text in 20.3%. The gap is consistent throughout all 16 models. BACKDROP formalizes these gaps and shows how an agent's score in a task's world is a ceiling on real-world performance.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GuardianAgentBench: Where Agents Fail and How to Guard Them

    Jul 23, 2026Vishal Ishwar Naik, Chenyu Xu, Donna Dong +5Agentic BenchmarksLarge Language Model Safety

  2. AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

    Apr 3, 2026Yunhao Feng, Yifan Ding, Yingshui Tan +6Computer-Use AgentsHazard

  3. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

    Jul 24, 2026Jiaqi Shao, Hanck Chen, Wei Zhang +2Agentic BenchmarksExploitation