cs.CROct 6, 2026

CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents

Authors: Rafid Ahmed, Joseph Fioresi, Mubarak Shah, Yuzhang Shang

Abstract

Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing. Safe execution requires distinguishing malicious requests from genuine ones without simply refusing to act. Despite its practical importance, this problem remains underexplored and it is unclear whether current agents or existing defenses can achieve it. To study this problem, we first propose CredLeak-Bench, a comprehensive benchmark designed to evaluate how effectively and securely agents automate human workflows when confronted with phishing and identity verification. The benchmark covers both user-directed authentication and autonomous inbox monitoring, where agents are not explicitly instructed to log in. It systematically varies deceptive cues and pairs phishing scenarios with legitimate counterparts, enabling joint evaluation of information leakage and utility on genuine tasks. Within a sandboxed environment, leakage is measured through actual submissions of information rather than agents' self-reported behavior. Our evaluation reveals that all tested models are vulnerable to leakage. Agents also disclose sensitive information during autonomous inbox monitoring, demonstrating that phishing can induce disclosure without a user request to authenticate. Furthermore, most evaluated mitigations that reduce leakage also impair performance on genuine tasks, exposing a security utility trade off in existing defenses. These findings show why reducing leakage alone is insufficient: effective defenses must prevent unauthorized disclosure while preserving legitimate task completion. CredLeak-Bench provides a controlled framework for measuring both objectives and evaluating progress toward secure, useful agents.

Explore similar work

CardsList
  1. PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say

    May 29, 2026Mingxuan Zhang, Jiahui Han, Dadi Guo +5Data LeakagePrivacy Auditing

  2. DecepEval: A Benchmark for Evaluating Deception in LLM Agents

    Oct 6, 2026Yiming Xu, Hongyue Yu, Beihua Yang +8Deception in Language ModelsLLM Agent Evaluation

  3. An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

    Jun 15, 2026Hankyul Baek, Jaewon Noh, Sang Seo +5Data LeakageLLM Agent Evaluation