cs.CRMay 21, 2026

Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard

Authors: Sahar AbdelnabiChris HicksKonrad RieckAhmad-Reza Sadeghi

Organizations: ELLIS Institute Tübingen & MPI-IS & Tübingen AI Center, Germany · The Alan Turing Institute, London, UK · BIFOLD & Technische Universität Berlin, Germany · Technische Universität Darmstadt, Germany

Abstract

The benchmarks used to evaluate AI agents in security-critical roles suffer from crucial weaknesses. Building on recent empirical evidence, we characterize three core challenges that undermine security evaluations: benchmark vulnerabilities, temporal staleness, and runtime uncertainty. We then outline practical directions toward building more robust and trustworthy evaluation frameworks.

Explore similar work

CardsList
  1. Agent Security is a Systems Problem

    May 18, 2026Mihai Christodorescu, Earlence Fernandes, Ashish Hooda +11SecurityAgentic Systems