cs.AIJul 31, 2026

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

Authors: Yuan GaoZeren YangJunnan LiShawnZhongAhmed DajaniMai ZhengAndrea Arpaci-Dusseau+1 more

Organizations: Wanxiang · University of Wisconsin–Madison · 2Iowa State University

Abstract

Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.

Explore similar work

CardsList
  1. Agents' Last Exam

    Jun 3, 2026Yiyou Sun, Xinyang Han, Weichen Zhang +307Artificial Intelligence AgentsTaxonomy