cs.MAJul 28, 2026

Towards a Systems Foundation for Agentic Cloud Management

Authors: Minghao LiZiqian LiuZiyu MaoDaqian DingYu KangQingwei LinTianyin XuYiming Qiu

Abstract

Agentic cloud management is emerging as a practice to automate laborious operations, minimize toil, and improve responsiveness. Despite the rapid development of autonomous management agents, we argue that the fundamental missing piece is a systems foundation to enable safe, effective operations across agents and between agents and human operators. In this paper, we advocate for the need of such a systems foundation and share our efforts on developing CloudWeaver, an agentic management substrate that works across existing cloud-user interfaces and future agent-native interfaces. Specifically, we discuss how CloudWeaver (1) scopes the context of individual agent sessions with local views of cloud resources and (2) coordinates concurrent management operations on shared cloud resources. CloudWeaver offers strong safety guarantees and attributable feedback in the presence of conflicting intents, while preserving concurrency between independent operations. We validate CloudWeaver using a representative Azure API workload.

Explore similar work

Apr 26, 2026cs.NI

Infrastructure for the Agentic Web: Gap Analysis and Architecture from the Agentverse Platform

The emergence of autonomous AI agents as first-class participants in digital infrastructure marks a fundamental inflection point in the evolution of the Web. While significant research has been directed at agent behaviour and reasoning, comparatively little attention has been paid to the infrastructure those agents require to operate reliably at scale. This paper addresses that gap with a systematic analysis of Agentverse, the agent cloud platform developed by Fetch.ai under the Artificial Superintelligence (ASI) Alliance, which represents one of the most mature production deployments of agent-native infrastructure available today. We make three principal contributions. First, we conduct an empirical audit of the Agentverse platform, cataloguing 204 API endpoints (Q1 2026) and characterising what is operational, partially deployed, or absent. From this audit we derive a Gap Taxonomy of eight categories encompassing 62 distinct missing capabilities, ranging from agent memory and observability to security, economic primitives, and enterprise scaling. Second, we propose a seven-layer Agent Cloud Stack -- a reference architecture for what a fully realised agent-native cloud should provide by 2030, grounded in the specific gaps we identify. Third, we characterise five critical evolution paths: from ephemeral storage to a full Agent Memory Cloud; from keyword discovery to a semantic, trust-weighted Agent DNS; from a single-protocol model to a multi-standard agent lingua franca; from single-instance hosting to Kubernetes-scale orchestration; and from simple token payments to rich agent economic primitives. Together these contributions provide a diagnostic of current agent infrastructure and a technically grounded vision for what the agent cloud must become to support the agentic web -- Web4 -- by 2030.
Robin Dey, Panyanon Viradecha
May 26, 2026cs.AI

Agyn: An Open-Source Platform for AI Agents with Scalable On-Demand Execution, Agent Definition as a Code, and Zero-Trust Access

As organizations move toward production deployments of AI agents, which execute non-deterministic workflows, maintain stateful sessions, and often operate with privileged access to internal services, the engineering challenge shifts from building individual agents to operating them at scale with proper isolation, governance, and security. In this paper we present Agyn, an open-source platform designed around three key principles tailored for agent workloads: a signal-driven, stateful serverless runtime on Kubernetes; a Terraform provider for agent and harness definition; and a security model grounded in zero-trust and least-privilege principles. Agyn is agent-agnostic, model-agnostic, and cloud-agnostic.
Nikita Benkovich, Vitalii Valkov
Jun 1, 2026cs.CR

Agent Operating Systems (AOS): Integrating Agentic Control Planes into, and Beyond, Traditional Operating Systems

Traditional operating systems were designed around deterministic programs, explicit control flow, and human initiated workflows. Their core abstractions processes, threads, system calls, files, and permissions assume bounded behavior and predictable interaction patterns. Agentic AI systems introduce a different execution model: long-lived, goal-directed entities that reason probabilistically, invoke tools dynamically, and adapt behavior based on feedback. While agents can be implemented as user-space applications today, their execution characteristics stress OS boundaries in scheduling, memory and state management, security, observability, and governance. This paper introduces the concept of an Agent Operating System (AOS), a systems architecture that integrates an agentic control plane into existing operating systems or, in some models, subsumes selected OS responsibilities over time. We provide a precise definition of an AOS, explicit assumptions and non-goals, and a structured decomposition of AOS responsibilities into schedulers, context and memory management, tool and capability registries, policy and trust enforcement, and observability and audit. We analyze limitations of classical OS abstractions for agent workloads, propose integration models from user-space runtimes to distributed control planes, and map AOS concepts onto Linux and Windows primitives. We present security and safety implications, including agent specific threat models, and define evaluation criteria that emphasize deterministic enforcement, auditability, and operator comprehensibility. The objective is not to replace operating systems wholesale, but to establish a rigorous systems foundation for agentic computation that remains controllable, accountable, and secure at scale.
Ankur Sharma, Deep Shah