cs.DCOct 2, 2026

EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures

Authors: Yuhai Long, Yuanxin Wei, Kai Wu, Jinhui Wei, Dan Huang, Jiangsu Du

Organizations: School of Computer Science and Engineering, Sun Yat-sen University · China Mobile Internet Company Ltd

Abstract

Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU co-execution. Furthermore, speculative decoding in multi-agent workloads faces extreme variance in drafting difficulty, alternating between complex reasoning and predictable structured generation. Compounded by frequent tool-induced stalls, this highly fragmented execution severely underutilizes hardware and defeats traditional static batching. We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. At the micro-architectural level, it bypasses rigid graph-compiler constraints to enable zero-copy UMA-aware tensor parallelism, utilizing asymmetric memory layouts to fully saturate both CPU and GPU compute units. At the scheduling level, it dynamically allocates draft budgets based on real-time sequence predictability to bound bandwidth waste. Concurrently, an asynchronous suspend-and-yield mechanism actively evicts stalled agents, ensuring continuous hardware saturation during unpredictable tool invocations. Extensive evaluations on an Apple M4 SoC demonstrate that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding. Adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x speedup under extreme tool-use latencies.

Explore similar work

CardsList
  1. Agent-X: Full Pipeline Acceleration of On-device AI Agents

    May 11, 2026Jinha Chung, Byeongjun Shin, Jiin Kim +1LLM Inference EfficiencyOn-Device Language Model Inference

  2. E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments

    Jun 2, 2026Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La +3LLM ServingEdge Computing

  3. PPAI: Enabling Personalized LLM Agent Interoperability for Collaborative Edge Intelligence

    May 18, 2026Zile Wang, Qianli Liu, Kaibin Guo +4Multi-Agent CollaborationLLM Agents