cs.AROct 7, 2025

Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM

Authors: Tianhao Zhu, Dahu Feng, Erhu Feng, Yubin Xia

Organizations: Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University · Department of Precision Instrument, Tsinghua University

Abstract

With the widespread adoption of Large Language Models (LLMs), the demand for high-performance LLM inference services continues to grow. Multi-core AI accelerators, such as Groq, Graphcore IPU, and Cerebras WSE, provide promising platforms for LLM serving, but their distributed memory systems require careful coordination between hardware configuration and serving policies. Otherwise, mismatched tensor partitioning, data placement, and memory management can substantially underutilize compute and communication resources. To address these challenges, we present WaferAI-SIM, a multi-level simulation framework that combines transaction-level simulation with an analytical performance model. WaferAI-SIM enables simulator-driven co-design of LLM serving strategies and multi-core accelerator architectures, targeting the early design stage in which emerging platforms are not yet broadly available for empirical serving studies. It captures how LLM serving policies interact with compute-core count, memory hierarchy, and interconnect topology, enabling architecture-aware exploration beyond GPU-centric assumptions. We evaluate representative LLMs across a range of chip configurations and serving scenarios. Across the evaluation, WaferAI-SIM reports 1.32×\times--6.03×\times latency improvements, where the lower endpoint comes from ring-based placement at TP=16 over the placement baselines, and the upper endpoint comes from K-dimension TP over MN-dimension TP for Qwen3_4B at TP=4 with sequence length 256. For LLM serving, our findings provide guidance for co-designing hardware architectures and serving strategies for multi-core AI accelerators across diverse LLM workloads.

Figures & tables

Explore similar work

May 20, 2026cs.DC

Frontier: Towards Comprehensive and Accurate LLM Inference Simulation

Modern LLM serving is no longer homogeneous or monolithic. Production systems now combine disaggregated execution, complex parallelism, runtime optimizations, and stateful workloads such as reasoning, agents, and RL rollouts. Simulation is attractive for exploring this growing design space, yet existing simulators lack the architectural completeness and decision-grade fidelity it demands. Their monolithic-replica abstractions are ill-suited to disaggregated serving, while average-case analytical proxies can distort SLA predictions and even reverse optimization conclusions. We present Frontier, a discrete-event simulator for modern LLM inference serving. Frontier features a disaggregated abstraction. It captures the structure and dynamics of modern serving systems by modeling co-location, Prefill-Decode Disaggregation (PDD), and Attention-FFN Disaggregation (AFD) with role-specific cluster workers, incorporating key runtime optimizations (e.g., CUDA Graphs, speculative decoding) within the scheduler-batch-engine loop, and supporting stateful requests for emerging workloads. It further provides accurate and generalizable predictions of computation, communication, and memory costs across diverse serving scenarios with complex workload compositions. On 16-H800 GPU testbed, Frontier achieves an average throughput error below 4%. Compared with state-of-the-art simulators, it reduces end-to-end latency error from 44.9% to 6.4% under co-location and from 51.7% to 2.6% under disaggregation. It scales to over 1K GPUs on commodity CPUs and enables new use cases such as SLA-dependent Pareto frontier exploration, heterogeneous disaggregated allocation, agentic reasoning scheduling validation, and RL post-training reconfiguration. We release Frontier at https://github.com/NetX-lab/Frontier.
Jun 8, 2026cs.CL

AGENTSERVESIM: A Hardware-aware Simulator for Multi-Turn LLM Agent Serving

Multi-turn LLM agents interleave model calls with external tool invocations, shifting serving from stateless request processing to stateful program execution. Serving these workloads requires scheduling, KV-cache management, and routing policies that use program-level context, including turn dependencies, tool-induced gaps, and reusable KV state. Evaluating such policies directly on real systems is costly, since each design point may require dedicated accelerator time across arrival rates, model scales, serving-instance counts, and memory hierarchies. Simulation offers a scalable alternative, but existing LLM serving simulators target stateless request-level workloads and therefore omit the core dynamics of agent serving: multi-turn program execution, cross-turn cache locality, and KV-cache residency during tool gaps. We present AGENTSERVESIM, a hardware-aware simulator for multi-turn LLM agent serving. AGENTSERVESIM evaluates serving policies at program granularity through composable modules: a Program Orchestrator preserves program identity and turn order, a Tool Simulator materializes tool-induced gaps, a Session-Aware Router maintains program-to-instance affinity for cache-aware dispatch, and a KV Residency Model tracks policy-defined KV placement across HBM, host DRAM/CXL, and eviction. Across real serving deployments and hardware configurations, AGENTSERVESIM reproduces real-system behavior within 6% error across key performance metrics while running entirely on commodity CPUs. These results show that AGENTSERVESIM enables controlled, repeatable exploration of agent-serving policies without requiring exhaustive deployment on costly accelerators.
Oct 4, 2026cs.DC

Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs

Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selecting the most effective parallelism strategy remains challenging due to complex interactions among computation, communication, pipeline utilization, sequence length, batch size, and model architecture. Existing approaches largely rely on empirical evaluation and provide limited analytical insight into the trade-offs among these strategies, particularly across the distinct prefill and decoding phases of inference. In this paper, we present a unified analytical framework for modeling distributed LLM inference under TP, PP, and HB. The framework decomposes end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead, and derives analytical models that capture TP collective communication, PP point-to-point communication, and pipeline utilization as functions of hardware, model, and workload characteristics. The model further characterizes the differing execution behavior of prefill and decoding, explaining why PP-oriented configurations favor compute-intensive prefill while TP-oriented configurations reduce decoding latency by eliminating pipeline bubbles. Experiments with modern LLMs on multi-GPU platforms validate the model and confirm the fundamental compute-communication trade-off across parallelism strategies. The framework provides practical guidance for parallelism selection, capacity planning, and optimization of future LLM serving systems.