cs.AROct 7, 2025

Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM

Authors: Tianhao Zhu, Dahu Feng, Erhu Feng, Yubin Xia

Organizations: Institute of Parallel and Distributed Systems, Shanghai Jiao Tong University · Department of Precision Instrument, Tsinghua University

Abstract

With the widespread adoption of Large Language Models (LLMs), the demand for high-performance LLM inference services continues to grow. Multi-core AI accelerators, such as Groq, Graphcore IPU, and Cerebras WSE, provide promising platforms for LLM serving, but their distributed memory systems require careful coordination between hardware configuration and serving policies. Otherwise, mismatched tensor partitioning, data placement, and memory management can substantially underutilize compute and communication resources. To address these challenges, we present WaferAI-SIM, a multi-level simulation framework that combines transaction-level simulation with an analytical performance model. WaferAI-SIM enables simulator-driven co-design of LLM serving strategies and multi-core accelerator architectures, targeting the early design stage in which emerging platforms are not yet broadly available for empirical serving studies. It captures how LLM serving policies interact with compute-core count, memory hierarchy, and interconnect topology, enabling architecture-aware exploration beyond GPU-centric assumptions. We evaluate representative LLMs across a range of chip configurations and serving scenarios. Across the evaluation, WaferAI-SIM reports 1.32×\times--6.03×\times latency improvements, where the lower endpoint comes from ring-based placement at TP=16 over the placement baselines, and the upper endpoint comes from K-dimension TP over MN-dimension TP for Qwen3_4B at TP=4 with sequence length 256. For LLM serving, our findings provide guidance for co-designing hardware architectures and serving strategies for multi-core AI accelerators across diverse LLM workloads.

Figures & tables

Explore similar work

CardsList
  1. Frontier: Towards Comprehensive and Accurate LLM Inference Simulation

    May 20, 2026Yicheng Feng, Xin Tan, Yangtao Deng +3LLM Inference OptimizationUser Simulation

  2. AGENTSERVESIM: A Hardware-aware Simulator for Multi-Turn LLM Agent Serving

    Jun 8, 2026Rakibul Hasan Rajib, Mengxin Zheng, Qian LouLarge Language Model AgentsHardware-Aware