cs.LGSep 30, 2026

Characterizing High Bandwidth Flash for LLM Serving

Authors: Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, +2 more

Organizations: University of California, Berkeley Berkeley, California, USA · FuriosaAI Seoul, South Korea · University of California, Berkeley; ICSI; LBNL Berkeley, California, USA

Abstract

Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.7% relative to HBM-only systems. Modeled energy savings reach 59.1%, with benefits depending on the workload and weight placement. Buffered cache-aware scheduling extends estimated HBF write lifetime from 1.21 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.

Figures & tables

Explore similar work

CardsList
  1. Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization

    Jul 9, 2026Jiantong Jiang, Peiyu Yang, Rui Zhang +1Kv-Cache ManagementLarge Language Model Serving

  2. Service-Induced Congestion in Memory-Constrained LLM Serving

    Jun 14, 2026Ruicheng Ao, Jing Dong, Gan Luo +1Large Language Model ServingLarge Language Model Memory

  3. LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

    Jul 29, 2026Ming-Yen Lee, Hanchen Yang, Faaiq Waqar +4Large Language Model MemoryVram