Characterizing High Bandwidth Flash for LLM Serving
Authors: Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, +2 more
Organizations: University of California, Berkeley Berkeley, California, USA · FuriosaAI Seoul, South Korea · University of California, Berkeley; ICSI; LBNL Berkeley, California, USA
Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.7% relative to HBM-only systems. Modeled energy savings reach 59.1%, with benefits depending on the workload and weight placement. Buffered cache-aware scheduling extends estimated HBF write lifetime from 1.21 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.
Figures & tables
Figure 1. Simulator versus profiled vLLM on H100: total input-plus-output token throughput. Simulated request lengths follow lognormal distributions with σ=0.5 ; host KV offload is disabled. Two side-by-side line plots compare simulated and measured vLLM throughput for 4,096-token prompts and output lengths of 128 and 512 tokens, across 32 to 4,096 requests.
Figure 2. Two memory architectures we considered in this study. The sites refer to the space where HBM/HBF stacks can connect to the accelerator. HBM and HBF compete for shared sites in the first architecture, whereas HBM and HBF are connected sequentially in the second architecture. We simulate both architectures. Two diagrams contrast four HBM and four HBF stacks sharing sites with eight HBF stacks connected behind eight HBM stacks.
Figure 3. Hierarchical placement of weights, activations, and KV state. Host reload returns to HBM, and colder device state can move through HBF to the host. The buffer drawn on the HBF read path is the latency hiding buffer of the H 3 architecture. A four-level diagram shows compute below HBM, HBF, and host storage, with offload arrows upward and a host-to-HBM reload path.
Figure 4. Vanilla cache-aware scheduling can lose a future cache hit. At time step 1, the scheduler selects the cache hit from session i , then admits session 1 in queue order, evicting session k ’s cached prefix. When session k returns at time step 2, its prefix is no longer resident. Two time steps show admission of a new session evicting the cached prefix of session k before its next request arrives.
Figure 5. Buffered cache-aware scheduling can preserve a future cache hit. Delaying session 1 leaves session k ’s cached prefix resident for reuse at time step 2. Dotted lines illustrate admission headroom. Session k is retained at time step 1, and session j completes and remains cached at time step 2. Two time steps show an admission buffer delaying session 1 so that session k can reuse its cached prefix when it returns.
Parameter
HBM
HBF
Capacity (GB)
24 ( Micron Technology, 0000 )
375 ( Ha et al., 2026 )
Read bandwidth (GB/s)
1,000 ( Lenovo, 0000 )
1,000 ( Ha et al., 2026 )
Write bandwidth (GB/s)
1,000 ( Lenovo, 0000 )
8
Read startup ( μ s)
0.1 ( Ma and Patterson, 2026 )
20 ( Ha et al., 2026 )
Write startup ( μ s)
0.1 ( Ma and Patterson, 2026 )
250
Read energy (pJ/bit)
3.4 ( Moon et al., 2023 )
6.8 ( Wu et al., 2026 )
Table 1. Per-stack memory assumptions. GB and TB use decimal units. HBF read energy is set to twice HBM's following MemExplorer ( Wu et al., 2026 ) , however H 3 reports up to four times the power consumption ( Ha et al., 2026 ) . Given limited published HBF write specifications, we use typical SSD write parameters and assume SLC endurance of 100,000 P/E cycles. Since activation writes remain in HBM, HBF sees marginal write traffic compared with read traffic, modeled energy is therefore not very sensitive to the exact HBF write cost assumption.
Operation type j
Energy ej (pJ/op)
FP16 multiply–accumulate
0.38
FP32 addition
0.22
FP32 multiplication
0.71
FP32 exponentiation
7.4
FP32 division
5.6
FP32 square root
11.1
Table 2. Arithmetic energy parameters at 4 nm, scaled from 7 nm by 0.546 ( Zhao et al., 2026 ; TSMC, 2020 ; TSMC, 2021 ) .
Model
GPUs
Turns
Parallelism
Qwen3-32B
1
25k
TP1
Llama-3.1-405B
8
100k
TP8
GLM-5.2
16
100k
TP16/EP16 + DPA
GLM-5.2
16
100k
2× TP8/EP8 + DPA
GLM-5.2
16
500k
TP16/EP16 + DPA
GLM-5.2
16
500k
2× TP8/EP8 + DPA
Table 3. Workload and parallelism configuration. TP/EP denotes tensor/expert parallelism with the stated group size; DPA denotes data-parallel attention. 2× TP8/EP8 uses two eight-GPU replicas. The 1 × and 5 × traces contain 100k and 500k turns. Qwen uses a separate 25k-turn workload.
Memory
Scheduling
Clock (s)
Energy (MJ)
HBM KV writes (TB)
HBF KV writes (TB)
Device hit (%)
Host hit (%)
Life (years)
8 HBM
FCFS
19,033.77
2.323
158.54
0.00
0.16
96.36
—
8 HBM
Cache aware
18,479.17
2.274
148.41
0.00
6.61
89.92
—
8 HBM
Cache aware, 10%
16,313.41
2.245
69.35
0.00
56.79
39.73
—
4 HBM + 4 HBF
FCFS
23,228.20
2.983
8.23
147.15
7.30
89.23
0.75
4 HBM + 4 HBF
Cache aware
20,821.73
2.849
7.91
93.00
41.72
54.80
1.06
4 HBM + 4 HBF
Cache aware, 10%
16,614.28
2.703
6.89
7.48
96.06
0.47
10.57
Table 4. Qwen3-32B scheduling on one GPU, 25k turns, and HBM-resident weights. Shared-site 4 HBM + 4 HBF divides eight stack sites. H 3 adds eight HBF stacks alongside eight HBM stacks. FCFS means first-come, first-served. Cache aware, 10% reserves 10% admission headroom (bold hybrid rows). Clock is workload completion time. Device and host hit rates are the fractions of admitted input tokens served from device memory (HBM or HBF) and host storage (DRAM or SSD), respectively. Life estimates HBF write endurance via Equation 8 . A dash indicates that HBF lifetime is not applicable to HBM-only configurations, as HBM is not subject to flash program/erase-cycle limits.
Configuration
Time (s) ↓
Energy (MJ) ↓
TPOT (ms) ↓
8+0, HBM weights
39,666
53.60
90.85
H 3 8+8, HBM weights
21,439
40.70
438.61
H 3 8+8, HBF weights
21,423
42.06
438.27
Table 5. Llama-3.1-405B weight placement on 8 GPUs with eight-way tensor parallelism (TP8), the 1 × trace (100k turns), and a 10% scheduling buffer. X+Y denotes X HBM and Y HBF stacks per GPU. 8+0 is HBM-only and H 3 adds HBF stacks alongside HBM. Time is workload completion time. TPOT is mean per-request time per output token after the first. Arrows indicate lower is better. Bold marks each column minimum.
1 × trace
5 × trace
Architecture
Stacks
Parallelism
Weights
Time (s)
Energy (MJ)
TPOT (ms)
Time (s)
Energy (MJ)
TPOT (ms)
HBM-only
8+0
TP16
HBM
11,354
9.11
218.72
198,411
91.99
191.93
H 3
8+8
TP16
HBM
4,543
7.19
279.15
28,279
37.62
355.63
H 3
8+8
TP16
HBF
4,538
9.54
278.01
27,036
48.17
333.66
H 3
8+8
2× TP8
HBF
4,204
10.31
230.67
24,385
52.06
280.43
H 3
8+8
2× TP8
Split
4,433
7.79
252.61
25,661
40.89
307.66
Table 6. GLM-5.2 architecture, placement, and parallelism on 16 GPUs with a 10% scheduling buffer. X+Y denotes X HBM and Y HBF stacks per GPU. H 3 adds HBF alongside HBM, while shared-site designs divide eight stack sites between them. TP16 uses one 16-GPU replica with TP16/EP16, while 2× TP8 uses two 8-GPU replicas with TP8/EP8. Both use data-parallel attention. HBM/HBF denotes all weights in that tier. Split places selected experts in HBM, remaining experts in HBF, and other weights in HBM, using offline expert lookup and ideal expert balance. The 1 × /5 × traces contain 100k/500k turns with original/extended contexts. Time is workload completion time, and TPOT is mean per-request time per output token after the first. On the 5 × trace, among the tested shared-site allocations, 4+4 minimizes completion time and 6+2 minimizes energy. Bold marks each column minimum across all configurations. Figure 6 visualizes the completion-time and energy tradeoffs for six of these configurations.
Figure 6. Completion-time and modeled-energy tradeoffs for six GLM-5.2 configurations on (a) the 1 × and (b) the 5 × traces (100k/500k turns). X+Y denotes HBM+HBF stacks per GPU; TP16 is one 16-GPU replica and 2× TP8 denotes two 8-GPU replicas. Lower is better on both axes. Table 6 gives placement details. Two scatter plots comparing six configurations on the original and extended GLM traces. HBF configurations finish sooner than HBM-only serving. Among the plotted configurations, HBF-resident weights with two TP8 replicas minimize time, while HBM-resident weights with TP16 minimize energy on both traces.
Contribution
Llama 8+8 1 ×
GLM 6+2 1 ×
GLM 8+8 1 ×
GLM 4+4 5 ×
GLM 8+8 5 ×
Reduced weight reads
+15.22
+14.06
+19.50
+51.28
+55.14
Increased KV/index reads
−13.16
−0.69
−0.63
−1.39
−0.61
Avoided repeated prefill
+20.25
0.00
0.00
0.00
0.00
Other net savings
+1.74
+2.21
+2.21
+4.49
+4.58
Net energy savings
24.06%
15.58%
21.09%
54.37%
59.10%
Table 7. Energy-savings contributions relative to same-model, same-workload HBM-only energy. Llama denotes Llama-3.1-405B and GLM denotes GLM-5.2. X+Y gives HBM+HBF stack counts per GPU; 1 × /5 × denotes the original/extended trace with 100k/500k turns. The 8+8 columns use H 3 with HBM weights. The 6+2 and 4+4 configurations use shared-site expert placement. Entries are percentage points: positive values save energy and negative values add cost. Avoided prefill counts computation only, other savings are the residual, and bold totals may differ due to rounding.
School of Computing and Information Systems, The University of Melbourne · School of Computer Science and Technology, Huazhong University of Science and Technology (www.ruizhang.info)
Institute for Data, Systems, and Society, Massachusetts Institute of Technology, Cambridge, MA 02139 · Columbia Business School, Columbia University, New York, NY 10027 · School of Mathematical Sciences, Peking University, Beijing, China