Organizations: Institute of Artificial Intelligence, China Telecom (TeleAI) China · Shanghai Jiao Tong University · Individual Researcher · Tsinghua University Institute of Artificial Intelligence, China Telecom (TeleAI) China
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbitrary prefix boundaries. Existing systems materialize recurrent-state checkpoints, restricting prefix reuse to checkpoint-aligned positions. We present SuffixReplay, the first prefix caching system that lets hybrid LLMs reuse cached prefixes at every cache-supported page boundary without materializing recurrent-state checkpoints. Our key insight is to just let linear states forget the distant past. Modern linear-attention mechanisms use recurrent decay and gating to attenuate the influence of old inputs. Therefore, instead of checkpointing every prefix boundary, SuffixReplay approximates the state at a matched boundary by replaying only a recent suffix of the layer's input hidden states, which we retain as anchors. At the algorithmic level, SuffixReplay combines layer-wise and token-wise anchor sparsity with a bounded replay budget to control storage, computation, and quality. At the system level, it uses an independently managed anchor sidecar and a pipelined replay path to overlap anchor movement and state reconstruction with the native serving pipeline. We evaluate SuffixReplay on three hybrid LLMs: OLMo-Hybrid-7B, Qwen3.5-4B, and Qwen3.6-27B-FP8. Across these models, SuffixReplay retains 91.4-100% of full-prefill quality on average across LongBench and RULER, while using only 0.36-0.51x the amortized per-token storage of SGLang's default 8192-token checkpoint cache. Integrated into SGLang, SuffixReplay reduces median TTFT by 15-70% on branching workloads, sustains 2.3-4.3x SGLang's throughput when the working set exceeds HBM, and matches SGLang on high-hit continuation traffic.
Figures & tables
Model
SGlang@8192
Naive
Ratio
OLMo-7B
6.5
180
27.6 ×
Qwen3.5-4B
6.1
120
19.5 ×
Qwen3.6-27B-FP8
18.4
480
26.2 ×
Table 1. Additional storage per token (KiB) on top of the KV cache. Appendix A derives both columns.
Model
A
SGlang
Naive
Ours
Ours/Native
OLMo-7B
7
6.5
180
3.3
0.50 ×
Qwen3.5-4B
7
6.1
120
2.2
0.36 ×
Qwen3.6-27B-FP8
15
18.4
480
9.4
0.51 ×
Table 2. Additional storage per token (KiB) with anchors at the entry of each group ( A layers) and anchor density ρ=1/16 (Ours), against SGLang’s cache at its 8192-token interval and naive replay (Table 1 ). SuffixReplay stores A⋅d⋅2ρ bytes per token; Appendix A gives the derivation.
OLMo-Hybrid-7B
Qwen3.5-4B
Qwen3.6-27B-FP8
Cell
k=16
32
64
128
256
512
k=16
32
64
128
k=16
32
64
128
LongBench QA
NarrativeQA@16K
85.8
88.1
91.1
92.5
93.6
97.2
97.5
98.5
100
100
98.3
97.2
100
99.3
NarrativeQA@32K
84.2
91.0
93.3
94.8
95.3
95.4
91.6
91.8
97.4
97.9
100
95.8
99.2
100
NarrativeQA@64K
85.6
82.9
86.8
92.8
93.3
96.8
94.1
96.5
100
99.2
95.7
100
97.2
99.7
HotpotQA@16K
86.8
89.9
90.9
95.5
92.5
94.1
99.2
99.5
99.6
100
99.2
98.1
97.5
97.4
Table 3. Anchor replay quality relative to full prefill (%full) for different replay budgets k . Each entry is a paired sample average; entries below 90% are bold. @ denotes context length.
throughput (req/s)
TTFT p50 (ms)
model
N
SGLang
ours
SGLang → ours
Qwen3.5-4B
8
21.9
23.1 (+5.4 %)
54 → 47
16
26.9
27.6 (+2.7 %)
62 → 57
32
32.7
32.2 ( −1.4 %)
68 → 65
Qwen3.6-27B-FP8
8
8.8
9.3 (+5.6 %)
97 → 86
16
11.4
12.1 (+6.5 %)
119 → 107
Table 4. High-hit agent continuation with 240 multi-turn traces and a 98% hit rate. Results use closed-loop HBM-only serving; N is the number of concurrent sessions.
SGLang prefill chunk
8192 (stock)
2048
1024
cold prefill TTFT p50 (ms)
677
823 (+22 %)
1,268 (+87 %)
throughput (req/s)
13.8
14.1
11.5 ( −16 %)
state bytes per cached token
6,286
25,144
50,288
Table 5. Effect of smaller SGLang prefill chunks on cold-prefill TTFT, throughput, and checkpoint storage. Results use Qwen3.5-4B in HBM-only serving.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
OLMo-7B
Qwen3.5-4B
Qwen3.6-27B
Llin
24
24
48
d
3840
2560
5120
Hk / Hv
30 / 30
16 / 32
16 / 48
dk / dv
96 / 192
128 / 128
128 / 128
C
11,520
8,192
10,240
Brec
2,211,840
2,097,152
3,145,728
Appendix
Table 6. Model parameters and the resulting sizes. All three models use convolution kernel width w=4 .