cs.CLOct 7, 2026

Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

Authors: Hanzuo Liu, Chunyu Liu, Chaofan Lin, Alex Lamb, Mingyu Gao

Organizations: Tsinghua University

Abstract

Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

Figures & tables

Explore similar work

CardsList
  1. InferScale: GPU-Native KV Injection for Personalized LLM Serving

    Jul 29, 2026Peter Li, Prashant PandeyLarge Language Model MemoryKv-Cache Management

  2. FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

    Jul 7, 2026Anna Córdoba, Adam Puente Tercero, Nerea Angulo Hijo +4Key-Value Cache CompressionDepthweave-Kv