Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.
Figures & tables
Figure 1: Reader adaptation, reuse depth, and storage precision. Qwen3-8B throughout. (a) Native LoCoMo full Judge, 1,986 items, fixed j=12 and evidence; only suffix adaptation changes. (b) Separately adapted splits: RULER-B on 1,500 examples versus fixed-pack prefill from three processes. (c) Fixed-reader FP16 H16/H8/H4: RULER means over 1,500 examples and persistent GPU Store on six 8k documents. Quality and storage use distinct supports. Lines connect measured settings; the depth and precision studies are not a joint sweep.
Figure 2: A persistent memory boundary inside the decoder. Independent lower-layer Write produces document residuals, stored in H16/H8/H4 format. Token-ID retrieval selects chunks; only their states are reconstructed. The native sink and query join the selected states for causal upper-layer Read with suffix LoRA. The diagram describes the attention-only Qwen3-8B implementation.
Method
RULER
LongEval
QA6 F1
BABILong
LoCoMo
Qwen3-8B; full attention; j=12 ; FP16 precision study
Dense, full source
51.91
31.6
8.97
58.52
32.40
KIVI2
42.43
9.0
9.37
53.38
29.81
KIVI4
50.64
31.2
8.89
58.90
32.60
Residual, no LoRA
16.01
0.0
9.98
38.38
25.45
EncBank-H16
97.27
63.0
12.05
58.05
44.03
Table 1: Quality across storage precisions and backbones. Scores are 0–100. Within each H16/H8/H4 group, the reader is fixed; other rows are whole-method references. 8B values follow the complete FP16 precision study, not the separate native-reuse cohort. Hybrid values are reported aggregates ( † ); see the scope below.
Path
RULER-B
MK
Prefill (ms)
Reusable
Same-adapter replay, j=0
99.19
100.0
931.9
No
Chunk-local H16, j=12
96.07
92.5
664.4
Yes
Continuous-prefix oracle
99.19
100.0
NR
No
Table 2: Matched native reuse measures a fidelity cost. RULER-B uses 1,500 paired examples; MK is the 200-example 8k/16k multikey subset. Prefill uses a separate fixed pack, three processes, three warmups and 20 reads/process; it excludes selection, Write, fetch and decode. NR: latency not reported for the per-query oracle.
Panel
Configuration
Measured outcomes
(a) Interaction
MK/8k
MK/16k
Full causal
96
94
Block diagonal
60
32
(b) Write context
MK
LE
Qasper F1
No overlap
92.5
72.67
11.37
32-token overlap
98.5
62.33
10.76
Table 3: Context interaction and local repair, with separate cohorts. (a) Multikey full/block-diagonal attention, same evidence, 50 examples/length; query sees all chunks. (b) Left-overlap Write at fixed stored shapes: MK has 200 examples; independent LongEval (LE) has 300, Qasper 200. Scores are 0–100.
(a) Method
Adapter
LongEval
Qasper F1
KiB/token
Replay
off
84.33
6.98
.008
Chunk-KV
off
73.00
4.88
144
Residual H16
off
0.00
10.45
8
Replay
on
77.67
4.65
.008
Chunk-KV
on
55.67
3.68
144
EncBank-H16
on
72.67
11.37
8
Table 4: Matched evidence exposes the residual–KV trade-off. (a) Quality on 500 examples: LongEval at 8k/16k/32k (100 each), Qasper (200). The shared adapter was trained for EncBank. (b) RTX 5090 costs on five prespecified LongEval/32k examples, shared adapter on, three processes and three repetitions/example. Preparation builds the whole-document pinned-CPU store; E2E includes 128 output tokens. Chunk-KV is a synchronous CacheBlend-style control.
Method
Store
Write (ms)
Warm TTFT (ms)
Tokens/s
Peak
Dense
1073.51
710.68
654.71
44.96
21.223
KIVI2
210.83
1127.73
1485.31
20.19
20.604
KIVI4
343.55
1128.38
1486.14
19.99
20.603
EncBank-H16
59.64
265.62
1382.98
33.56
20.648
EncBank-H8
31.69
274.12
1392.93
32.95
20.627
EncBank-H4
16.78
271.76
1378.10
32.97
20.617
Table 5: Persistent storage savings versus request costs. RTX 5090, FP16, six RULER/8k documents, 18 fresh Write/Read pairs, 32 fixed output tokens. Only H variants share the reader. Warm TTFT excludes document Write. Store is persistent GPU MiB; peak is sampled whole-device GiB, including weights. Write times are component means; decode throughput is pooled.
Figure 3: Bounded online allocation and residency. (a) Native H16/CPU-pinned online TTFT-phase peak on RTX 5090; full-source Dense fails at 32k/128k under a 28 GB allocator cap. (b) Separate FP16 GPU-resident test with distinct documents: H16 fails at Write 92, H4 completes 96. Levels are measured, not maximum-capacity estimates.
Source / store
G=1
G=32
G=128
G=512
32k / CPU pinned
8.9
9.2
10.9
94.0
32k / GPU resident
8.4
7.7
5.5
∞
128k / CPU pinned
27.6
29.7
37.4
520.1
128k / GPU resident
25.8
26.8
27.2
164.2
1M / CPU pinned
198.2
266.7
421.3
302.6
1M / GPU resident
180.2
183.0
190.6
574.7
Table 6: Native H16 preparation payback depends on reuse. Q⋆ in queries from measured Write, fetch, Read and decode; G is output tokens. Three process medians; stable stores and cache hits. Round up for an integer threshold. ∞ means nonpositive per-query savings. This grid is not an H4 amortization measurement.
Method
Pass / tasks
Success (%)
Dense, full context
61/89
68.54
EncBank-H16, k=12
70/89
78.65
Table 7: Terminal-Bench 2.1 verifier results on all 89 tasks. Each method retains one verified outcome per task; adapters and serving backends differ.