Hypernetworks that map a context directly to a LoRA adapter let a large language model carry that context in its weights, but prior work has demonstrated them only on base models of up to 14 billion parameters. We present the Internalizer, a state-of-the-art, portable Context-to-Parameter Mapping hypernetwork that generates document-specific LoRA adapters for the frozen 284B-parameter DeepSeek v4 Flash, a target two orders of magnitude larger than in any previous work. Most of its parameters live in a model-agnostic trunk with only thin entry and exit layers per base model, so it trains cheaply against small models before being ported to the large one. On unseen documents of up to 4096 tokens, the generated adapters reach 84.9% top-1 and 97.8% top-5 teacher-forced accuracy against 63.4% and 83.5% for the base model, with nothing in the context window but a three-word instruction. Once the hypernetwork is trained, a single forward pass turns any document into an adapter for such a model, which could be served alone for speed or alongside the document in the window to raise accuracy further.
Figures & tables
Figure 1: System architecture, with the shared trunk in light green and the thin model-specific entry and exit layers in dark green.
Perplexity
Top-1 (%)
Top-5 (%)
Length
Docs
Rank
Adapter
Base
Adapter
Base
Adapter
Base
FinePDFs (train distribution)
256
100
32
1.27
7.85
92.6
57.5
99.5
79.3
512
87
32
1.55
6.89
86.2
59.1
98.7
80.6
1024
67
48
1.55
5.99
86.4
61.3
98.7
82.1
2048
49
80
1.58
5.58
85.8
61.8
98.6
83.0
Table 1: Document reconstruction by document set and context length, comparing Adapter-only regurgitation against the unmodified base model. Rank is the adapter’s combined LoRA rank, 16(k+1) for k chunks. Accuracies are means over the row’s documents, perplexity is the exponential of the mean per-document cross-entropy, and the better value is in bold. † Beyond the 4096-token, eight-chunk training cap, so the hypernetwork generates adapters at chunk counts it never saw in training.
Figure 2: Adapter versus base model across context lengths, one document set per row. FinePDFs (train distribution) covers 100 unique documents (24–100 per context length, 368 evaluations), FinePDFs (long) holds the same 32 documents at every context length (192 evaluations), and PG-19, an out-of-domain set, covers 100 unique documents (95–100 per context length, 594 evaluations). Error bars show the standard error of the mean. The 16384-token bucket is omitted, as its adapter perplexities would dominate the axis. It is discussed in Section 6.2 .
Figure 3: Evaluation loss and top-1 accuracy while porting to DeepSeek v4 Flash. The trunk-frozen and fully unfrozen runs start from the warmed shared trunk, while the cold-start run trains its hypernetwork from scratch.
Institute for Artificial Intelligence, Peking University · School of Electronics Engineering and Computer Science, Peking University · University of Oxford +1