General-purpose models can adapt to many tasks from context, while specialised models can execute individual functions with less capacity. Yet obtaining such specialists requires task-specific training or adaptation. We ask whether they can instead be generated directly from a few demonstrations. Using ARC-1D as a controlled testbed, we show that individual transformations can be represented by tiny specialist models, and that a hypernetwork can generate their parameters from context. The generated parameters form a structured weight space, while the resulting specialists show partial compositional generalisation and generalisation to transformations not seen during training. In both settings, removing explicit task identifiers improves generalisation beyond the training transformations. Together, these results provide a proof of concept that few-shot task context can be compiled on-the-fly into compact executable model parameters, and that the resulting weight space can support reuse and generalisation beyond known functions.
Figures & tables
Figure 1: Hypernetwork weight generation. The support set ZE={(xi,yi)}i=1K is encoded by the task encoder ϕω into an episode representation zE , which the weight decoder ψρ maps to the complete parameters wE of the episode-specific predictor f(⋅;wE) . Once generated, the predictor operates independently of the support set, receiving only a new input x′ to produce its output.
Figure 2: Flip transformation: Example ARC-1D episode The support set ZE contains three input–output pairs specifying the transformation, which is applied to the query input in ZE′ . See Appendix G for more examples.
Figure 3: Specialisation requires less capacity. Individual models saturate earlier than joint models, even when provided with transformation identity. Bands show ±1 s.d. across 5 seeds.
Figure 4: The joint-training penalty is concentrated in a few transformations. Per-transformation exact-match validation accuracy at 2.4K parameters. Error bars show ±1 s.d. across 5 seeds.
Figure 5: Generated models recover specialist performance. Per-task validation exact-match accuracy for individually trained and hypernetwork-generated models, with and without task IDs, using the smallest 1.4K-parameter target model. Error bars show ±1 s.d. across 5 seeds.
Figure 6: Generated parameters form a task-structured weight space. t-SNE projection of generated target-model parameters, coloured by transformation. Left: With a frozen task embedding, tasks are fully separable (100% linear probe accuracy). Right: Without task identity, structure emerges from the support examples alone (60.9% linear probe accuracy).
Figure 7: Generated specialists generalise across episodes. Exact-match accuracy on the query paired with the generating support set and on queries from other episodes of the same transformation. Bars show the macro-average over 14 transformations with ±1 s.d.
Figure 8: Fill ∘ Mirror: Example compositional ARC-1D episode The support pairs specify the composition of Mirror followed by Fill, where both transformations are observed independently during training but their composition is held out. See Appendix H for more examples.
In-Distribution
Generalisation
Exact Match ↑
Token Accuracy ↑
Exact Match ↑
Compositional
Task ID
93.0±2.4
67.7±1.3
0.05±0.10
w/o Task ID
57.4±5.1 ( -35.6 )
75.5±1.1 ( +7.8 )
1.00±0.71 ( +0.95 )
Unseen
Task ID
97.0±0.5
68.5±4.8
0.29±0.57
Table 1: Task-ID conditioning improves in-distribution specialisation but hinders generalisation. Results for compositional and unseen transformations, reported as mean ± s.d. over 5 seeds.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Direct models
Hypernetworks
Muon learning rate
0.005
0.005
Muon momentum
0.95
0.95
AdamW learning rate
0.0005
0.001
Weight decay
0.01
0.01
Warmup steps
200
800
Cosine minimum LR
0
10−4
Appendix
Table 2: Optimisation hyperparameters for directly trained models and hypernetworks.
Hidden dimension
4
6
10
14
Direct model
Layers / heads
3 / 1
3 / 1
3 / 1
3 / 1
Token embedding dim.
10
10
10
10
Canon kernel
5
5
5
5
Parameters
1,398
2,444
5,400
9,508
Hypernetwork
Appendix
Table 3: Model configurations. Direct models are the three-layer RoPE Canon Transformer at four widths. The hypernetwork generates every weight of the target except the token embedder, which it shares with its encoder. Only the dimension-4 hypernetwork is trained; grey values marked with † show the corresponding parameter counts for wider targets, where only the decoder output layer grows.
Runs
Min./run
GPU-h
1
Capacity and ablations
960
3–9
64
2
Hypernetwork, in distribution
10
24
4
3
Weight reuse (evaluation only)
10
0.4
<1
4
Compositional generalisation (CPU evaluation)
10
0.7
0
5
Leave-one-out
140
23
54
6
Data efficiency
870
7–109
340
Appendix
Table 4: Compute per experiment. Wall-clock time per run and total compute on a single NVIDIA A40 GPU unless otherwise indicated. Times include data setup and evaluation.
In-Distribution
Generalisation
Exact Match ↑
Unseen
Compositional
Token Accuracy ↑
Exact Match ↑
Token Accuracy ↑
None
57.4±5.7
81.6±0.9
10.2±3.6
75.5±1.2
Latent conditioning
Frozen
93.0±2.6 ( +35.6 )
79.1±1.6 ( -2.5 )
0.6±0.5 ( -9.6 )
63.2±8.3 ( -12.3 )
Learned
96.5±0.8 ( +39.1 )
77.3±2.9 ( -4.3 )
1.1±1.9 ( -9.1 )
66.4±4.8 ( -9.1 )
Appendix
Table 5: Task-ID conditioning improves in-distribution specialisation but hinders generalisation. Results across task-ID placements and embedding types, reported as mean ± s.d. over 5 seeds. Parentheses show the change relative to no task-ID conditioning.
Figure 9: Task representations depend on task-ID placement. t-SNE of the task representation z for validation examples from each transformation (seed 1; axes are arbitrary). Latent conditioning produces clearer task-aligned structure, while input conditioning forms coarser groups.
Figure 10: Canon layers improve performance, particularly at low capacity. Test exact-match accuracy with and without Canon layers for individual and joint models across model sizes. The gain diminishes for larger individual models but remains substantial for jointly trained models. Bands show ±1 s.d. across 5 seeds.
Figure 11: Muon improves marginally over AdamW Test exact-match accuracy with Muon (solid) and AdamW (dashed) across model sizes. Bands show ±1 s.d. across 5 seeds.
Figure 12: The joint-training penalty is concentrated in a few transformations. Per-transformation exact-match validation accuracy at 1.4K parameters. Error bars show ±1 s.d. across 5 seeds.
Figure 13: The joint-training penalty is concentrated in a few transformations. Per-transformation exact-match validation accuracy at 2.4K parameters. Error bars show ±1 s.d. across 5 seeds.
Figure 14: The joint-training penalty is concentrated in a few transformations. Per-transformation exact-match validation accuracy at 5.4K parameters. Error bars show ±1 s.d. across 5 seeds.
Figure 15: The joint-training penalty is concentrated in a few transformations. Per-transformation exact-match validation accuracy at 9.5K parameters. Error bars show ±1 s.d. across 5 seeds.
Figure 16: Data efficiency across models. Individual models and task-ID-conditioned hypernetworks require similar amounts of training data, while the larger task-ID-conditioned joint model learns with fewer episodes. Without task IDs, joint models and hypernetworks follow similar curves. Bands show ±1 s.d. across 5 seeds.
Task ID
w/o Task ID
Composition
Token Acc.
EM
Token Acc.
EM
Mirror∘Denoise MC
47.7
0.0
85.6 ( +37.9 )
8.0
Copy∘Denoise MC
69.6
0.0
67.3 ( -2.3 )
0.0
Denoise 1C∘Denoise MC
76.7
0.0
68.3 ( -8.3 )
0.5
Mirror∘Fill
61.3
0.0
73.1 ( +11.7 )
1.5
Shift 3∘Denoise 1C
75.5
0.5
77.0 ( +1.5 )
0.0
Appendix
Table 6: Performance on held-out compositions. Each composition combines two transformations observed independently during training, while the composition itself is held out. B∘A denotes applying A first, followed by B . Token accuracy is macro-averaged over the 10 compositions and 5 seeds; exact match is pooled over all 2,000 held-out instances.
Task ID
w/o Task ID
Category
Token Acc.
EM
Token Acc.
EM
Denoise Multicolor
47.9
0.00
81.1 ( +33.2 )
4.00
Scaling
57.1
0.00
86.9 ( +29.8 )
0.00
Pattern Copy Multicolor
57.6
0.00
84.7 ( +27.2 )
0.00
Move 3 Pixels
68.4
0.00
91.8 ( +23.4 )
0.00
Pattern Copy
77.3
0.00
99.5 ( +22.2 )
96.00
Appendix
Table 7: Performance on held-out task categories. Each row holds out one of the 14 base transformations entirely during training and evaluates it zero-shot. Token accuracy and exact match are each the mean across 5 seeds for that category’s own held-out score.
Figure 17: Generated parameters form a task-structured weight space. PCA projection of generated target-model parameters, coloured by transformation. Left: With a frozen task embedding, tasks are fully separable (100% linear probe accuracy). Right: Without task identity, structure emerges from the support examples alone (60.9% linear probe accuracy).
Figure 18: Generated parameters form a task-structured weight space. UMAP projection of generated target-model parameters, coloured by transformation. Left: With a frozen task embedding, tasks are fully separable (100% linear probe accuracy). Right: Without task identity, structure emerges from the support examples alone (60.9% linear probe accuracy).
Figure 19: ARC-1D movement and scaling tasks. Examples of the movement and scaling task families, including one-point (1P), two-point (2P), two-point dynamic-position (2P-DP), three-point (3P), dynamic-position (DP), and scaling variants.
Figure 20: ARC-1D object-transformation tasks. Examples of the fill, hollow, flip, and mirror task families.
Figure 21: ARC-1D denoising and pattern-copy tasks. Examples of the denoising and pattern-copy task families, shown for both the single-color (1C) and multicolor (MC) variants.
Figure 22: ARC-1D compositional task examples. Examples of held-out compositions formed by combining two transformations observed independently during training.
Figure 23: ARC-1D compositional task examples. Further examples of held-out compositions formed from transformations observed independently during training.
Neural network checkpoints have quietly become a large-scale data resource: millions of trained weight vectors now exist, each encoding task-, domain-, and architecture-specific knowledge. This position paper argues that model checkpoints should be treated as a first-class data modality, and that generative modeling in weight space should be standardized as a core machine learning primitive. Recent advances demonstrate that neural weights can be synthesized on demand, often matching fine-tuning performance while reducing adaptation cost by orders of magnitude. We contend that these results reflect an underlying structural fact: high-performing models occupy low-dimensional, highly structured regions of weight space shaped by symmetry, flatness, modularity, and shared subspaces. Building on this view, we organize existing methods into a five-stage pipeline, survey applications where the approach is already practical, and clarify current limits: adapter-scale and conditional generation are advancing rapidly, while unrestricted frontier-scale checkpoint synthesis remains open. Our goal is to shift the community's default mindset from optimizing models per task to sampling models from learned weight distributions, accelerating toward an era in which AI systems routinely improve or create other AI systems.
Scaling laws hold that language models grow more capable with more parameters and more training data. Mixture-of-Experts (MoE) architectures are a remarkable demonstration of these laws, activating only a fraction of an enormous parameter bank for each token. But this success is built on static pretraining data --- the facts and corrections supplied by users during live interactions are a significant untapped source of potential improvement for a deployed model, but cannot be exploited by conventional architectures whose weights are frozen after training. Instead, this newfound knowledge must be placed in the context (by instruction or retrieval) and re-read on every request, only to be discarded afterwards. We seek instead to learn from live interactions by dynamically updating model weights. Inspired by MoEs, we propose the \textbf{Infinite-Parameter LLM}. A compact hypernetwork turns the online data into low-rank modulations of a shared base network, so feed-forward weights are generated from live data, not read from static memory. Whereas existing weight generators are held fixed after reading the context once, we form a Bayesian belief over the generator's latent state and update it online, such that the effective weights are re-derived as our belief evolves during the session. Although the model's memory footprint is constant, the feasible space of generated weights is thus effectively infinite. Representing live data in the weights rather than the prompt amortises compute, frees the context window, persists updates across turns, and can generalise better than in-context use. Our evaluation protocol applies this methodology to in-context learning and retrieval.
We propose SHINE (Scalable Hyper In-context NEtwork), a scalable hypernetwork that can map diverse meaningful contexts into high-quality LoRA adapters for large language models (LLMs). By reusing the frozen LLM's own parameters in an in-context hypernetwork design and introducing architectural innovations, SHINE overcomes key limitations of prior hypernetworks and achieves strong expressive power with a relatively small number of parameters. We introduce a pretraining and instruction fine-tuning pipeline, and train our hypernetwork to generate high quality LoRA adapters from diverse meaningful contexts in a single forward pass. It updates LLM parameters without any fine-tuning, and immediately enables complex question answering tasks related to the context without directly accessing the context, effectively transforming in-context knowledge to in-parameter knowledge in one pass. Our work achieves outstanding results on various tasks, greatly saves time, computation and memory costs compared to SFT-based LLM adaptation, and shows great potential for scaling. Our code is available at https://github.com/MuLabPKU/SHINE
Yewei Liu, Xiyuan Wang, Yansheng Mao +3
Institute for Artificial Intelligence, Peking University · School of Electronics Engineering and Computer Science, Peking University · University of Oxford +1