Learning to Predict Distributions over Weight Updates for Test-Time Adaptation
Organizations: University of Minnesota · IBM Research AI
Abstract
Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Regression Hypernet | Mixing Hypernet + Bank | |||||
|---|---|---|---|---|---|---|
| Base model | FullFT | LoRA | Det. | Dist. | Det. | Dist. |
| Qwen3-1.7B | 1,700M | 5.5M | 194M | 295M | 368.0M | 368.1M |
| Qwen3-4B | 4,000M | 10.6M | 245M | 396M | 624.1M | 624.2M |
| Gemma-3-1B | 1,000M | 5.0M | 192M | 291M | 343.3M | 343.4M |
| Llama-3.2-1B | 1,000M | 3.9M | 219M | 345M | 289.3M | 289.4M |
| Llama-3.2-3B | 3,000M | 7.6M | 233M | 372M | 472.4M | 472.5M |
| Qwen3-1.7B | ||||
|---|---|---|---|---|
| Approach | Rank 2 | Rank 8 | Rank 32 | Rank 128 |
| FullFT trainable | 1,700M | 1,700M | 1,700M | 1,700M |
| LoRA trainable / generated update | 1.4M | 5.5M | 22.0M | 88.1M |
| Deterministic hypernet trainable | 118M | 194M | 497M | 1,707M |
| Qwen3-4B | ||||
| Approach | Rank 2 | Rank 8 | Rank 32 | Rank 128 |
| Base model | Hypernet size | |||
|---|---|---|---|---|
| Qwen3-1.7B | 25M | 44.7 | 51.9 | 53.3 |
| 100M | 45.3 | 53.0 | 53.7 | |
| 500M | 44.5 | 53.2 | 53.9 | |
| 1B | 46.7 | 53.5 | 54.2 | |
| Qwen3-4B | 25M | 51.3 | 60.1 | 61.2 |
| 100M | 50.1 | 60.5 | 61.3 |
| Qwen3 1.7B | Qwen3 4B | Gemma 3 1B | Llama 3.2 1B | Llama 3.2 3B | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Regression | Mixing | Regression | Mixing | Regression | Mixing | Regression | Mixing | Regression | Mixing | |||||||||||
| Benchmark | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. | Det. | Dist. |
| GSM8K | 56.8 | 64.0 | 78.0 | 69.2 | 64.4 | 62.4 | 88.0 | 82.4 | 29.2 | 28.8 | 30.4 | 41.6 | 44.0 | 40.8 | 49.2 | 52.0 | 68.8 | 72.4 | 82.8 | 84.4 |
| MATH500 | 34.4 | 42.0 | 46.0 | 47.2 | 52.8 | 56.0 | 56.4 | 45.2 | 20.0 | 21.2 | 18.4 | 26.4 | 22.0 | 23.6 | 24.8 | 23.2 | 39.2 | 38.0 | 40.4 | 37.6 |
| HumanEval | 45.1 | 47.6 | 48.2 | 59.2 | 59.2 | 61.0 | 66.5 | 65.2 | 7.9 | 12.8 | 15.8 | 24.4 | 30.5 | 31.7 | 36.0 | 38.4 | 45.7 | 48.8 | 48.8 | 51.2 |
| ARC-C | 59.2 | 68.0 | 71.6 | 76.0 | 86.0 | 85.2 | 83.6 | 86.4 | 23.6 | 30.4 | 22.0 | 43.6 | 19.6 | 27.2 | 50.0 | 54.8 | 65.2 | 72.8 | 76.0 | 74.4 |
| Model | Method | Samp. | GSM8K | MATH | HumanEval | ARC-C | MMLU | MedQA | GPQA | Macro |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3 1.7B | Full FT | T | .508/.700/.768 | .316/.444/.540 | .529/.643/.714 | .608/.740/.788 | .496/.664/.700 | .272/.424/.428 | .217/.273/.354 | .421/.555/.613 |
| LoRA | T | .568/.620/.632 | .316/.428/.500 | .667/.607/.635 | .520/.760/.800 | .492/.664/.708 | .172/.336/.372 | .177/.283/.268 | .416/.528/.559 | |
| MoL | T | .468/.560/.608 | .292/.428/.480 | .707/.701/.739 | .596/.704/.740 | .512/.688/.684 | .112/.320/.352 | .202/.293/.263 | .413/.528/.552 | |
| Reg. Det. | T | .524/.588/.612 | .312/.420/.504 | .633/.702/.738 | .528/.736/.748 | .496/.676/.688 | .196/.356/.368 | .207/.303/.308 | .414/.540/.567 | |
| Reg. Dist. | W | .620/.684/.740 | .396/.528/.556 | .688/.705/.706 | .708/.760/.784 | .608/.692/.720 | .308/.444/.428 | .197/.278/.258 | .504/.584/.599 | |
| Mix. Det. | T | .776/.844/.864 | .456/.528/.552 | .739/.781/.791 | .648/.748/.752 | .608/.744/.760 | .280/.388/.416 | .212/.298/.293 | .531/.619/.633 |