cs.LGOct 1, 2026

Learning to Predict Distributions over Weight Updates for Test-Time Adaptation

Authors: Azal Ahmad Khan, Keshav Ramji, Tahira Naseem, Ali Anwar, Ramón Fernandez Astudillo

Organizations: University of Minnesota · IBM Research AI

Abstract

Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

    Jul 21, 2026Nischay Dhankhar, Dos Baha, Abulhair SaparovHypernetworkScaling Laws

  2. SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

    Feb 6, 2026Yewei Liu, Xiyuan Wang, Yansheng Mao +3Large Language Model AdaptationLow-Rank Adaptation Adapters

  3. Query-Conditioned Test-Time Self-Training for Large Language Models

    May 13, 2026Chaehee Song, Minseok Seo, Yeeun Seong +2Test-Time TrainingTest-Time Scaling