Instruction-tuned models are deployed into environments where domains are heterogeneous and evolve, yet adding new domains or data typically requires costly retraining. We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters. The framework automatically discovers latent domains, uses them to train per-domain Low-Rank Adaptation (LoRA) adapters independently in parallel and performs parameter-free routing. Across 14 domain-specific benchmarks and GPT-4o pairwise judgements, AutoAdapt achieves parity with a LoRA adapter trained on all domains without requiring full-model retraining. We also find evidence of specialisation effect convergence across independent discovery methods. Overall, training each adapter on its own domain prevents domain interference by construction, thus enabling modular, taxonomy-free domain specialisation without aggregate performance loss or full model retraining.
Figures & tables
Dataset
Domain
#Samples
Avg Tokens
Alpaca-cleaned
General
51,760
∼ 120
CodeAlpaca-20K
Code
20,022
∼ 40
WebGPT
QA/Retrieval
18,994
∼ 169
Databricks Dolly 15K
General
15,011
∼ 129
Finance-Alpaca
Finance
68,912
∼ 162
AlpaCare-MedInstruct
Medicine
52,002
∼ 184
Table 1: Composition and statistics of the instruction-tuning dataset.
Figure 1: Domain overlap between discovery methods on LLM-as-a-judge evaluation with GPT-4o.
Benchmark
K-means
MNLI-24
BERTopic
Agreement
GSM8K
+2.50
+0.90
+4.00
WIN (all)
Spider
+3.30
+2.50
+2.80
WIN (all)
MedMCQA
+0.70
+5.20
+4.00
WIN (all)
BoolQ
-0.60
-1.60
-2.50
LOSE (all)
MultiNLI
-7.80
-8.50
-13.70
LOSE (all)
HumanEval
-7.32
-4.27
-4.27
LOSE (all)
Table 2: Cross-method benchmark deltas (percentage points vs. single LoRA). Sample size is N=1000 for all datasets except HumanEval ( N=164 ), MBPP ( N=500 ), and CRUXEval ( N=800 ). K-means ( −0.27 pp, 95% CI [−1.36,+0.82] , p=0.63 ) and MNLI ( +0.19 pp, 95% CI [−0.90,+1.28] , p=0.73 ) methods do not differ significantly from single LoRA. BERTopic shows a small statistically significant negative aggregate effect ( −2.18 pp, 95% CI [−3.28,−1.08] , p<0.001 ). CIs and p -values use an inverse-variance weighted mean of per-benchmark deltas.
Held-out LM
GPT-4o judge
Model
NLL ↓
PPL ↓
Upd. wins
Other wins
Ties
Updated c15 (+ Magicoder)
0.1926
1.212
–
–
–
Old c15 (no Magicoder)
0.2805
1.324
104
60
36
Single LoRA (frozen)
0.2813
1.325
119
50
30
Single LoRA (retrained)
0.2024
1.224
85
76
38
Full FT (retrained)
0.4488
1.566
167
16
17
Table 3: Results of the post-deployment domain update scenario simulations. We report the average negative log-likelihood (NLL) and perplexity (PPL) as held-out language modelling metrics. We also report LLM-as-a-judge results with GPT-4o.
Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method that places trainable low-rank adapters into frozen pre-trained models. Recent studies show that using fewer LoRA adapters may still maintain or even improve performance, but existing methods still distribute adapters broadly, leaving \emph{where to place a limited number of adapters to maximize performance} largely open. To investigate this, we introduce \textbf{PAGE} (\textbf{P}rojected \textbf{A}dapter \textbf{G}radient \textbf{E}nergy), a gradient-based sensitivity probe that estimates the initial trainable gradient energy available to each candidate LoRA adapter. Surprisingly, we find that PAGE is highly concentrated on a single shallow FFN down-projection across two model families and four downstream tasks. We term this module the \textbf{dominant adaptation module} and show that its layer index is architecture-dependent but task-stable. Motivated by this finding, we propose \textbf{DomLoRA}, a placement method that places a single adapter at the dominant adaptation module. With only \textbf{0.7%} of vanilla LoRA's trainable parameters, DomLoRA outperforms it on average across downstream tasks, including instruction following, mathematical reasoning, coding, and multi-turn conversation. This method also matches or improves other LoRA variants and reduces training time by up to \textbf{2.74}× compared with broad placement, supporting the dominant adaptation module perspective as a practical placement guideline.
Suoxin Zhang, Run He, Di Fang +3
South China University of Technology, China · Zhejiang University, China
Low-Rank Adaptation (LoRA) has become the standard tool for parameter-efficient fine-tuning of large pretrained models. When applied sequentially across tasks in Continual Learning (CL), the standard assumption is that each new task requires a dedicated low-rank adapter. In this work, we challenge this assumption empirically and structurally. We show that task-specific LoRA adapters in CL exhibit significant low-rank redundancy: the subspaces spanned by adapters trained on different tasks substantially overlap, and in many cases earlier adapters can faithfully represent later tasks. Building on this observation, we propose LiteLoRA, a plug-and-play gating mechanism that learns at train time whether to recruit a new adapter or reuse existing low-rank representations. Our method reduces the number of active adapters by 20-70% while matching or exceeding state-of-the-art performance on standard CL benchmarks, revealing that structural redundancy is pervasive and that selective learning is sufficient to achieve stability without sacrificing plasticity.
We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction, code, and natural text. Using singular value decomposition of weight matrices, we find that CPT leaves singular value spectra largely invariant, with adaptation driven mainly by changes in singular vectors. An analysis of attention-head projection matrices reveals strong, domain-dependent head heterogeneity, which we exploit to define a head importance criterion: up to 60% of head updates can be removed without measurable quality loss. Selectively rewinding low-importance heads to their pre-trained state improves benchmark accuracy by up to 4% versus the fully trained baseline. Finally, we identify domain connectivity - linear interpolation between CPT checkpoints yields smooth domain-quality interpolation without notable degradation on either domain - and release Diffract, an open-source toolkit for scalable spectral analysis of billion-parameter models.
Nikita Borodin, Maria Krylova, Artem Zabolotnyi +6