Instruction-tuned models are deployed into environments where domains are heterogeneous and evolve, yet adding new domains or data typically requires costly retraining. We present AutoAdapt, a modular framework that incorporates new domains and data via targeted single-adapter training without modifying other adapters. The framework automatically discovers latent domains, uses them to train per-domain Low-Rank Adaptation (LoRA) adapters independently in parallel and performs parameter-free routing. Across 14 domain-specific benchmarks and GPT-4o pairwise judgements, AutoAdapt achieves parity with a LoRA adapter trained on all domains without requiring full-model retraining. We also find evidence of specialisation effect convergence across independent discovery methods. Overall, training each adapter on its own domain prevents domain interference by construction, thus enabling modular, taxonomy-free domain specialisation without aggregate performance loss or full model retraining.
Figures & tables
Dataset
Domain
#Samples
Avg Tokens
Alpaca-cleaned
General
51,760
∼ 120
CodeAlpaca-20K
Code
20,022
∼ 40
WebGPT
QA/Retrieval
18,994
∼ 169
Databricks Dolly 15K
General
15,011
∼ 129
Finance-Alpaca
Finance
68,912
∼ 162
AlpaCare-MedInstruct
Medicine
52,002
∼ 184
Table 1: Composition and statistics of the instruction-tuning dataset.
Figure 1: Domain overlap between discovery methods on LLM-as-a-judge evaluation with GPT-4o.
Benchmark
K-means
MNLI-24
BERTopic
Agreement
GSM8K
+2.50
+0.90
+4.00
WIN (all)
Spider
+3.30
+2.50
+2.80
WIN (all)
MedMCQA
+0.70
+5.20
+4.00
WIN (all)
BoolQ
-0.60
-1.60
-2.50
LOSE (all)
MultiNLI
-7.80
-8.50
-13.70
LOSE (all)
HumanEval
-7.32
-4.27
-4.27
LOSE (all)
Table 2: Cross-method benchmark deltas (percentage points vs. single LoRA). Sample size is N=1000 for all datasets except HumanEval ( N=164 ), MBPP ( N=500 ), and CRUXEval ( N=800 ). K-means ( −0.27 pp, 95% CI [−1.36,+0.82] , p=0.63 ) and MNLI ( +0.19 pp, 95% CI [−0.90,+1.28] , p=0.73 ) methods do not differ significantly from single LoRA. BERTopic shows a small statistically significant negative aggregate effect ( −2.18 pp, 95% CI [−3.28,−1.08] , p<0.001 ). CIs and p -values use an inverse-variance weighted mean of per-benchmark deltas.
Held-out LM
GPT-4o judge
Model
NLL ↓
PPL ↓
Upd. wins
Other wins
Ties
Updated c15 (+ Magicoder)
0.1926
1.212
–
–
–
Old c15 (no Magicoder)
0.2805
1.324
104
60
36
Single LoRA (frozen)
0.2813
1.325
119
50
30
Single LoRA (retrained)
0.2024
1.224
85
76
38
Full FT (retrained)
0.4488
1.566
167
16
17
Table 3: Results of the post-deployment domain update scenario simulations. We report the average negative log-likelihood (NLL) and perplexity (PPL) as held-out language modelling metrics. We also report LLM-as-a-judge results with GPT-4o.