Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert's update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations never see the output factor. If each expert is split into a reader (down-projection) and a writer (up-projection), the usefulness of every reader--writer pair becomes an inner product in the shared rank-r space, and all NANB pairs can be scored from NA+NB vectors. We build VANE on this identity. A low-rank compass predicts the descent direction of each token. VANE scores every pair by the alignment between its update and the compass without forming any update, activates the top-k pairs with additive gates, and gives every pair its exact first-order router gradient. On single-domain commonsense reasoning and a four-domain multi-task mixture with Llama-3.2-3B and Llama-3.1-8B, VANE attains the best average among twelve PEFT and MoE-LoRA baselines, by 0.9--1.1 and 1.3--1.5 points respectively, with less than half the trainable parameters of an 8-expert MoE-LoRA. Its router scores also track the measured usefulness of experts far more closely than token routers do.
Figures & tables
Figure 1: What a router can see (exact computations in a controlled linear-response model with 64 reader–writer pairs; Appendix E ). (a) Squared correlation between router scores and the expected usefulness E[uab∣x] as the token mean grows. Token routers follow the ceiling of Theorem 1 (line: closed form; dots: least-squares fit) and see nothing for centered tokens. Reader norms are writer-blind. The update score is exact, and Vane ’s calibrated version stays at 0.87 for every κ . (b) Routers trained online with the same dense first-order signal: held-out fraction of the attainable first-order gain captured by their top-4 pairs. (c) FLOPs per token to score every pair at width 4096. Vane scores 1024 pairs 77× more cheaply than a linear router.
Figure 2: Vane for one token. Readers produce rank- r latents. A low-rank compass forecasts the descent direction at the layer output, and each writer turns it into a query in the same latent space. Every reader–writer pair is scored by the alignment between its update and the compass without materializing the update. The top- k pairs are summed with additive gates. In the backward pass, the exact first-order usefulness of every pair comes from the same latents and gives the router a dense gradient.
Method
Param %
BoolQ
PIQA
SIQA
HellaS.
WinoG.
ARC-e
ARC-c
OBQA
Avg.
Δ
Llama-3.2-3B
LoRA
0.76
67.9
83.2
76.8
88.6
79.4
84.7
71.3
79.6
78.9
–
DoRA
0.78
68.5
83.8
77.4
88.9
80.1
85.1
72.2
80.2
79.5
+0.6
MoSLoRA
0.76
68.5
83.5
77.3
88.9
80.0
85.1
72.0
80.2
79.4
+0.5
CoMoL
0.77
69.2
84.1
77.8
89.4
80.6
85.4
73.0
80.7
80.0
+1.1
FlyLoRA
0.77
69.1
84.2
78.0
89.3
80.8
85.8
73.5
81.1
80.2
+1.3
Table 1: Single-domain commonsense reasoning (accuracy, %, mean of three seeds; trained on Commonsense170K). Param % : trainable parameters relative to the backbone. Δ : average gain over LoRA. Groups: single-adapter PEFT, mixtures of LoRA experts, ours. Bold / underline : best/second best.
Method
Param %
GSM8K
MATH
HumanE.
MBPP
MedMC.
PubMed.
ARC-c
WinoG.
Avg.
Δ
Llama-3.2-3B
LoRA
0.76
52.7
12.6
36.0
44.8
50.2
72.4
69.5
73.8
51.5
–
DoRA
0.78
53.6
13.3
37.0
45.5
50.9
72.9
70.0
74.3
52.2
+0.7
MoSLoRA
0.76
53.6
13.3
36.8
45.5
50.7
72.8
70.0
74.1
52.1
+0.6
CoMoL
0.77
54.8
14.0
38.0
46.3
51.3
73.3
71.1
75.2
53.0
+1.5
FlyLoRA
0.77
55.2
14.7
39.5
46.8
51.9
73.8
71.1
75.1
53.5
+2.0
Table 2: Multi-task adaptation : one adapter trained on a four-domain mixture (math, code, medicine, commonsense) and evaluated on two benchmarks per domain (accuracy or pass@1, %, mean of three seeds). Columns as in Table 1 .
Configuration
CS
MT
ΔMT
ρ
Vane (full)
86.9
63.9
–
0.61
Router score
Token router over pairs
86.0
62.5
−1.4
0.22
Product router
85.6
61.9
−2.0
0.12
Latent-linear router
86.1
62.6
−1.3
0.24
Reader-norm self-routing
85.7
62.0
−1.9
0.06
Table 3: Ablations on Llama-3.1-8B. CS/MT: single-domain and multi-task averages. ρ : Spearman correlation between router scores and the exact first-order usefulness of all 32 pairs on held-out multi-task tokens.
Figure 3: How Vane routes (Llama-3.1-8B, multi-task, held-out tokens). (a) Spearman correlation between router scores and the exact first-order usefulness of all 32 pairs, measured in the backward pass ( Vane : ±1 s.d. over seeds). (b) Cosine between the compass and the true descent direction, overall and after projection onto the span of the candidate updates. (c) Fraction of the oracle top- k first-order gain captured by the selected pairs. (d) Pair usage (readers × writers) per domain, averaged over layers.
Figure 4: Efficiency, structure, sparsity, and scale (multi-task average). (a) Accuracy vs. trainable parameters on Llama-3.1-8B (gray: single adapters; blue: mixtures). (b) Readers × writers at r=4 (box: default). (c) Number of active pairs k out of 32 (dotted: best baseline). (d) Gain over LoRA across Llama scales.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
G(x)
r
Router score
LoRA
[1]
r
–
Paired MoE-LoRA
diagonal
r
token
HydraLoRA
one column
r
token
MoSLoRA
fixed dense mixer
r
–
CoMoL
routed dense core
r
latent-linear
SMoRA, FlyLoRA
diagonal
1
token, reader
Appendix
Table 4: Low-rank mixtures as instances of Eq. ( 6 ). “Reader” scores depend on the token only through the input-side activations; “update” scores depend on the writer and on a forecast of the loss gradient.
Check
Result
Latent alignment vs. materialized ⟨e,BbAax⟩
2.7×10−15
Latent calibrated score vs. materialized norm
6.7×10−16
Loss change vs. −ηuab ( η=10−5 ), rel. error
4.9×10−4
Autograd ∂ℓ/∂sab vs. −rασ′uab , all pairs, rel.
2.1×10−16
Gradient to unselected pairs without the ST core
0
cos(Δe,δ) after a step on U : min / mean (200 draws)
+0.224 / +0.486
Appendix
Table 5: Numerical verification of Propositions 1 – 4 and Theorem 1 (output of code/theory_checks.py ). The 10−5 score change comes from the ϵ in Eq. ( 4 ).
Figure 5: Controlled linear-response study, continued (exact computations). (a) Squared correlation with expected usefulness vs. compass rank. The raw update score is exact once re≥rankΓ=6 , and calibration caps it at 0.87 . (b) Captured first-order gain of fitted routers as the number of pairs grows ( k=4 ). (c) Captured gain vs. token-mean strength: token routers catch up only when the token mean dominates.
Table 10: Single-domain (CS) and multi-task (MT) averages on Qwen backbones.
NA\NB
2
4
8
16
1
61.2 (0.46%)
62.3 (0.60%)
63.1 (0.87%)
63.3 (1.42%)
2
61.7 (0.52%)
62.8 (0.66%)
63.5 (0.93%)
63.6 (1.48%)
4
62.0 (0.65%)
63.1 (0.78%)
63.9 (1.06%)
64.0 (1.61%)
8
62.1 (0.89%)
63.2 (1.03%)
63.8 (1.31%)
63.9 (1.85%)
Appendix
Table 11: Multi-task average on Llama-3.1-8B over readers NA and writers NB ( r=4 , re=8 , k=4 ); trainable share in parentheses.
top- k
1
2
4
8
16
32
Vane
62.6
63.4
63.9
63.7
63.2
62.4
Comb. + token router
61.6
62.2
62.5
62.3
61.9
61.3
compass rank re
1
2
4
8
16
32
MT avg.
62.9
63.3
63.7
63.9
64.0
63.9
ρ
0.38
0.47
0.56
0.61
0.63
0.63
Appendix
Table 12: Multi-task average on Llama-3.1-8B vs. active pairs k (top) and compass rank re (bottom, with the score–usefulness correlation ρ ).
λbal
0
10−3
10−2
10−1
MT avg.
62.8
63.6
63.9
63.4
writer init std
0 (+ tie noise)
10−5
10−4
10−3
MT avg.
63.5
63.8
63.9
63.6
Appendix
Table 13: Sensitivity to the balancing weight and the writer initialization (multi-task average, Llama-3.1-8B).
Method
Params
Train
Mem.
Decode
MT
(M)
(h)
(GB)
(ms/tok)
LoRA
41.9
3.1
38.2
24.1
60.4
HydraLoRA
108.0
3.9
41.0
27.9
61.9
MixLoRA
125.8
5.2
44.7
33.6
62.0
GOAT
177.7
5.6
46.1
34.8
62.4
LD-MoLE
177.7
5.9
46.9
36.2
62.6
Appendix
Table 14: Cost of multi-task training on Llama-3.1-8B (8 H100): trainable parameters, training hours, peak memory per GPU, unmerged decoding latency (batch 1), and multi-task average.
Data used
10%
25%
50%
100%
LoRA
54.9
57.6
59.3
60.4
LD-MoLE
55.8
59.1
61.2
62.6
Vane
57.4
60.8
62.7
63.9
Appendix
Table 15: Multi-task average on Llama-3.1-8B when training on a fraction of the mixture.
MMLU
MT avg. by placement
Method
(0-shot)
Attn.
MLP
All
Base model
65.0
–
–
–
LoRA
63.1
59.3
60.0
60.4
MixLoRA
63.8
–
–
–
LD-MoLE
64.0
–
–
–
Vane
64.6
62.1
63.2
63.9
Appendix
Table 16: Zero-shot MMLU after multi-task training (knowledge retention) and multi-task average when adapting only attention, only MLP, or all projections (Llama-3.1-8B).
Score class
ρ
Captured gain
Reader-norm self-routing
0.06
0.04
Product router
0.12
0.17
Token router over pairs
0.22
0.41
Latent-linear router
0.24
0.43
Vane (update score)
0.61
0.78
Appendix
Table 17: Spearman correlation with the exact usefulness of all 32 pairs and captured first-order gain of the top-4 pairs (Llama-3.1-8B, multi-task, held-out tokens). All routers use Vane ’s reader–writer structure.
Figure 6: Training dynamics and gate statistics (Llama-3.1-8B, multi-task). (a) Training loss. (b) Held-out multi-task average during training. (c) Rank of the per-token update among the top-4 pairs, in units of r . Most tokens combine pairs with distinct readers and writers and reach rank 4r (Proposition 4 ). (d) Sum of the active gates per token. Complementary pairs stack well above the softmax constraint of 1.
Figure 7: Where the compass points (Llama-3.1-8B, multi-task). (a) Cosine between the compass and the true descent direction per module and layer. Alignment is highest in the middle layers and in the MLP projections. (b) Share of selected pairs per reader and writer. Marginal balancing keeps every factor in use while pair usage stays specialized (Figure 3 d); without it, two readers and three writers absorb most of the traffic.
Department of Computer Science and Engineering, POSTECH, South Korea · Microsoft Research Asia, Beijing, China · Graduate School of Artificial Intelligence, POSTECH, South Korea