Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them route by scoring the token: the router picks experts without seeing what any of them would write. We argue that the router should score the update. To first order, adding an expert's update to a layer output lowers the loss by the inner product between that update and the negative loss gradient at the output. This usefulness is quadratic in the token, so a router that is linear in the token sees only the part of it that runs through the token mean, and routers that rank experts by the norm of their own activations never see the output factor. If each expert is split into a reader (down-projection) and a writer (up-projection), the usefulness of every reader--writer pair becomes an inner product in the shared rank-r space, and all NANB pairs can be scored from NA+NB vectors. We build VANE on this identity. A low-rank compass predicts the descent direction of each token. VANE scores every pair by the alignment between its update and the compass without forming any update, activates the top-k pairs with additive gates, and gives every pair its exact first-order router gradient. On single-domain commonsense reasoning and a four-domain multi-task mixture with Llama-3.2-3B and Llama-3.1-8B, VANE attains the best average among twelve PEFT and MoE-LoRA baselines, by 0.9--1.1 and 1.3--1.5 points respectively, with less than half the trainable parameters of an 8-expert MoE-LoRA. Its router scores also track the measured usefulness of experts far more closely than token routers do.
Figures & tables
Figure 1: What a router can see (exact computations in a controlled linear-response model with 64 reader–writer pairs; Appendix E ). (a) Squared correlation between router scores and the expected usefulness E[uab∣x] as the token mean grows. Token routers follow the ceiling of Theorem 1 (line: closed form; dots: least-squares fit) and see nothing for centered tokens. Reader norms are writer-blind. The update score is exact, and Vane ’s calibrated version stays at 0.87 for every κ . (b) Routers trained online with the same dense first-order signal: held-out fraction of the attainable first-order gain captured by their top-4 pairs. (c) FLOPs per token to score every pair at width 4096. Vane scores 1024 pairs 77× more cheaply than a linear router.
Figure 2: Vane for one token. Readers produce rank- r latents. A low-rank compass forecasts the descent direction at the layer output, and each writer turns it into a query in the same latent space. Every reader–writer pair is scored by the alignment between its update and the compass without materializing the update. The top- k pairs are summed with additive gates. In the backward pass, the exact first-order usefulness of every pair comes from the same latents and gives the router a dense gradient.
Method
Param %
BoolQ
PIQA
SIQA
HellaS.
WinoG.
ARC-e
ARC-c
OBQA
Avg.
Δ
Llama-3.2-3B
LoRA
0.76
67.9
83.2
76.8
88.6
79.4
84.7
71.3
79.6
78.9
–
DoRA
0.78
68.5
83.8
77.4
88.9
80.1
85.1
72.2
80.2
79.5
+0.6
MoSLoRA
0.76
68.5
83.5
77.3
88.9
80.0
85.1
72.0
80.2
79.4
+0.5
CoMoL
0.77
69.2
84.1
77.8
89.4
80.6
85.4
73.0
80.7
80.0
+1.1
FlyLoRA
0.77
69.1
84.2
78.0
89.3
80.8
85.8
73.5
81.1
80.2
+1.3
Table 1: Single-domain commonsense reasoning (accuracy, %, mean of three seeds; trained on Commonsense170K). Param % : trainable parameters relative to the backbone. Δ : average gain over LoRA. Groups: single-adapter PEFT, mixtures of LoRA experts, ours. Bold / underline : best/second best.
Method
Param %
GSM8K
MATH
HumanE.
MBPP
MedMC.
PubMed.
ARC-c
WinoG.
Avg.
Δ
Llama-3.2-3B
LoRA
0.76
52.7
12.6
36.0
44.8
50.2
72.4
69.5
73.8
51.5
–
DoRA
0.78
53.6
13.3
37.0
45.5
50.9
72.9
70.0
74.3
52.2
+0.7
MoSLoRA
0.76
53.6
13.3
36.8
45.5
50.7
72.8
70.0
74.1
52.1
+0.6
CoMoL
0.77
54.8
14.0
38.0
46.3
51.3
73.3
71.1
75.2
53.0
+1.5
FlyLoRA
0.77
55.2
14.7
39.5
46.8
51.9
73.8
71.1
75.1
53.5
+2.0
Table 2: Multi-task adaptation : one adapter trained on a four-domain mixture (math, code, medicine, commonsense) and evaluated on two benchmarks per domain (accuracy or pass@1, %, mean of three seeds). Columns as in Table 1 .
Configuration
CS
MT
ΔMT
ρ
Vane (full)
86.9
63.9
–
0.61
Router score
Token router over pairs
86.0
62.5
−1.4
0.22
Product router
85.6
61.9
−2.0
0.12
Latent-linear router
86.1
62.6
−1.3
0.24
Reader-norm self-routing
85.7
62.0
−1.9
0.06
Table 3: Ablations on Llama-3.1-8B. CS/MT: single-domain and multi-task averages. ρ : Spearman correlation between router scores and the exact first-order usefulness of all 32 pairs on held-out multi-task tokens.
Figure 3: How Vane routes (Llama-3.1-8B, multi-task, held-out tokens). (a) Spearman correlation between router scores and the exact first-order usefulness of all 32 pairs, measured in the backward pass ( Vane : ±1 s.d. over seeds). (b) Cosine between the compass and the true descent direction, overall and after projection onto the span of the candidate updates. (c) Fraction of the oracle top- k first-order gain captured by the selected pairs. (d) Pair usage (readers × writers) per domain, averaged over layers.
Figure 4: Efficiency, structure, sparsity, and scale (multi-task average). (a) Accuracy vs. trainable parameters on Llama-3.1-8B (gray: single adapters; blue: mixtures). (b) Readers × writers at r=4 (box: default). (c) Number of active pairs k out of 32 (dotted: best baseline). (d) Gain over LoRA across Llama scales.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
G(x)
r
Router score
LoRA
[1]
r
–
Paired MoE-LoRA
diagonal
r
token
HydraLoRA
one column
r
token
MoSLoRA
fixed dense mixer
r
–
CoMoL
routed dense core
r
latent-linear
SMoRA, FlyLoRA
diagonal
1
token, reader
Appendix
Table 4: Low-rank mixtures as instances of Eq. ( 6 ). “Reader” scores depend on the token only through the input-side activations; “update” scores depend on the writer and on a forecast of the loss gradient.
Check
Result
Latent alignment vs. materialized ⟨e,BbAax⟩
2.7×10−15
Latent calibrated score vs. materialized norm
6.7×10−16
Loss change vs. −ηuab ( η=10−5 ), rel. error
4.9×10−4
Autograd ∂ℓ/∂sab vs. −rασ′uab , all pairs, rel.
2.1×10−16
Gradient to unselected pairs without the ST core
0
cos(Δe,δ) after a step on U : min / mean (200 draws)
+0.224 / +0.486
Appendix
Table 5: Numerical verification of Propositions 1 – 4 and Theorem 1 (output of code/theory_checks.py ). The 10−5 score change comes from the ϵ in Eq. ( 4 ).
Figure 5: Controlled linear-response study, continued (exact computations). (a) Squared correlation with expected usefulness vs. compass rank. The raw update score is exact once re≥rankΓ=6 , and calibration caps it at 0.87 . (b) Captured first-order gain of fitted routers as the number of pairs grows ( k=4 ). (c) Captured gain vs. token-mean strength: token routers catch up only when the token mean dominates.
Table 10: Single-domain (CS) and multi-task (MT) averages on Qwen backbones.
NA\NB
2
4
8
16
1
61.2 (0.46%)
62.3 (0.60%)
63.1 (0.87%)
63.3 (1.42%)
2
61.7 (0.52%)
62.8 (0.66%)
63.5 (0.93%)
63.6 (1.48%)
4
62.0 (0.65%)
63.1 (0.78%)
63.9 (1.06%)
64.0 (1.61%)
8
62.1 (0.89%)
63.2 (1.03%)
63.8 (1.31%)
63.9 (1.85%)
Appendix
Table 11: Multi-task average on Llama-3.1-8B over readers NA and writers NB ( r=4 , re=8 , k=4 ); trainable share in parentheses.
top- k
1
2
4
8
16
32
Vane
62.6
63.4
63.9
63.7
63.2
62.4
Comb. + token router
61.6
62.2
62.5
62.3
61.9
61.3
compass rank re
1
2
4
8
16
32
MT avg.
62.9
63.3
63.7
63.9
64.0
63.9
ρ
0.38
0.47
0.56
0.61
0.63
0.63
Appendix
Table 12: Multi-task average on Llama-3.1-8B vs. active pairs k (top) and compass rank re (bottom, with the score–usefulness correlation ρ ).
λbal
0
10−3
10−2
10−1
MT avg.
62.8
63.6
63.9
63.4
writer init std
0 (+ tie noise)
10−5
10−4
10−3
MT avg.
63.5
63.8
63.9
63.6
Appendix
Table 13: Sensitivity to the balancing weight and the writer initialization (multi-task average, Llama-3.1-8B).
Method
Params
Train
Mem.
Decode
MT
(M)
(h)
(GB)
(ms/tok)
LoRA
41.9
3.1
38.2
24.1
60.4
HydraLoRA
108.0
3.9
41.0
27.9
61.9
MixLoRA
125.8
5.2
44.7
33.6
62.0
GOAT
177.7
5.6
46.1
34.8
62.4
LD-MoLE
177.7
5.9
46.9
36.2
62.6
Appendix
Table 14: Cost of multi-task training on Llama-3.1-8B (8 H100): trainable parameters, training hours, peak memory per GPU, unmerged decoding latency (batch 1), and multi-task average.
Data used
10%
25%
50%
100%
LoRA
54.9
57.6
59.3
60.4
LD-MoLE
55.8
59.1
61.2
62.6
Vane
57.4
60.8
62.7
63.9
Appendix
Table 15: Multi-task average on Llama-3.1-8B when training on a fraction of the mixture.
MMLU
MT avg. by placement
Method
(0-shot)
Attn.
MLP
All
Base model
65.0
–
–
–
LoRA
63.1
59.3
60.0
60.4
MixLoRA
63.8
–
–
–
LD-MoLE
64.0
–
–
–
Vane
64.6
62.1
63.2
63.9
Appendix
Table 16: Zero-shot MMLU after multi-task training (knowledge retention) and multi-task average when adapting only attention, only MLP, or all projections (Llama-3.1-8B).
Score class
ρ
Captured gain
Reader-norm self-routing
0.06
0.04
Product router
0.12
0.17
Token router over pairs
0.22
0.41
Latent-linear router
0.24
0.43
Vane (update score)
0.61
0.78
Appendix
Table 17: Spearman correlation with the exact usefulness of all 32 pairs and captured first-order gain of the top-4 pairs (Llama-3.1-8B, multi-task, held-out tokens). All routers use Vane ’s reader–writer structure.
Figure 6: Training dynamics and gate statistics (Llama-3.1-8B, multi-task). (a) Training loss. (b) Held-out multi-task average during training. (c) Rank of the per-token update among the top-4 pairs, in units of r . Most tokens combine pairs with distinct readers and writers and reach rank 4r (Proposition 4 ). (d) Sum of the active gates per token. Complementary pairs stack well above the softmax constraint of 1.
Figure 7: Where the compass points (Llama-3.1-8B, multi-task). (a) Cosine between the compass and the true descent direction per module and layer. Alignment is highest in the middle layers and in the MLP projections. (b) Share of selected pairs per reader and writer. Marginal balancing keeps every factor in use while pair usage stays specialized (Figure 3 d); without it, two readers and three writers absorb most of the traffic.
Mixture-of-Experts (MoE) variants of Low-Rank Adaptation (LoRA) route every token to a fixed number of experts k. Tokens differ in how uncertain the model is about them, so a single k over-spends on easy tokens and under-serves hard ones. We observe that the router's output distribution is already a per-token uncertainty signal: peaked mass indicates confidence, while a flat distribution indicates ambiguity. We introduce CARE (Confidence-Adaptive Routing of Experts), which admits experts in a nucleus fashion. Experts are activated in decreasing router weight until their cumulative mass reaches a threshold, with a small extension when the admitted experts disagree. A budget thermostat calibrates the threshold so that the average number of active experts matches any target. CARE is a drop-in, single-forward-pass rule with no extra parameters. Across eight commonsense benchmarks on LLaMA-3.1-8B and Qwen2.5-7B, as well as math, code, and knowledge tasks, CARE improves over fixed top-k MoE-LoRA at matched compute and matches the fixed-k=4 baseline while activating fewer experts. The same confidence and disagreement signals also improve out-of-distribution detection over MSP, entropy, and multi-pass proxies. We support the design with nucleus fidelity, budget optimality, and an epistemic reading of disagreement, and we release code.
Tom Saliencro, Rohan Desai, Priya Nair +2
University of California, Irvine · University of Washington
Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8×7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.
Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2
Data Science Institute, Columbia University, New York, NY, USA · Department of Computer Science, Columbia University, New York, NY, USA
Mixture-of-Experts (MoE) language models route each token to a small subset of experts, but whether the routes selected by a trained top-k router are good ones is rarely evaluated directly. Holding the model fixed, we compare each standard route against sampled equal-compute alternatives for the same token and score each by the next-token probability it assigns to the realized token in a verified reasoning trajectory. The result is sharply token-conditional: the standard router is well-aligned with route utility on confident tokens but uninformative on the fragile tokens that drive hard reasoning, where lower-loss equal-compute routes consistently exist inside the frozen model but are not selected. The same pattern holds across Qwen3-30B-A3B, GPT-OSS-20B, DeepSeek-V2-Lite, and OLMoE-1B-7B, and follows structurally from how standard top-k training evaluates routing decisions: the language modeling loss scores only the executed route, and load balancing depends only on aggregate routing statistics. A minimal router-only update to the final-layer router, leaving every expert and every other router frozen, is sufficient to shift pass@K on AIME 2024+2025 and HMMT 2025 for both Qwen3-30B-A3B and GPT-OSS-20B, suggesting that at least part of the failure reflects router-reachable misallocation rather than expert capacity alone.
Youngsik Yoon, Siwei Wang, Wei Chen +1
Department of Computer Science and Engineering, POSTECH, South Korea · Microsoft Research Asia, Beijing, China · Graduate School of Artificial Intelligence, POSTECH, South Korea