Conditional Rank Allocation for Taxonomy-Aware Medical Language Model Adaptation
Authors: Guangyuan Dong, Ziwei Hong, Xuehao Zhou, Zidong Yu, Bingchen Liu, Kehan Liu, Chuang Liu, Rong Fu, +1 more
Organizations: National University of Singapore, Singapore · University of Pennsylvania, USA · Department of Urology, Shanghai General Hospital, Shanghai, China · Shanghai Jiao Tong University School of Medicine, Shanghai, China · Syracuse University, USA · Shandong University, China
Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An additive gate combines question representations, specialty tags, operation tags, and their interaction; a learned coefficient scales the adapter residual. An illustrative separation under orthogonal, equiprobable subtasks shows how conditional selection can avoid an approximation floor faced by a fixed update with the same active rank. This result motivates the design without asserting a corresponding bound for medical corpora. On Qwen3-8B across CMB, CMExam, MedQA, and MedMCQA, five-seed experiments yield 69.69% mean accuracy across benchmarks, exceeding LoRA r16 and MoELoRA by 1.26 and 1.30 percentage points, respectively. The reported advantage over LoRA r16 increases from 0.08 to 1.94 points as training expands from one to seven specialties. Tag perturbations and atom masking support the usefulness of clinical routing, while atom clusters align with the supplied specialty labels (adjusted Rand index 0.62). Calibration, transfer, and measured costs further characterize the method. These findings support structured conditional adaptation for medical QA, while leaving clinical safety and broader deployment untested.
Figures & tables
Fig. 1: Arbor overview. Clinical priors and frozen question features feed four additive routing branches. Straight-through top- κ selection activates four of sixteen stored rank-one atoms; a learned coefficient scales their contribution alongside the frozen backbone. Mask and scale are shared across target layers. The training objective and question-level routing record are shown below. Highlighted mask positions are schematic.
Axis
Accuracy
Coverage
Majority
Agreement
Specialty (7)
91.4%
100%
28.6%
0.86
Operation (8)
84.6%
100%
24.2%
0.79
TABLE I: Clinical-prior validation on 500 annotated questions. Agreement is inter-annotator κ .
Method
Trainable (M)
CMB
CMExam
MedQA
MedMCQA
Macro
Frozen Qwen3-8B
–
71.68
68.95
62.86
54.86
64.59
LoRA r4
10.91
77.62±.21
75.67±.24
62.91±.42
57.19±.38
68.35±.18
LoRA r16
43.65
77.79±.19
75.98±.22
62.60±.45
57.34±.41
68.43±.21
LoRA r24
65.47
77.95±.23
75.88±.25
62.71±.44
57.62±.37
68.54±.22
DoRA r24
66.87
78.48±.20
75.86±.28
62.19±.47
56.90±.44
68.36±.17
HydraLoRA
44.90
78.02±.22
75.92±.24
63.15±.46
57.34±.40
68.61±.20
TABLE II: Qwen3-8B medical QA accuracy (%). Trainable methods share the CMB-source protocol; values are mean ± standard deviation over five seeds. Macro is the unweighted benchmark average. Bold marks the best trainable method. Adapter libraries compose pretrained adapters and are reported separately without across-seed uncertainty.
Fig. 2: Accuracy improvements under matched source training. (a) Benchmark gains over three baselines share a 0–2.5 percentage-point radial scale, with zero at the inner ring; the center reports Arbor ’s macro accuracy. (b) Paired macro differences against seven trainable baselines, with reported 95% intervals across five training seeds. All intervals exclude zero.
Fig. 3: Clinical routing and component analysis. (a) Tag-source controls with other components held fixed. (b) Aggregate accuracy under test-time tag noise (25 runs per level), with the reported training-heterogeneity endpoints shown separately. (c) Component ablations relative to full Arbor ; annotation uncertainty denotes standard deviation across five seeds. (d) Alignment with supplied specialty labels. Only available numerical summaries are shown; heterogeneity endpoints are not interpolated.
Method
Params
Active
Train
Latency
Memory
(M)
rank
(min)
(ms)
(GB)
LoRA r4
10.91
4
38
416
19.6
LoRA r16
43.65
16
41
418
19.8
LoRA r24
65.47
24
44
421
20.0
DoRA
66.87
24
58
435
21.0
HydraLoRA
44.90
16
49
452
20.6
TABLE III: Reported Qwen3-8B costs on one A100-80GB. Parameters are total trainable parameters; memory is peak inference memory. Latency is per question under greedy decoding.
Medical multiple-choice question answering requires parameter-efficient adaptation across heterogeneous knowledge domains and reasoning operations. A medication question, a diagnostic decision, a public-health item, and a nursing-action item may require different low-rank updates, while some recall items should preserve the base model's representation with only mild adapter intervention. We propose BiRG-LoRA, a single-adapter rank-gated LoRA method for medical question answering. BiRG-LoRA keeps one LoRA module per target layer but makes its rank dimension input-conditioned: for each question, a biaxial gate combines hidden semantic evidence with specialty/profession priors, clinical-operation priors, and their interaction to select a sparse top-k subset of rank atoms. A scalar injection coefficient further controls the strength of the selected adapter update. Under a matched Qwen3-8B CMB-source protocol, BiRG-LoRA achieves the highest four-benchmark macro-average accuracy among trainable PEFT baselines and matched routing controls: 69.31% averaged over CMB, CMExam, MedQA, and MedMCQA. It improves over MoELoRA by 0.89 percentage points while using 28.1% fewer trainable parameters; a paired, benchmark-stratified bootstrap over final predictions gives a 95% confidence interval of [0.42, 1.37] for this macro-average gain. Basic controls show that BiRG-LoRA also improves over vanilla LoRA r16 and active-rank-matched LoRA r4 by 0.83 macro points, and an evaluation-time weak-axis perturbation check suggests that performance is not brittle to moderate tag noise. The results support a bounded claim: clinically structured rank allocation improves cross-benchmark medical QA under a matched single-seed protocol, while training-seed variance remains future work.
Hao Gong, Ruilin Gong, Yining Huang
University of Chinese Academy of Sciences · Henan University of Technology · Meta Emergence Laboratory
Medical large language models are commonly adapted with a fixed low-rank budget, even though medical questions differ substantially in confidence, clinical coverage, and cross-domain difficulty. We study adaptive rank budgeting for parameter-efficient medical question answering: for each question, the adapter decides whether to activate a small, medium, or large subset of LoRA rank channels. The central challenge is that a naive adaptive budget router can collapse to unstable choices or spend capacity without improving shifted benchmarks. We propose TriageRA-CCF, a source-side teacher for adaptive rank-budgeted LoRA. It combines three signals computed only from source training data: base-model answer confidence, metadata-cell clinical coverage, and a counterfactual close-miss proxy. These signals supervise a straight-through budget router over active ranks {2,4,8}, together with budget-cost, entropy, and rank-balance regularization. Under a matched CMB-source training protocol, TriageRA-CCF achieves the best average accuracy among LoRA, DoRA, and MoELoRA baselines on both Qwen3-8B and Llama3.1-8B. The gains are modest and non-uniform across benchmarks: +0.21 average points over the strongest external baseline on Qwen3-8B and +0.16 on Llama3.1-8B. Component ablations show that confidence, coverage, and counterfactual signals all provide useful budget supervision, but their combination is not monotonically best on every backbone.
Shucan Ji, Yining Huang, Hongliang Guo
College of Computer Science, Sichuan University, Chengdu, China · Meta Emergence Laboratory
Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited training data. We systematically investigate this assumption via tuned lens probing and causal activation patching, and find that Arabic medical knowledge is present in intermediate model representations but fails to surface at the output. This mechanistic insight motivates a targeted adaptation strategy: rather than fine-tuning the full network, we propose Targeted Low-Rank Adaptation (TLoRA), restricted to the layer window where cross-lingual representations diverge, upstream of the output layers where the failure manifests. We evaluate TLoRA on multiple-choice medical QA, where our approach outperforms full-network LoRA, zero-shot, and few-shot baselines. We further evaluate it on short-answer generation and multi-turn clinical dialogue, where it performs competitively without the need for task-specific finetuning. We additionally introduce AraClinicDialog, a clinician-constructed Arabic medical dialogue benchmark in MSA with validated variants across four Arabic dialects. Together, these contributions demonstrate that mechanistic diagnosis can serve as a practical guide for targeted adaptation in underrepresented-language medical LLMs.
Chaimae Abouzahir, Musa Khan, Hala Ali-Hassan +7
New York University Abu Dhabi · Cleveland Clinic Abu Dhabi