Conditional Rank Allocation for Taxonomy-Aware Medical Language Model Adaptation
Authors: Guangyuan Dong, Ziwei Hong, Xuehao Zhou, Zidong Yu, Bingchen Liu, Kehan Liu, Chuang Liu, Rong Fu, +1 more
Organizations: National University of Singapore, Singapore · University of Pennsylvania, USA · Department of Urology, Shanghai General Hospital, Shanghai, China · Shanghai Jiao Tong University School of Medicine, Shanghai, China · Syracuse University, USA · Shandong University, China
Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An additive gate combines question representations, specialty tags, operation tags, and their interaction; a learned coefficient scales the adapter residual. An illustrative separation under orthogonal, equiprobable subtasks shows how conditional selection can avoid an approximation floor faced by a fixed update with the same active rank. This result motivates the design without asserting a corresponding bound for medical corpora. On Qwen3-8B across CMB, CMExam, MedQA, and MedMCQA, five-seed experiments yield 69.69% mean accuracy across benchmarks, exceeding LoRA r16 and MoELoRA by 1.26 and 1.30 percentage points, respectively. The reported advantage over LoRA r16 increases from 0.08 to 1.94 points as training expands from one to seven specialties. Tag perturbations and atom masking support the usefulness of clinical routing, while atom clusters align with the supplied specialty labels (adjusted Rand index 0.62). Calibration, transfer, and measured costs further characterize the method. These findings support structured conditional adaptation for medical QA, while leaving clinical safety and broader deployment untested.
Figures & tables
Fig. 1: Arbor overview. Clinical priors and frozen question features feed four additive routing branches. Straight-through top- κ selection activates four of sixteen stored rank-one atoms; a learned coefficient scales their contribution alongside the frozen backbone. Mask and scale are shared across target layers. The training objective and question-level routing record are shown below. Highlighted mask positions are schematic.
Axis
Accuracy
Coverage
Majority
Agreement
Specialty (7)
91.4%
100%
28.6%
0.86
Operation (8)
84.6%
100%
24.2%
0.79
TABLE I: Clinical-prior validation on 500 annotated questions. Agreement is inter-annotator κ .
Method
Trainable (M)
CMB
CMExam
MedQA
MedMCQA
Macro
Frozen Qwen3-8B
–
71.68
68.95
62.86
54.86
64.59
LoRA r4
10.91
77.62±.21
75.67±.24
62.91±.42
57.19±.38
68.35±.18
LoRA r16
43.65
77.79±.19
75.98±.22
62.60±.45
57.34±.41
68.43±.21
LoRA r24
65.47
77.95±.23
75.88±.25
62.71±.44
57.62±.37
68.54±.22
DoRA r24
66.87
78.48±.20
75.86±.28
62.19±.47
56.90±.44
68.36±.17
HydraLoRA
44.90
78.02±.22
75.92±.24
63.15±.46
57.34±.40
68.61±.20
TABLE II: Qwen3-8B medical QA accuracy (%). Trainable methods share the CMB-source protocol; values are mean ± standard deviation over five seeds. Macro is the unweighted benchmark average. Bold marks the best trainable method. Adapter libraries compose pretrained adapters and are reported separately without across-seed uncertainty.
Fig. 2: Accuracy improvements under matched source training. (a) Benchmark gains over three baselines share a 0–2.5 percentage-point radial scale, with zero at the inner ring; the center reports Arbor ’s macro accuracy. (b) Paired macro differences against seven trainable baselines, with reported 95% intervals across five training seeds. All intervals exclude zero.
Fig. 3: Clinical routing and component analysis. (a) Tag-source controls with other components held fixed. (b) Aggregate accuracy under test-time tag noise (25 runs per level), with the reported training-heterogeneity endpoints shown separately. (c) Component ablations relative to full Arbor ; annotation uncertainty denotes standard deviation across five seeds. (d) Alignment with supplied specialty labels. Only available numerical summaries are shown; heterogeneity endpoints are not interpolated.
Method
Params
Active
Train
Latency
Memory
(M)
rank
(min)
(ms)
(GB)
LoRA r4
10.91
4
38
416
19.6
LoRA r16
43.65
16
41
418
19.8
LoRA r24
65.47
24
44
421
20.0
DoRA
66.87
24
58
435
21.0
HydraLoRA
44.90
16
49
452
20.6
TABLE III: Reported Qwen3-8B costs on one A100-80GB. Parameters are total trainable parameters; memory is peak inference memory. Latency is per question under greedy decoding.