Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.
Figures & tables
Figure 1: Two routes to the intermediate blocks (Split ImageNet-R). (a) Placement search with no observation (48 slots; seeds 0–2) peaks at blocks 4–7 (AugReg) and 3–6 (iBOT); dots show blocks selected by internal criteria. (b) LS-B: a frozen fMRI encoding model turns each of the first K=3 tasks into a readout Dt ; ρl=ml/sl commits k=4 blocks and the other eight share one slot. Example: AugReg seed 0 commits blocks 2, 6, 7, 8 (72 of 120 slots); allocations vary across seeds.
Figure 2: LS-B end to end (AugReg, Split ImageNet-R). (a) Offline, once per backbone: an encoding model fitted to NSD subj01 maps frozen block features to cortex (flatmaps: preferred block per vertex; R2 of the retained vertices). (b) Observation ( t<K=3 ): every block trains a task-specific slot, while unlabeled task images pass forward through an adapter-free copy and the encoding model to give Dt . (c) Dt=C+Et yields sl , ml and ρl ; the top k=4 blocks are committed once. (d) Later tasks add slots only to committed blocks; the others update one shared slot. Matrices: diagnostic readout; slots: seed 0 (chosen by index; committed blocks vary across seeds).
Figure 3: Both routes lead to the intermediate blocks. (a) Top: final accuracy of a four-block window at 48 slots without observation (large: mean of seeds 0–2; small: single seeds; open: five further seeds of window 8–11). Window 0–3 extends the eight registered windows. Bottom: readout window score ∑l∈wρl ( K=3 diagnostic readout, min–max scaled). (b) Blocks selected by each criterion; LS-B: fraction of seeds committing each block (solid: all seeds); boxes: best window of (a). (c) Matched storage and observation ( K=3 , 72 slots): only the committed blocks differ. Thin lines: single seeds; star: LS-B mean over 16 seeds.
Figure 4: Storage–accuracy trade-off ( a AugReg, b iBOT); storage: retained slots relative to full BiLoRA. LS-B with k=1 –8 ( K=3 ) lies above the static shallow-first allocations for k≤5 and within 0.42 pp of them for k≥6 ; the default uses 60%. Stars: the default over 16 seeds. Dashes at 40%: all other four-block allocations (all runs). Dashed line and band: full BiLoRA ± 0.42 pp.
Figure 5
Figure 7: Committed blocks on visual cortex (flattened fsaverage, NSD visual mask; lines: region outlines). (a) Layer map of the frozen encoding model: each vertex colored by the block with the largest layer-selection weight (a possibly weak peak). (b) Footprint of LS-B’s allocation, ∑lflsel[v,l] , with fl the fraction of seeds (16, 16, 8) committing block l ; uniform: 4/12 . (c) Share of the task residual of the four top-ranked blocks of the K=3 diagnostic readout per region (dark: V1–V4; dotted: 1/12 ). Panels (b) and (c) use different committed sets (seed frequencies vs. diagnostic top four; overlap 3/4, 4/4, 4/4). Overlapping vertices count in both regions.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Blocks (AugReg / iBOT)
K
Slots
Runs
LS-B ( k=4 , K=3 )
per run
3
72
16/16/8
LS-B, k∈{1,2,3,5,6,7,8}
per run
3
48+6k
3/3/3 a
LS-B, K∈{1,2} , k=4
per run b
1, 2
56, 64
3/3/3 a
Matched observation, commit 0–3
0–3
3
72
3/3/–
Matched observation, commit 8–11
8–11
3
72
3/3/–
Window w…w+3 , w=1,…,7
w…w+3
0
48
3/3/–
Appendix
Table 1: Configurations on Split ImageNet-R. Blocks are indexed 0…11 ; “Slots” is the number of task–layer slots retained after T=10 tasks ( 12+9k without observation, 48+6k for LS-B with K=3 ); one slot holds 6,000 parameters. Runs: AugReg/iBOT/DINO; “–”: not run. a On DINO only k∈{2,4} and K∈{1,3} . b At K=1 the score is identically zero and the rule commits blocks 0–3. c On DINO only l∈{4,6,8} . d Coincides with window 2–5: the same experiment. e Coincides with the deepest four blocks; no separate runs. f Plus 16 further independent draws per backbone on AugReg and iBOT.
Figure 8: The frozen fMRI encoding model behind the readout (NSD subj01, flattened cortex). (a) The twelve visual regions; vertices in two or more regions are grey. (b) Retained vertices (500 best-fitting per region) colored by validation R2 , and mean R2 per region. (c) Weights with which the readout aggregates each region over the blocks ( R2 -weighted mean layer selector; rows sum to one). Dots mark the weighted mean block, which rises from V1 to V4 and is largest in EBA, FBA and FFA on every backbone.
Figure 9: Worked example of the specialization score on the diagnostic readout of the three observation tasks. Rows: backbones. Columns 1–3: two-way-centered readouts Dt (12 regions × 12 blocks); column 4: their mean C ; column 5: residual energy averaged over tasks (square-root scale). Colors are scaled per backbone. Right: ml , sl and ρl ; colored bars mark the four top-ranked blocks ( {2,6,7,8} , {2,5,6,7} , {1,3,5,6} ). Block 11 carries the largest shared energy on every backbone and would be selected by ml alone; the ratio does not select it. The task residual is 5–7% of the centered readout. This diagnostic readout is not the per-run readout; committed sets vary across seeds (Table 1 ).
Figure 10: Task-evoked cortical readout during observation (retained vertices; AugReg and iBOT, whose vertex-level readouts were archived). Column 1: the selector-weighted, reference-standardized dispersion ∑lsel[n,l]zt[n,l] averaged over the three observation tasks (negative: less dispersion than on NSD images). Columns 2–4: deviation of each task from that mean. Column 5: residual energy on the four top-ranked blocks (square-root scale). These maps precede region aggregation and centering; most of the pattern is shared across tasks (pairwise correlations 0.93–0.998).
Configuration
Slots
Storage
AugReg
iBOT
DINO
Gain (AugReg / iBOT)
All shared
12
10%
64.95
62.67
58.62
–
Shallow-first, l=4
48
40%
68.44
66.53
62.40
–
Best window (4–7 / 3–6)
48
40%
71.20
68.58
–
–
LS-B k=1
54
45%
70.83
68.57
–
+1.63 / +1.40
LS-B k=4 , K=1
56
47%
70.10
67.73
63.43
+0.64 / +0.35
LS-B k=2
60
50%
71.43
69.01
63.97
+1.47 / +1.20
Appendix
Table 2: LS-B operating points and static references (ACC in %, seeds 0–2; storage relative to full BiLoRA). “Gain”: storage-matched gain over the interpolation of static shallow-first allocations (pp, seed-paired). Default over all seeds: 72.29 / 68.88 / 64.39.
Figure 11: Accuracy after each task ( a AugReg, b iBOT). Top: accuracy on all classes seen so far; bottom: difference from full BiLoRA; grey: observation stage. After the last task, LS-B with k=4 is 1.46 and 0.47 pp below full BiLoRA, with k=2 2.25 and 1.05 pp, and the static allocation of blocks 0–3 5.25 and 3.53 pp.
Configuration
Slots
CIFAR-100
CUB-200
CUB-200
CUB-200
AugReg
AugReg
iBOT
DINO
All shared
12
84.11 (1.32)
78.50 (1.61)
45.93 (0.89)
45.35 (4.07)
Shallow-first, l=4
48
84.96 (1.89)
76.69 (1.65)
49.24 (1.02)
47.13 (4.50)
LS-B k=4 , K=1
56
85.82 (1.85)
78.46 (2.12)
49.92 (1.31)
47.99 (4.37)
LS-B k=2 , K=3
60
85.88 (1.81)
77.96 (2.29)
50.79 (1.57)
48.94 (3.82)
Shallow-first, l=6
66
85.80 (1.75)
78.39 (2.42)
50.47 (1.06)
48.73 (4.54)
Appendix
Table 3: Supplementary ten-task conditions (ACC in %, mean of seeds 0–2; in parentheses: range over the three seeds). Slots are retained task–layer slots after T=10 tasks (full BiLoRA: 120); LS-B with K=1 commits blocks 0–3 and so adds only observation. Bottom: registered readouts for LS-B with k=4 / k=2 ( K=3 ) in pp: mean ACC minus the upper convex hull of the other seven configurations at the same slot count; seed-paired gain over the same-seed interpolation of the static shallow-first allocations; seed-paired difference from full BiLoRA. n=3 throughout; no significance claims.
Figure 12: Supplementary conditions (seeds 0–2; bars: range over the three seeds). (a–d) Final accuracy against storage relative to full BiLoRA on CIFAR-100 (AugReg) and CUB-200 (AugReg, iBOT, DINO). Gray: all shared, shallow-first prefixes of 4, 6 and 8 blocks, and full BiLoRA (also the dashed line); colored: LS-B with k=2 and k=4 ( K=3 ), and LS-B with k=4 and K=1 (diamond), which commits blocks 0–3 and so adds only observation. (e) ImageNet-R with AugReg at T=10 (filled) and T=20 (open; prefixes of 2, 4 and 6 blocks); dashed and dotted lines: full BiLoRA at T=10 and T=20 . (f) Blocks committed by LS-B ( k=4 , K=3 ) for seeds 0–2 of every condition; on ImageNet-R ( T=10 ) these are the default configuration’s seeds 0–2.
Configuration
Slots
Storage
Final
Avg
Gain
All shared
12
5.0%
58.91 (0.18)
63.62 (0.66)
–
Shallow-first, l=2
50
20.8%
60.12 (0.90)
64.50 (0.82)
–
LS-B k=2 , K=3
80
33.3%
64.89 (0.33)
69.46 (0.31)
+3.12
Shallow-first, l=4
88
36.7%
62.22 (0.07)
65.79 (0.34)
–
LS-B k=4 , K=3 (default)
112
46.7%
66.02 (0.95)
70.18 (0.64)
+2.20
Shallow-first, l=6
126
52.5%
64.75 (0.58)
67.90 (0.11)
–
Appendix
Table 4: Split ImageNet-R with T=20 tasks of ten classes, AugReg (seeds 0–2; in parentheses: range). Final: ACC after the last task; Avg: mean of the accuracies after each of the 20 tasks. Storage relative to full BiLoRA (240 slots). Gain: seed-paired gain over the same-seed interpolation of the static shallow-first allocations (pp).
Quantity
Backbone
ImageNet-R
CIFAR-100
CUB-200
Order holds
Observation only
AugReg
+1.66
+0.85
+1.77
no
iBOT
+1.21
–
+0.68
yes
DINO
+1.02
–
+0.86
yes
Gain over interpolation, k=4
AugReg
+0.89
−0.04
−0.40
yes
iBOT
+0.77
–
+1.20
no
DINO
+0.29
–
+0.75
no
Appendix
Table 5: Registered direction test: both quantities were predicted to order the datasets as ImageNet-R ≥ CIFAR-100 ≥ CUB-200 on every backbone (pp, seed-paired, seeds 0–2). Observation only: LS-B with K=1 , which commits blocks 0–3 (56 slots), minus the static blocks 0–3 (48 slots). Gain over interpolation: as in Table 3 ; on ImageNet-R the k=4 point is the default configuration’s seeds 0–2. “–”: no runs (the registered iBOT × CIFAR-100 condition was not run, Appendix G ; DINO × CIFAR-100 was not registered).
Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combinations of vision blocks can be skipped without fine-tuning. A shared K-block route skips one searched set of exactly K blocks for every question, while a capability-specific K-block policy selects one same-size route using a known capability label. We introduce a source-balanced evolutionary search and compare it with independent ranking, contiguous removal, and random routes at matched budgets. Experiments use Qwen2.5-VL-3B-Instruct, SmolVLM2-2.2B-Instruct, and an 876-example image-disjoint selection split. Search transfers across architectures: on SmolVLM2, the searched shared four-block route beats independent construction by 4.91 percentage points. Capability specialization is less stable. On Qwen, the six-block capability policy beats the shared route by 2.17 points, driven by a 7.10-point OCR gain. On sealed IIIT5K, however, the SmolVLM2 OCR-specific route trails its shared route by 13.6 points. Combinatorial search reliably improves route construction, but capability labels do not define universally transferable vision pathways.
Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation site shapes what a model learns, how well that learning generalizes, and how selectively it is applied. We introduce a controlled benchmark spanning five objectives (lexical binding, factual association, behavioral policy learning, causal mapping, and procedural reasoning) and define each objective's "adaptation geometry" as its profile of acquisition, transfer, and boundedness under full-stack and early-, middle-, or late-layer LoRA. The objectives exhibit distinct geometries. Lexical binding favors early-layer adaptation for acquisition and boundedness but requires broader updates for transfer; factual association favors later layers among localized adapters; behavioral learning separates late-layer action acquisition from middle-layer policy gating; and causal and procedural transfer benefit most from middle- or full-stack adaptation. These patterns largely persist under parameter-matched controls, and most corresponding directional contrasts replicate across five model families. These findings establish adaptation site as a key design variable for controlling what models learn, generalize, and leave unchanged.
Low-rank adaptation (LoRA) is a widely used parameter-efficient fine-tuning method that places trainable low-rank adapters into frozen pre-trained models. Recent studies show that using fewer LoRA adapters may still maintain or even improve performance, but existing methods still distribute adapters broadly, leaving where to place a limited number of adapters to maximize performance largely open. To investigate this, we introduce PAGE (Projected Adapter Gradient Energy), a gradient-based sensitivity probe that estimates the initial trainable gradient energy available to each candidate LoRA adapter. Surprisingly, we find that PAGE is highly concentrated on a single shallow FFN down-projection across two model families and four downstream tasks. We term this module the dominant adaptation module and show that its layer index is architecture-dependent but task-stable. Motivated by this finding, we propose DomLoRA, a placement method that places a single adapter at the dominant adaptation module. With only ~0.7% of vanilla LoRA's trainable parameters, DomLoRA outperforms it on average across various downstream tasks, including instruction following, mathematical reasoning, code generation, and multi-turn conversation. This method also improves other LoRA variants, supporting the dominant adaptation module perspective as a practical placement guideline.
Suoxin Zhang, Run He, Di Fang +3
South China University of Technology, China · Zhejiang University, China