Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models
Authors: Dayan Pan, Jingyuan Wang, Xie Yu
Organizations: School of Computer Science and Engineering, MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University Beijing, China · School of Computer Science and Engineering, School of Economics and Management, Beihang University Beijing, China
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.
Figures & tables
Figure 1 . Activation heterogeneity in a pretrained Llama-3.2-3B model. The x-axis in (a), (b), and (d) indexes query dimensions of attention mechanism. Activation patterns vary across layers (a), heads (b), and input token types (c, d), indicating that attention operates heterogeneously across dimensions, heads, layers, and tokens.
Figure 2 . Position-dependent responses across attention dimensions induced by RoPE. (a) Different dimensions respond differently to relative positional distances. (b) Heatmap of all dimensions, showing non-uniform positional sensitivity.
Figure 3 . The architecture of DyPAM framework. DyPAM applies input-conditioned, dimension-wise modulation together with head-wise and layer-wise structural biases to the query and key representations before RoPE, enabling fine-grained adaptation of positional attention within the PEFT paradigm.
Backbone LLM
Method
Param(%)
MultiArith
GSM8K
AddSub
AQuA
SingleEq
SVAMP
MAWPS
micro-avg(%) ↑
macro-avg(%) ↑
LLaMA 3.2 3B
LoRA
1.12
71.50
33.21
78.48
22.44
81.50
54.10
76.47
54.96
59.67
AdaLoRA
2.22
75.67
36.32
80.51
22.83
87.80
55.60
78.57
57.90
62.47
OFT
0.73
87.17
40.18
85.82
24.02
86.42
61.50
84.03
62.75
67.02
Bone
1.14
87.50
39.73
85.57
23.62
86.61
63.70
81.93
63.03
66.95
IA 3
0.02
58.33
27.37
68.61
20.47
72.83
47.90
58.82
46.89
50.62
LN-Tuning
0.01
58.00
26.38
66.58
21.26
74.80
44.90
60.08
46.01
50.29
Table 1. Comparison of DyPAM with PEFT baselines on mathematical reasoning benchmarks across three backbone models. Micro-avg and macro-avg denote micro- and macro-averaged performance. Best results are highlighted in bold, and second-best results are underlined. ∗ indicates statistically significant improvements over the best baseline (two-sided t-test, p<0.05 ).
Baseline
Qwen 3 0.6B
Qwen 3 1.7B
Qwen 3 4B
Qwen 3 8B
LoRA
64.06
66.64
75.60
80.37
OFT
65.96
67.81
75.54
80.45
SHiRA
63.95
64.65
73.33
81.04
RoSA
63.99
67.38
77.92
81.29
DyPAM (ours)
66.13
69.24
78.24
83.20
Table 2 . Macro-averaged accuracy on mathematical reasoning benchmarks across Qwen3 model scales. The comparison includes representative strong PEFT baselines from the main experiments.
Backbone LLM
Method
Param(%)
BoolQ
PIQA
SocialIQA
ARC-C
ARC-E
OpenBookQA
HellaSwag
WinoGrande
Micro Avg
Macro Avg
LLaMA 3.2 3B
LoRA
1.12
63.61
79.71
66.94
69.45
84.05
67.00
73.94
55.56
71.94
70.03
AdaLoRA
2.22
63.52
78.94
67.09
68.94
85.14
70.20
78.11
56.35
73.95
71.04
OFT
0.73
65.63
79.54
70.37
70.39
85.06
71.80
83.15
66.38
77.52
74.04
Bone
1.14
64.56
75.68
69.34
64.42
79.76
70.20
75.92
65.75
72.77
70.70
IA 3
0.02
62.32
77.09
59.67
57.94
77.10
57.40
50.48
52.25
58.66
61.78
LN Tuning
0.01
62.51
76.99
59.52
59.81
76.52
59.00
52.02
52.17
59.42
62.32
Table 3. Comparison of DyPAM with PEFT baselines on commonsense reasoning benchmarks across three backbone models. Micro-avg and macro-avg denote micro- and macro-averaged performance. Best results are highlighted in bold, and second-best results are underlined. ∗ indicates statistically significant improvements over the best baseline (two-sided t-test, p<0.05 ).
Figure 4 . Ablation and hyperparameter sensitivity of DyPAM. The ablation removes individual modulation components, while the sensitivity study varies the modulation strength α .
Figure 5 . Learned layer-wise query bias over attention dimensions on LLaMA-3.2-3B under commonsense and mathematical reasoning settings. The heatmaps show that the structural bias varies across layers and dimensions rather than following a uniform pattern.
Figure 6 . Layer-wise modulation range of query and key representations on LLaMA-3.2-3B under commonsense and mathematical reasoning settings. The mean scale remains close to 1, while the effective range changes across layers and training data.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 . Cross-model comparison of query activation patterns across layers and attention dimensions. RoPE-based LLaMA-3.2-3B shows structured, dimension-dependent patterns, while ALiBi-based BLOOM-560M and OPT-350M with learned embeddings are more homogeneous.
Figure 8 . Token-type activation patterns across layers and attention dimensions for LLaMA-3.2-3B, BLOOM-560M, and OPT-350M. Token-dependent variation is most pronounced in the RoPE-based model.
Dataset
Samples
Total Tokens
Avg. Tokens/Sample
Math10K
9,919
2,273,016
229.16
Commonsense15K
15,119
1,778,782
117.65
Appendix
Table 4. Statistics of the training datasets for commonsense and mathematical reasoning tasks.
Dataset
Samples
Answer Type
MultiArith
600
Numeric
GSM8K
1,319
Numeric
AddSub
395
Numeric
AQuA
254
Multiple Choice (A–E)
SingleEq
508
Numeric
SVAMP
1,000
Numeric
Appendix
Table 5. Statistics of Mathematical Reasoning Test Datasets.
Dataset
Samples
Answer Format
BoolQ
3,270
true / false
PIQA
1,838
solution1 / solution2
SIQA
1,954
answer1 / answer2 / answer3
ARC-Challenge
1,172
answer1 / answer2 / answer3 / answer4
ARC-Easy
2,376
answer1 / answer2 / answer3 / answer4
OBQA
500
answer1 / answer2 / answer3 / answer4
Appendix
Table 6. Statistics of Commonsense Reasoning Test Datasets.
Method
BS=1 (ms/tok)
BS=32 (ms/tok)
Mem (GB)
FLOPs overhead
Base
14.49
0.49
6.00
—
LoRA (unmerged)
27.64
1.52
6.06
73.0M (+1.13%)
RoSA
22.40
1.45
6.03
34.9M (+0.54%)
DyPAM (ours)
24.98
1.49
6.06
66.1M (+1.03%)
Appendix
Table 7. Inference overhead of DyPAM compared with the base model and representative PEFT baselines on LLaMA-3.2-3B. Latency is reported per token. The per-token value is lower at batch size 32 because more tokens are processed in parallel.