Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models
Authors: Dayan Pan, Jingyuan Wang, Xie Yu
Organizations: School of Computer Science and Engineering, MOE Engineering Research Center of Advanced Computer Application Technology, Beihang University Beijing, China · School of Computer Science and Engineering, School of Economics and Management, Beihang University Beijing, China
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.
Figures & tables
Figure 1 . Activation heterogeneity in a pretrained Llama-3.2-3B model. The x-axis in (a), (b), and (d) indexes query dimensions of attention mechanism. Activation patterns vary across layers (a), heads (b), and input token types (c, d), indicating that attention operates heterogeneously across dimensions, heads, layers, and tokens.
Figure 2 . Position-dependent responses across attention dimensions induced by RoPE. (a) Different dimensions respond differently to relative positional distances. (b) Heatmap of all dimensions, showing non-uniform positional sensitivity.
Figure 3 . The architecture of DyPAM framework. DyPAM applies input-conditioned, dimension-wise modulation together with head-wise and layer-wise structural biases to the query and key representations before RoPE, enabling fine-grained adaptation of positional attention within the PEFT paradigm.
Backbone LLM
Method
Param(%)
MultiArith
GSM8K
AddSub
AQuA
SingleEq
SVAMP
MAWPS
micro-avg(%) ↑
macro-avg(%) ↑
LLaMA 3.2 3B
LoRA
1.12
71.50
33.21
78.48
22.44
81.50
54.10
76.47
54.96
59.67
AdaLoRA
2.22
75.67
36.32
80.51
22.83
87.80
55.60
78.57
57.90
62.47
OFT
0.73
87.17
40.18
85.82
24.02
86.42
61.50
84.03
62.75
67.02
Bone
1.14
87.50
39.73
85.57
23.62
86.61
63.70
81.93
63.03
66.95
IA 3
0.02
58.33
27.37
68.61
20.47
72.83
47.90
58.82
46.89
50.62
LN-Tuning
0.01
58.00
26.38
66.58
21.26
74.80
44.90
60.08
46.01
50.29
Table 1. Comparison of DyPAM with PEFT baselines on mathematical reasoning benchmarks across three backbone models. Micro-avg and macro-avg denote micro- and macro-averaged performance. Best results are highlighted in bold, and second-best results are underlined. ∗ indicates statistically significant improvements over the best baseline (two-sided t-test, p<0.05 ).
Baseline
Qwen 3 0.6B
Qwen 3 1.7B
Qwen 3 4B
Qwen 3 8B
LoRA
64.06
66.64
75.60
80.37
OFT
65.96
67.81
75.54
80.45
SHiRA
63.95
64.65
73.33
81.04
RoSA
63.99
67.38
77.92
81.29
DyPAM (ours)
66.13
69.24
78.24
83.20
Table 2 . Macro-averaged accuracy on mathematical reasoning benchmarks across Qwen3 model scales. The comparison includes representative strong PEFT baselines from the main experiments.
Backbone LLM
Method
Param(%)
BoolQ
PIQA
SocialIQA
ARC-C
ARC-E
OpenBookQA
HellaSwag
WinoGrande
Micro Avg
Macro Avg
LLaMA 3.2 3B
LoRA
1.12
63.61
79.71
66.94
69.45
84.05
67.00
73.94
55.56
71.94
70.03
AdaLoRA
2.22
63.52
78.94
67.09
68.94
85.14
70.20
78.11
56.35
73.95
71.04
OFT
0.73
65.63
79.54
70.37
70.39
85.06
71.80
83.15
66.38
77.52
74.04
Bone
1.14
64.56
75.68
69.34
64.42
79.76
70.20
75.92
65.75
72.77
70.70
IA 3
0.02
62.32
77.09
59.67
57.94
77.10
57.40
50.48
52.25
58.66
61.78
LN Tuning
0.01
62.51
76.99
59.52
59.81
76.52
59.00
52.02
52.17
59.42
62.32
Table 3. Comparison of DyPAM with PEFT baselines on commonsense reasoning benchmarks across three backbone models. Micro-avg and macro-avg denote micro- and macro-averaged performance. Best results are highlighted in bold, and second-best results are underlined. ∗ indicates statistically significant improvements over the best baseline (two-sided t-test, p<0.05 ).
Figure 4 . Ablation and hyperparameter sensitivity of DyPAM. The ablation removes individual modulation components, while the sensitivity study varies the modulation strength α .
Figure 5 . Learned layer-wise query bias over attention dimensions on LLaMA-3.2-3B under commonsense and mathematical reasoning settings. The heatmaps show that the structural bias varies across layers and dimensions rather than following a uniform pattern.
Figure 6 . Layer-wise modulation range of query and key representations on LLaMA-3.2-3B under commonsense and mathematical reasoning settings. The mean scale remains close to 1, while the effective range changes across layers and training data.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 . Cross-model comparison of query activation patterns across layers and attention dimensions. RoPE-based LLaMA-3.2-3B shows structured, dimension-dependent patterns, while ALiBi-based BLOOM-560M and OPT-350M with learned embeddings are more homogeneous.
Figure 8 . Token-type activation patterns across layers and attention dimensions for LLaMA-3.2-3B, BLOOM-560M, and OPT-350M. Token-dependent variation is most pronounced in the RoPE-based model.
Dataset
Samples
Total Tokens
Avg. Tokens/Sample
Math10K
9,919
2,273,016
229.16
Commonsense15K
15,119
1,778,782
117.65
Appendix
Table 4. Statistics of the training datasets for commonsense and mathematical reasoning tasks.
Dataset
Samples
Answer Type
MultiArith
600
Numeric
GSM8K
1,319
Numeric
AddSub
395
Numeric
AQuA
254
Multiple Choice (A–E)
SingleEq
508
Numeric
SVAMP
1,000
Numeric
Appendix
Table 5. Statistics of Mathematical Reasoning Test Datasets.
Dataset
Samples
Answer Format
BoolQ
3,270
true / false
PIQA
1,838
solution1 / solution2
SIQA
1,954
answer1 / answer2 / answer3
ARC-Challenge
1,172
answer1 / answer2 / answer3 / answer4
ARC-Easy
2,376
answer1 / answer2 / answer3 / answer4
OBQA
500
answer1 / answer2 / answer3 / answer4
Appendix
Table 6. Statistics of Commonsense Reasoning Test Datasets.
Method
BS=1 (ms/tok)
BS=32 (ms/tok)
Mem (GB)
FLOPs overhead
Base
14.49
0.49
6.00
—
LoRA (unmerged)
27.64
1.52
6.06
73.0M (+1.13%)
RoSA
22.40
1.45
6.03
34.9M (+0.54%)
DyPAM (ours)
24.98
1.49
6.06
66.1M (+1.03%)
Appendix
Table 7. Inference overhead of DyPAM compared with the base model and representative PEFT baselines on LLaMA-3.2-3B. Latency is reported per token. The per-token value is lower at batch size 32 because more tokens are processed in parallel.
Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show -- both empirically and theoretically -- that heads with different functional roles require distinct frequency ranges and attention scaling factors to operate effectively. Ignoring this structure leads to suboptimal utilization of embedding dimensions and degraded performance, particularly under long-context settings. To address these limitations, we propose AdaRoPE, which equips each attention head with learnable rotation frequencies and attention scaling factors. Pretrained LLMs with AdaRoPE consistently outperform existing RoPE variants, including partial RoPE and NoPE baselines. For context extension, we further show that uniform frequency and attention scaling, used in methods such as YaRN, are suboptimal. By applying head-specific scaling, AdaRoPE enables better context extension while better preserving short-context performance in both the extrapolation setting and the long-context continued pretraining setting. These results highlight the importance of optimizing rotary position embedding at the level of individual attention heads.
Shaowen Wang, Yuke Zheng, Tansheng Zhu +4
Tsinghua University · Hunyuan Team, Tencent · Xiongan AI Institute +1
Rotary Positional Encodings (RoPE) are currently the most popular positional encodings used in modern language models. RoPE rotates two-dimensional chunks of query and key vectors, operating as a function of their relative positional offset. The position-wise rates of rotation in RoPE typically follow a geometric sequence specified by a fixed base-frequency hyperparameter. Prior work has improved performance by either increasing this parameter to slow rotation or by applying RoPE to only a subset of QK dimensions. In this work we modify RoPE by learning a scalar per frequency, treating frequencies as learnable parameters rather than hyperparameters. We validate Learned RoPE by training a ladder of language models from scratch, ranging from 52M to 2.5B parameters. We observe and analyze the emergence of a high-norm, positional LeRoPE band. LeRoPE consistently outperforms RoPE and partial RoPE across all scales, with RoPE requiring 3.4% more compute (FLOPs) to match LeRoPE at the largest scale.
Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
Muhammad Azeem Lodhi, Chao Zhou, Rebekka Burkholz
Saarland University, Saarbrücken, Germany · CISPA Helmholtz Center for Information Security, Saarbrücken, Germany