On-device learning is necessary when the model encounters user-,sensor-, or environment-specific shifts after deployment. Although parameter-efficient fine-tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA) variants, enable efficient adaptation at the edge, the limiting resource for Convolutional Neural Network (CNN) adaptation is often not the number of trainable parameters but the activation state that must be retained until the backward pass. This paper introduces Memory-Floor LoRA (MemFLoRA), a low-rank CNN adapter built around a memory-first design principle rather than a direct application of transformer-oriented LoRA. Instead of merely reducing trainable weights, we define an activation-memory-floor criterion: trainable backward computations must not depend on full-width layer inputs. The resulting adapter freezes the down-projection, trains a scale-matched up-projection, and combines eval-mode backbone normalization with activation-minimal backward rules, reducing saved state to the low-rank branch. Evaluated on three Human Activity Recognition (HAR) datasets and two CNN backbones under subject, body-location, and sensor-placement shifts, MemFLoRA reduces saved-activation memory by 98.5-98.7% and peak training-state memory by 94.9-97.3% relative to full fine-tuning, while matching or exceeding CNN PEFT baselines.
Figures & tables
Figure 1. Brief overview of MemFLoRA computation and saved-state graph, reducing persistent saved state from O(BCℓTℓ) to O(BrTℓ) plus a bit-packed ReLU mask. MemFLoRA computation and saved-state flow MemFLoRA computation and saved-state flow. An input feature tensor splits into a frozen backbone branch and a low-rank adapter branch. The backbone branch applies a frozen convolution followed by evaluation-mode batch normalization. The adapter branch applies a frozen pointwise down-projection initialized from a normal distribution, optionally applies bottleneck batch normalization, then applies a trainable up-projection and a fixed per-channel scale inherited from the source batch normalization. The two branches are added and passed through a rectified linear unit. The diagram shows that only the bottleneck activation, whose channel count equals the adapter rank, and a bit-packed rectified-linear-unit sign mask are retained. The full-width input, backbone activations, and batch-normalization states are dropped rather than persistently saved for backward computation.
Figure 2. Peak training-state memory decomposition of MobileNetV2 at rank-2 on Opportunity with a batch size of 64. Labels report total peak memory, with saved-activation memory in parentheses. MobileNetV2 peak training-memory decomposition Stacked horizontal bars compare peak training memory for MobileNetV2 on Opportunity at rank two and batch size sixty-four. The five bars represent Full Fine-Tuning, LoRA-C, LoRA-Edge, MemFLoRA, and MemFLoRA-SG. Each bar is divided into model state, optimizer state, and live training state. Live training state occupies most of the memory for Full Fine-Tuning, LoRA-C, and LoRA-Edge. Total peak memory and saved-activation memory, both in megabytes, are 666.64 and 639.56 for Full Fine-Tuning, 592.61 and 583.05 for LoRA-C, 491.22 and 481.79 for LoRA-Edge, 17.83 and 8.40 for MemFLoRA, and 26.30 and 8.40 for MemFLoRA-SG. The MemFLoRA bars are substantially shorter than the three conventional training methods, while MemFLoRA-SG has a slightly larger total than MemFLoRA because of its additional replay state.
Dataset
Backbone
Zero Shot
Full FT
Bias-Tuning
BN-Tuning
r
LoRA-C
LoRA-Edge
MemFLoRA
MemFLoRA-SG
Opportunity
MobileNetV2
2
0.698 ± 0.037
0.730 ± 0.032
0.747 ± 0.025
0.758 ± 0.028
0.585 ± 0.088
0.820 ± 0.025
0.632 ± 0.054
0.668 ± 0.043
4
0.728 ± 0.032
0.760 ± 0.027
0.764 ± 0.028
0.774 ± 0.024
8
0.756 ± 0.028
0.789 ± 0.025
0.771 ± 0.028
0.785 ± 0.023
T-ResNet
2
0.746 ± 0.044
0.710 ± 0.052
0.768 ± 0.046
0.773 ± 0.040
0.511 ± 0.087
0.847 ± 0.026
0.655 ± 0.054
0.684 ± 0.052
4
0.773 ± 0.039
0.746 ± 0.046
0.793 ± 0.034
0.797 ± 0.035
8
0.801 ± 0.030
0.774 ± 0.038
0.812 ± 0.033
0.815 ± 0.031
Table 1. Headline Macro-F1 comparison at 50 adaptation steps. Mean ± sample standard deviations across 20 seeds are reported.
Dataset
Shift
Subj.
Act.
Hz
Win./Str.
Opportunity ( Chavarriaga et al., 2013 )
Subject
4
17
30
60/30
RealWorld ( Sztyler and Stuckenschmidt, 2016 )
Body location
15
8
50
500/250
RealDisp ( Banos et al., 2014 )
Sensor placement
17
33
50
250/125
Table 2. Datasets, domain shifts, and preprocessing settings.
Method
B1
B8
B32
B64
Opt. [MB]
T-ResNet
Full FT
9.02/0.90
13.84/7.09
35.08/28.33
63.40/56.65
4.49
LoRA-C
5.51/3.10
11.53/9.12
32.20/29.79
59.76/57.35
0.10
LoRA-Edge
2.93/0.60
7.05/4.72
21.19/18.86
40.04/37.71
0.02
MemFLoRA
2.44/0.02
2.50/0.11
2.81/0.42
3.23/0.83
0.08
MemFLoRA-SG
2.50/0.02
2.78/0.11
3.74/0.42
5.01/0.83
0.10
MobileNetV2
Full FT
43.20/16.12
109.00/81.91
346.93/319.85
666.64/639.56
17.97
Table 3. For T-ResNet and MobileNetV2, batch-size entries report peak/saved-activation memory [MB]; optimizer-state memory is reported separately.
Figure 3. Adaptation performance at rank=2 across adaptation steps for T-ResNet (left) and MobileNetV2 (right). Adaptation trajectories across datasets and backbones Six line charts show the macro-averaged F1 score as the number of adaptation steps increases from fifty to two hundred fifty. The left column presents T-ResNet and the right column presents MobileNetV2. The rows present Opportunity, RealWorld, and RealDisp, respectively. Each panel compares Zero Shot, Full Fine-Tuning, Bias-Tuning, Batch-Normalization Tuning, LoRA-C, LoRA-Edge, MemFLoRA, and MemFLoRA-SG. Zero Shot appears as a constant dashed baseline, while the adapted methods generally improve during the first adaptation checkpoints and then level off. Full Fine-Tuning is highest or close to highest across the panels. MemFLoRA-SG is generally the strongest low-rank method and is closely followed by MemFLoRA. The RealWorld T-ResNet panel shows the widest separation among methods, whereas the MobileNetV2 trajectories are more tightly clustered, particularly on RealDisp.
Factor
Variant
r=2
r=4
r=8
Init.
Normal P
0.759
0.777
0.802
Orthogonal P
0.673
0.709
0.751
Geom.
1×1 down, k×k up
0.760
0.777
0.802
k×k down, 1×1 up
0.731
0.755
0.773
Scale
None
0.200
0.197
0.187
BNr
0.728
0.772
0.805
Table 4. Ablations on Opportunity with T-ResNet at 50 adaptation steps.
Backbone
Method
Par. [103]
MACs [106]
Relative cost [%]
Steps to 85%[n]
T-ResNet
Full FT
560.9
320.0
–
–
LoRA-C
12.8
320.2
100
22
LoRA-Edge
2.3
218.6
68
24
MemFLoRA
10.6
222.5
70
19
MemFLoRA-SG
10.6
331.7
104
15
MobileNetV2
Full FT
2250
403.0
–
–
Table 5. Training-performance summary at rank 2. Total adaptation MACs are reported as AdaBN calibration MACs plus 50 adaptation steps.
Method
T-ResNet
MobileNetV2
Full FT
54.52 / 47.77
666.72 / 639.55
LoRA-C
52.37 / 49.96
592.61 / 583.04
LoRA-Edge
35.12 / 32.79
491.21 / 481.78
MemFLoRA
3.22 / 0.83
17.83 / 8.39
Table 6. Training-state memory measured at rank r=2 and batch size B=64 . Entries report peak concurrent training-state storage / saved-backward state [MB].
Method
T-ResNet
MobileNetV2
Full FT
66.20
673.39
LoRA-C
64.08
601.45
LoRA-Edge
44.80
503.32
MemFLoRA
19.43
95.43
Table 7. Peak requested CUDA memory at r=2 , B=64 [MB], from the same runs as Table 6 . Values include inputs, temporaries, and workspaces in addition to training state.
Figure 4. Requested GPU memory during one MobileNetV2 update at r=2 , B=64 on Jetson Orin Nano. Two stacked area plots of requested GPU memory over one training update. Full FT rises during the forward pass to a peak of 673.4 MB dominated by saved activations, then declines during the backward pass. MemFLoRA stays low, with a narrow 95.4 MB spike of working buffers early in the forward pass and a second 95.2 MB spike at the end of the backward pass.
As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challenging. Low-Rank Adaptation (LoRA), currently one of the most widely adopted parameter-efficient fine-tuning (PEFT) methods, mitigates this challenge by optimizing only low-rank adaptation matrices, thereby greatly reducing the number of trainable parameters. With the parameter overhead substantially reduced, the activations retained for backpropagation have emerged as the primary remaining memory bottleneck during LoRA fine-tuning. To address this, we propose CARE-LoRA, a data-aware Compressed Activation REconstruction framework. By exploiting the inherent projection structure of LoRA, CARE-LoRA replaces the full input activation with the low-rank compressed activation naturally produced by the LoRA branch. It further computes a lightweight reconstruction matrix during the forward pass with negligible additional computation cost, which is used during backpropagation to reconstruct the gradient signal, thereby keeping LoRA matrices fully trainable. Extensive experiments across diverse models and downstream tasks demonstrate that, while substantially reducing the overall memory footprint, CARE-LoRA achieves competitive or even superior performance compared with standard LoRA and representative LoRA variants. Our code is publicly available at https://github.com/fishandyu/CARE-LoRA .
Gengyu Zhang, Haiyin Ran, Zhengbao He +4
Institute of Image Processing and Pattern Recognition, Shanghai Jiao Tong University, Shanghai 200240, China
Fine-tuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) on an end-user's data offers personalized experiences while keeping data private, but faces severe memory constraints on consumer hardware. Peak memory during fine-tuning often exceeds device limits, especially for models with billions of parameters and long-context training data. This paper introduces a suite of complementary techniques to reduce memory footprint without sacrificing model quality: (1) base model quantization with on-the-fly dequantization, (2) memory-efficient checkpointing combining selective activation caching and disk offloading, (3) softmax approximation using semantically relevant token subsets, and (4) logits masking. Experiments on Llama-3.2 3B and Qwen-2.5 3B demonstrate up to 26× and 28× reduction in peak memory, enabling fine-tuning on resource-constrained devices.
Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
Muhammad Azeem Lodhi, Chao Zhou, Rebekka Burkholz
Saarland University, Saarbrücken, Germany · CISPA Helmholtz Center for Information Security, Saarbrücken, Germany