DADP: Dynamic Activity-Dependent Pruning, A Reverse Hebbian-Inspired Structural Pruning Method
Organizations: Independent Researcher, Pune, India
Abstract
Modern neural networks are heavily over-parameterized. This redundancy incurs substantial compute and memory overhead during training and inference. Existing pruning methods rely on post-hoc magnitude thresholds or static initialization heuristics. Consequently, they often require manual per-layer sparsity targets or expensive retraining cycles. We propose Dynamic Activity-Dependent Pruning (DADP), a biologically inspired structural plasticity mechanism. During training, DADP measures connection importance via the accumulated product of pre-synaptic activations and post-synaptic error gradients. Using a single global threshold instead of fixed layer budgets, DADP dynamically allocates sparsity across network depth while naturally inducing neuron- and channel-level pruning. Across MLP, VGG-16, ResNet-18, BiLSTM-CRF, and MiniBERT architectures, DADP matches or outperforms Magnitude, SNIP and RigL, retaining 73.67% accuracy (dense baseline: 76.06%) at 99% sparsity on ResNet-18. Finally, matrix-based Shannon entropy and effective rank measurements confirm that DADP preserves latent feature diversity at extreme sparsities without representation collapse.
Figures & tables
| Architecture | Method | Sparsity (%) | Accuracy / F1 (%) | Acc. Change vs. Dense |
|---|---|---|---|---|
| MLP (MNIST) | Dense Baseline | 0.00% | 98.39% | – |
| Magnitude | 80.00% | 98.63% | +0.24% | |
| SNIP | 80.00% | 97.85% | -0.54% | |
| RigL | 80.00% | 97.97% | -0.42% | |
| DADP ( ) | 84.90 0.56% | 97.91 0.12% | -0.48 0.12% | |
| VGG-16 (CIFAR-10) | Dense Baseline | 0.00% | 85.21% | – |
| ResNet-18 (CIFAR-10) | VGG-16 (CIFAR-10) | |||
| Initialization Scheme | Sparsity (%) | Test Acc (%) | Sparsity (%) | Test Acc (%) |
| Kaiming Normal | ||||
| Kaiming Uniform | ||||
| Xavier Normal | ||||
| Xavier Uniform | ||||
| Orthogonal | ||||
| ResNet-18 (CIFAR-10) | VGG-16 (CIFAR-10) | MiniBERT (SST-2) | ||||
|---|---|---|---|---|---|---|
| Prune Interval ( ) | Sparsity (%) | Test Acc (%) | Sparsity (%) | Test Acc (%) | Sparsity (%) | Test Acc (%) |
| (Default) | ||||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | MLP | VGG-16 | ResNet-18 | BiLSTM-CRF | MiniBERT |
|---|---|---|---|---|---|
| Dataset | MNIST | CIFAR-10 | CIFAR-10 | CoNLL-2003 | SST-2 |
| Task Modality | Vision (Digits) | Vision (Objects) | Vision (Objects) | Sequence (NER) | NLP (Sentiment) |
| Total Parameters | 667,146 | 14,728,266 | 11,173,962 | 1,514,841 | 4,386,178 |
| Optimizer | Adam | Adam | Adam | Adam | Adam |
| Learning Rate ( ) | |||||
| Weight Decay |
| Method | Configuration / Threshold | Final Sparsity (%) | Final Test Acc (%) | Peak Test Acc (%) |
|---|---|---|---|---|
| Dense Baseline | Unpruned (0%) | 0.00% | 98.39% | 98.44% |
| DADP (Ours) | 84.90 0.56% | 97.91 0.12% | 98.18 0.02% | |
| DADP (Ours) | 94.60 0.14% | 97.55 0.09% | 98.02 0.03% | |
| DADP (Ours) | 96.19 0.11% | 97.38 0.15% | 97.91 0.08% | |
| DADP (Ours) | 98.35 0.04% | 96.69 0.04% | 96.92 0.05% | |
| DADP (Ours) | (100 Epoch Limit Test) | 98.00% | 96.94% | 97.69% |
| Method | Configuration / Threshold | Final Sparsity (%) | Final Test Acc (%) | Peak Test Acc (%) |
|---|---|---|---|---|
| Dense Baseline | Unpruned (0%) | 0.00% | 76.06% | 77.44% |
| DADP (Ours) | 84.65 0.21% | 76.72 0.26% | 76.97 0.25% | |
| DADP (Ours) | 89.66 0.10% | 76.64 0.75% | 77.36 0.17% | |
| DADP (Ours) | 91.41 0.08% | 76.76 0.16% | 77.25 0.24% | |
| DADP (Ours) | 95.63 0.05% | 76.15 0.34% | 77.08 0.16% | |
| DADP (Ours) | 96.92 0.01% | 76.03 0.61% | 76.80 0.21% |
| Method | Configuration / Threshold | Final Sparsity (%) | Final Test F1 (%) | Peak Test F1 (%) |
|---|---|---|---|---|
| Dense Baseline | Unpruned (0%) | 0.00% | 85.16% | 85.16% |
| DADP (Ours) | 12.10 3.16% | 93.77 0.12% | 93.94 0.13% | |
| DADP (Ours) | 26.19 3.80% | 93.82 0.06% | 93.96 0.15% | |
| DADP (Ours) | 33.57 4.06% | 93.82 0.14% | 93.95 0.13% | |
| DADP (Ours) | 95.00 3.60% | 84.28 1.34% | 93.82 0.23% | |
| DADP (Ours) | 95.09 0.98% | 87.99 1.73% | 93.25 0.05% |
| Method | Configuration / Threshold | Final Sparsity (%) | Final Test Acc (%) | Peak Test Acc (%) |
|---|---|---|---|---|
| Dense Baseline | Unpruned (0%) | 0.00% | 80.62% | 82.68% |
| DADP (Ours) | 71.52 4.93% | 79.01 0.57% | 82.19 0.39% | |
| DADP (Ours) | 91.69 2.03% | 78.67 0.66% | 81.57 0.30% | |
| DADP (Ours) | 96.03 0.79% | 80.85 1.06% | 82.11 0.37% | |
| DADP (Ours) | (Over-threshold) | 96.87 0.16% | 64.18 12.70% | 64.22 12.76% |
| DADP (Ours) | (Over-threshold) | 96.96% | 50.92% | 50.92% |
| Method | Configuration / Threshold | Final Sparsity (%) | Final Test Acc (%) | Peak Test Acc (%) |
|---|---|---|---|---|
| Dense Baseline | Unpruned (0%) | 0.00% | 85.21% | 85.21% |
| DADP (Ours) | 74.26 1.25% | 85.09 0.17% | 85.11 0.16% | |
| DADP (Ours) | 89.77 0.08% | 84.70 0.17% | 85.47 0.34% | |
| DADP (Ours) | 91.65 0.67% | 82.65 0.65% | 84.29 0.53% | |
| DADP (Ours) | 91.86 0.54% | 83.22 0.00% | 83.22 0.00% | |
| DADP (Ours) | 92.16 0.28% | 83.06 2.01% | 84.02 1.36% |
| Layer Name | Original Shape | DADP Compressed Shape | SNIP Compressed Shape | Magnitude Compressed Shape | RigL Compressed Shape |
|---|---|---|---|---|---|
| features.0 | (0 dead) | (0 dead) | (0 dead) | (5 dead) | |
| features.3 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| features.7 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| features.10 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| features.14 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| features.17 | (0 dead) | (0 dead) | (0 dead) | (0 dead) |
| Layer Name | Original Shape | DADP Compressed Shape | SNIP Compressed Shape | Magnitude Compressed Shape | RigL Compressed Shape |
|---|---|---|---|---|---|
| conv1 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| layer1.0.conv1 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| layer1.0.conv2 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| layer1.1.conv1 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| layer1.1.conv2 | (0 dead) | (0 dead) | (0 dead) | (0 dead) | |
| layer2.0.conv1 | (0 dead) | (0 dead) | (0 dead) | (0 dead) |
| Model | Initialization Scheme | Pruning Threshold | Final Sparsity (%) | Final Test Acc (%) |
|---|---|---|---|---|
| ResNet-18 | Kaiming Normal | 95.41% | 75.10% | |
| ResNet-18 | Kaiming Uniform | 95.35% | 76.08% | |
| ResNet-18 | Xavier Normal | 95.66% | 77.05% | |
| ResNet-18 | Xavier Uniform | 95.54% | 76.72% | |
| ResNet-18 | Orthogonal | 95.75% | 76.51% | |
| ResNet-18 | Normal ( ) | 95.96% | 77.26% |
| Layer Index | ResNet-18 (CIFAR-10) | VGG-16 (CIFAR-10) | MiniBERT (SST-2) |
|---|---|---|---|
| 1 | conv1 | features.0 (conv1_1) | embeddings.word_embeddings |
| 2 | layer1.0.conv1 | features.3 (conv1_2) | embeddings.position_embeddings |
| 3 | layer1.0.conv2 | features.7 (conv2_1) | encoder.layer.0.attention.query |
| 4 | layer1.1.conv1 | features.10 (conv2_2) | encoder.layer.0.attention.key |
| 5 | layer1.1.conv2 | features.14 (conv3_1) | encoder.layer.0.attention.value |
| 6 | layer2.0.conv1 | features.17 (conv3_2) | encoder.layer.0.output.dense |