Neural Network Compression

Latest papers 128

All topics
CardsList
  1. Quantized Machine Learning Models for Medical Imaging in Low-Resource Healthcare Settings

    May 19, 2026Sumanth Meenan Kanneti, Aryan ShahEfficient Neural Network InferenceMedical Image Classification

  2. SparseSAM: Structured Sparsification of Activations in Segment Anything Models

    May 17, 2026Hoai-Chau Tran, Chi H. Nguyen, Duy M. H. Nguyen +3Efficient Transformer InferenceImage Segmentation

  3. Perforated Neural Networks for Keyword Spotting

    May 15, 2026Vishy Gopal, Aris Ilias Goutis, Ralph Crewe +2Keyword SpottingEdge ML

  4. IO-SVD: Input-Output Whitened SVD for Adaptive-Rank LLM Compression

    May 15, 2026Ali Abbasi, Chayne Thrash, Haoran Qin +2LLM CompressionLow-Rank Matrix Decomposition

  5. Two-Valued Symmetric Circulant Matrices: Applications in Deep Learning

    May 15, 2026Jayakrishna Amathi, Venkata Prasanth Yanambaka, Saraju P. Mohanty +1Structured SparsityEdge ML

  6. Neural Video Compression with Domain Transfer

    May 13, 2026Tiange Zhang, Rongqun Lin, Xiandong Meng +4Test-Time AdaptationCross-Domain Transfer Learning

  7. Robust Basis Spline Decoupling for the Compression of Transformer Models

    May 11, 2026Joppe De Jonghe, Van Tien Pham, Mariya IshtevaModel CompressionTensor Decomposition

  8. Evolving Knowledge Distillation for Lightweight Neural Machine Translation

    May 11, 2026Xuewen Zhang, Haixiao Zhang, Xinlong HuangMulti-Teacher Knowledge DistillationKnowledge Distillation

  9. Selection Plateau and a Sparsity-Dependent Hierarchy of Pruning Features

    May 10, 2026Guangqi Li, Yongxin LiSparse Neural NetworksNeural Network Compression

  10. Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning

    May 9, 2026Chen Wang, Siyu Hu, Guangming Tan +1Equivariant GNNsStructured Pruning

  11. SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

    May 9, 2026Shengkun Tang, Zekun Wang, Bo Zheng +7Mixture-of-Experts PruningLLM Pruning

  12. DiBA: Diagonal and Binary Matrix Approximation for Neural Network Weight Compression

    May 7, 2026Nobutaka OnoEfficient Neural Network InferenceMatrix Optimization

  13. Elastic Spiking Transformers for Efficient Gesture Understanding

    May 4, 2026Alberto Ancilotto, Gianluca Amprimo, Stefano Di Carlo +1Efficient Neural Network InferenceSpiking Transformers

  14. Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm

    May 2, 2026Wen-Da Wei, Han-Bin Fang, Yang-Di Liu +3LLM CompressionGradient Compression

  15. Auto-FlexSwitch: Efficient Dynamic Model Merging via Learnable Task Vector Compression

    Apr 30, 2026Junqi Gao, Dazhi Zhang, Zhichang Guo +3Task VectorsTask Vector Merging

  16. Parameter-Efficient Architectural Modifications for Translation-Invariant CNNs

    Apr 30, 2026Nuria Alabau-Bosque, Jorge Vila-Tomas, Paula Dauden-Oliver +2Neural Network RobustnessConvolutional Neural Networks

  17. QYOLO: Lightweight Object Detection via Quantum Inspired Shared Channel Mixing

    Apr 29, 2026Garvit Kumar Mittal, Sahil Tomar, Sandeep KumarEfficient Neural Network InferenceQuantum-Inspired ML

  18. Hierarchical Spatio-Channel Clustering for Efficient Model Compression in Medical Image Analysis

    Apr 25, 2026Sisipho Hamlomo, Marcellin Atemkeng, Habte Tadesse Likassa +5Low-Rank CompressionLow-Rank Approximation

  19. LTBs-KAN: Linear-Time B-splines Kolmogorov-Arnold Networks

    Apr 23, 2026Eduardo Said Merin-Martinez, Andres Mendez-Vazquez, Eduardo Rodriguez-TelloEfficient Neural Network InferenceKolmogorov-Arnold Networks

  20. Leveraging Kernel Symmetry for Joint Compression and Error Mitigation in Edge Model Transfer

    Apr 19, 2026Anis Hamadouche, Mathini SellathuraiNeural Network Compression

  21. Towards Joint Quantization and Token Pruning of Vision-Language Models

    Apr 19, 2026Xinqing Li, Xin He, Xindong Zhang +3Vision-Language ModelsVLM Quantization

  22. LASER: Low-Rank Activation SVD for Efficient Recursion

    Apr 19, 2026Ege Çakar, Ketan Ali Raghu, Lia ZhengLow-Rank Matrix DecompositionRecurrent Neural Networks

  23. A Comparative Study of CNN Optimization Methods for Edge AI: Exploring the Role of Early Exits

    Apr 16, 2026Nekane Fernandez, Ivan Valdes, Steven Van Vaerenbergh +2Efficient InferenceEdge Inference

  24. Exploiting Correlations in Federated Learning: Opportunities and Practical Limitations

    Apr 16, 2026Adrian Edin, Michel Kieffer, Mikael Johansson +1Communication-Efficient Distributed TrainingGradient Compression