LLM Pruning

LLM: Large Language Model

Momentum

12 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 113

All topics
CardsList
  1. TAP: Efficient Long-Horizon Agent Pruning via Trajectory-Anchored Recovery

    Oct 6, 2026Yuanzhe Li, Pengxin Wang, Yuxin Ren +5LLM PruningLLM Inference Acceleration

  2. A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

    Oct 6, 2026Davis Wertheimer, Haochen Shen, Ahan Gupta +6KV-Cache CompressionTransformer Attention

  3. MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

    Oct 1, 2026Xudong Wang, Hao Wu, Haozhe Hu +5LLM PruningEfficient VLM Inference

  4. Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs

    Sep 29, 2026Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi ShekharLLM PruningAlgorithmic Fairness

  5. Output-aware Residual Stream Pruning for Large Language Models

    Sep 28, 2026Chayne Thrash, Kevin Chen, Soheil KolouriLLM PruningLLM Compression

  6. GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning

    Sep 27, 2026Zhengao Li, Shuoqiu Li, Xiaofang Zhang +7LLM PruningLLM Compression

  7. RAZOR: Pruning Replaceable Experts in LLMs

    Sep 24, 2026Mingyang Song, Mao ZhengMixture-of-Experts PruningLLM Pruning

  8. Speech Block Influence: Component-Specific Layer Scoring for Pruning Speech LLMs

    Sep 24, 2026Siyu Yao, Du Q. Huynh, Lian Xu +1LLM PruningSpeech Language Models

  9. OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    Sep 15, 2026Ha Lan Nguyen, Huy Hoang Tran, Trac-Duy Tran +1LLM PruningLarge Reasoning Models

  10. What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

    Sep 15, 2026Congjing Zhang, Vashishtha Patil, Henning Lange +1LLM GroundingLLM Pruning

  11. ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression

    Sep 14, 2026Liu O. Martin, Lucas Bandarkar, Nanyun PengModel CompressionMixture-of-Experts Pruning

  12. LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

    Sep 11, 2026Sankar Behera, Dhruv Singh, Anshika Agnihotri +3LLM PruningLLM Compression

  13. Forward-Free LLM Depth Pruning via Weight Redundancy

    Sep 9, 2026Vincent-Daniel Yun, Woosang LimLLM PruningLLM Inference Acceleration

  14. Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

    Sep 2, 2026Irina Proskurina, Guillaume Metzler, Antoine Gourru +1LLM PruningNeural Network Pruning

  15. Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

    Aug 13, 2026Palaash Goel, Ayan Sengupta, Akshay Nambi +1LLM PruningStructured Pruning

  16. Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

    Aug 12, 2026Haokun Lin, Kaijie Zhu, Haobo Xu +4LLM QuantizationLLM Pruning

  17. The Sparsity Whisperer

    Aug 6, 2026Linghao Kong, Inimai Subramanian, Micah Adler +3LLM PruningLLM Inference Acceleration

  18. BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning

    Aug 5, 2026Sajib Hossain, Md Kamrus Samad, Anan Ghosh +2LLM PruningFew-Shot Learning

  19. TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning

    Aug 4, 2026Wonpyo Park, Seung-won HwangLLM PruningKV Caching

  20. WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

    Jul 30, 2026Haozhe Hu, Hao Wu, Peiran Yin +3LLM PruningToken Pruning

  21. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

    Jul 30, 2026Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña LópezLLM PruningLLM Interpretability

  22. Unified Static-Dynamic Pruning for Efficient LLM Inference

    Jul 24, 2026Jinhyeok Kim, Yejoon Lee, Jaeyoung DoEfficient Neural Network InferenceLLM Pruning

  23. Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

    Jul 23, 2026Yidu Wu, Xiang Wang, Kejie Zhao +3LLM PruningLLM Inference Acceleration

  24. CausalGate: Causal Importance Distillation for Transformer Module Pruning

    Jul 21, 2026Kiran Nair, Smriti Regmi, Rodrigue RizkEfficient Transformer InferenceLLM Pruning

  25. SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

    Jul 20, 2026Yuhang Wang, Yuling Shi, Shaoqiu Zhang +6LLM PruningAI Coding Agents

  26. CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning

    Jul 20, 2026Zhiren Gong, Zihao Zeng, Zijie Wang +3LLM PruningLLM Compression

  27. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

    Jul 14, 2026Qingyu Zhang, Qianhao Yuan, Hongyu Lin +7Open-Ended GenerationLLM Pruning

  28. Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

    Jul 9, 2026Bishmoy Paul, Youngmin Yi, Hoeseok YangLLM PruningActivation Sparsity

  29. Super Weights in LLMs and the Failure of Selective Training

    Jul 9, 2026Shreyas Subramanian, Adewale Akinfaderin, Akarsha SehwagLLM PruningFine-Tuning

  30. It Takes a MAESTRO To Prune Bad Experts

    Jul 9, 2026Palaash Goel, Ayush Maheshwari, Tanmoy ChakrabortyMixture-of-Experts PruningLLM Pruning

  31. Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

    Jul 9, 2026Ryota Kobayashi, Tsubasa Hirakawa, Takayoshi Yamashita +4LLM PruningLLM Compression

  32. PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

    Jul 8, 2026Yazdan Jamshidi, Alexey ShvetsLLM PruningLLM Compression

  33. Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

    Jul 5, 2026Akhiad Bercovich, Talor Abramovich, Daniel Afrimi +67Mixture-of-Experts PruningLLM Pruning

  34. FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

    Jun 26, 2026Fan Mo, Yuxuan Han, Geng Zhang +2Mixture-of-Experts PruningLLM Pruning

  35. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT

    Jun 25, 2026Jinghan Wang, Yanjun Chen, Wei Zhang +3LLM PruningOn-Device Language Model Inference

  36. Scaling Laws for Task-Specific LLM Distillation

    Jun 23, 2026Lavinia Ghita, Dhruv Desai, Ioana BoierLLM PruningLLM Compression

  37. Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

    Jun 23, 2026Pietro Tropeano, Maria Maistro, Tuukka Ruotsalo +1LLM PruningLLM Interpretability

  38. SVD-Surgeon: Optimal Singular-Value Surgery for Large Language Model Compression

    Jun 22, 2026Mahmoud Safari, Frank HutterLLM PruningLLM Compression

  39. The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer

    Jun 16, 2026Rui Wen, Lu Sun, Jiayang Liu +3Open-Ended GenerationLLM Evaluation

  40. Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

    Jun 16, 2026Yifu Ding, Jiacheng Wang, Ge Yang +4Mixture-of-Experts PruningLLM Pruning

  41. How to Score Experts for One-Shot MoE Expert Pruning: A Unified Formulation and Selection Principle

    Jun 14, 2026Zongfang Liu, Jinghui Zhang, Zijian Ma +2Mixture-of-Experts PruningLLM Pruning

  42. Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

    Jun 13, 2026Tao Lu, Haoyu Wang, Zonghui Wang +3LLM PruningGPU Kernel Optimization

  43. Persona-Pruner: Sculpting Lightweight Models for Role-Playing

    Jun 12, 2026Jinsu Kim, Jihoon Tack, Noah Lee +1LLM PruningLarge Language Model-Based Role-Play Simulation

  44. Small LLMs: Pruning vs. Training from Scratch

    Jun 12, 2026Yufeng Xu, Taiming Lu, Kunjun Li +3Language Model PretrainingLLM Pruning

  45. SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference

    Jun 9, 2026Jaeseong Lee, Seung-won Hwang, Samyam RajbhandariLLM PruningHigh-Performance Computing

  46. BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference

    Jun 8, 2026Yuhua Zhou, Shaoqi Yu, Shichao Weng +4LLM PruningCost-Aware Inference

  47. Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

    Jun 8, 2026Haozhe Hu, Hao Wu, Anhao Zhao +4LLM PruningLLM Inference Acceleration

  48. Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression

    Jun 5, 2026Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1Mixed-Precision QuantizationLLM Quantization

  49. Less is MoE: Trimming Experts in Domain-Specialist Language Models

    Jun 4, 2026Haoze He, Xinkai Zou, Xuan Jiang +4Mixture-of-Experts PruningLLM Pruning

  50. TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

    Jun 3, 2026Jiangyang He, Shaolin Zhu, Deyi XiongMixture-of-Experts PruningLLM Pruning