LLM Pruning

LLM: Large Language Model

Momentum

12 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 113

All topics
CardsList
  1. SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

    Jul 20, 2026Yuhang Wang, Yuling Shi, Shaoqiu Zhang +6LLM PruningAI Coding Agents

  2. CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning

    Jul 20, 2026Zhiren Gong, Zihao Zeng, Zijie Wang +3LLM PruningLLM Compression

  3. ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

    Jul 14, 2026Qingyu Zhang, Qianhao Yuan, Hongyu Lin +7Open-Ended GenerationLLM Pruning

  4. Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

    Jul 9, 2026Bishmoy Paul, Youngmin Yi, Hoeseok YangLLM PruningActivation Sparsity

  5. Super Weights in LLMs and the Failure of Selective Training

    Jul 9, 2026Shreyas Subramanian, Adewale Akinfaderin, Akarsha SehwagLLM PruningFine-Tuning

  6. It Takes a MAESTRO To Prune Bad Experts

    Jul 9, 2026Palaash Goel, Ayush Maheshwari, Tanmoy ChakrabortyMixture-of-Experts PruningLLM Pruning

  7. Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

    Jul 9, 2026Ryota Kobayashi, Tsubasa Hirakawa, Takayoshi Yamashita +4LLM PruningLLM Compression

  8. PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

    Jul 8, 2026Yazdan Jamshidi, Alexey ShvetsLLM PruningLLM Compression

  9. Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

    Jul 5, 2026Akhiad Bercovich, Talor Abramovich, Daniel Afrimi +67Mixture-of-Experts PruningLLM Pruning

  10. FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

    Jun 26, 2026Fan Mo, Yuxuan Han, Geng Zhang +2Mixture-of-Experts PruningLLM Pruning

  11. Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT

    Jun 25, 2026Jinghan Wang, Yanjun Chen, Wei Zhang +3LLM PruningOn-Device Language Model Inference

  12. Scaling Laws for Task-Specific LLM Distillation

    Jun 23, 2026Lavinia Ghita, Dhruv Desai, Ioana BoierLLM PruningLLM Compression

  13. Don't Go Breaking My LLM: The Impact of Pruning Attention Layers on Explanation Faithfulness and Confidence Calibration

    Jun 23, 2026Pietro Tropeano, Maria Maistro, Tuukka Ruotsalo +1LLM PruningLLM Interpretability

  14. SVD-Surgeon: Optimal Singular-Value Surgery for Large Language Model Compression

    Jun 22, 2026Mahmoud Safari, Frank HutterLLM PruningLLM Compression

  15. The Benchmark Illusion: Pruned LLMs Can Pass Multiple Choice but Fail to Answer

    Jun 16, 2026Rui Wen, Lu Sun, Jiayang Liu +3Open-Ended GenerationLLM Evaluation

  16. Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

    Jun 16, 2026Yifu Ding, Jiacheng Wang, Ge Yang +4Mixture-of-Experts PruningLLM Pruning

  17. How to Score Experts for One-Shot MoE Expert Pruning: A Unified Formulation and Selection Principle

    Jun 14, 2026Zongfang Liu, Jinghui Zhang, Zijian Ma +2Mixture-of-Experts PruningLLM Pruning

  18. Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

    Jun 13, 2026Tao Lu, Haoyu Wang, Zonghui Wang +3LLM PruningGPU Kernel Optimization

  19. Persona-Pruner: Sculpting Lightweight Models for Role-Playing

    Jun 12, 2026Jinsu Kim, Jihoon Tack, Noah Lee +1LLM PruningLarge Language Model-Based Role-Play Simulation

  20. Small LLMs: Pruning vs. Training from Scratch

    Jun 12, 2026Yufeng Xu, Taiming Lu, Kunjun Li +3Language Model PretrainingLLM Pruning

  21. SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference

    Jun 9, 2026Jaeseong Lee, Seung-won Hwang, Samyam RajbhandariLLM PruningHigh-Performance Computing

  22. BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference

    Jun 8, 2026Yuhua Zhou, Shaoqi Yu, Shichao Weng +4LLM PruningCost-Aware Inference

  23. Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy

    Jun 8, 2026Haozhe Hu, Hao Wu, Anhao Zhao +4LLM PruningLLM Inference Acceleration

  24. Joint Structural Pruning and Mixed-Precision Quantization for LLM Compression

    Jun 5, 2026Hoang-Loc La, Truong-Thanh Le, Amir Taherkordi +1Mixed-Precision QuantizationLLM Quantization

  25. Less is MoE: Trimming Experts in Domain-Specialist Language Models

    Jun 4, 2026Haoze He, Xinkai Zou, Xuan Jiang +4Mixture-of-Experts PruningLLM Pruning

  26. TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

    Jun 3, 2026Jiangyang He, Shaolin Zhu, Deyi XiongMixture-of-Experts PruningLLM Pruning