LLM Pruning

LLM: Large Language Model

Momentum

12 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 116

All topics
CardsList
  1. Less is MoE: Trimming Experts in Domain-Specialist Language Models

    Jun 4, 2026Haoze He, Xinkai Zou, Xuan Jiang +4Mixture-of-Experts PruningLLM Pruning

  2. TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts

    Jun 3, 2026Jiangyang He, Shaolin Zhu, Deyi XiongMixture-of-Experts PruningLLM Pruning

  3. From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression

    Jun 1, 2026Elia Cunegatti, Marcus Vukojevic, Erik Nielsen +1LLM PruningLLM Compression

  4. Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

    Jun 1, 2026Yujia Tong, Yuxi Wang, Yunyang Wan +3LLM QuantizationLLM Evaluation

  5. Before Parc Fermé: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving

    May 29, 2026Luca Benfenati, Ali Azimi, Matteo Risso +3LLM PruningLanguage Model-Based Control

  6. AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference

    May 28, 2026Yilin Feng, Ahmed Burak Gulhan, Mahmut Taylan KandemirLLM PruningVision-Language Models

  7. Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

    May 28, 2026Junlin He, Yihong Tang, Tong Nie +5LLM PruningLLM Compression

  8. Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

    May 27, 2026Karim Galliamov, Rochelle Choenni, Ivan TitovLLM PruningLLM Compression

  9. PrunePath: Towards Highly Structured Sparse Language Models

    May 27, 2026Zhexuan Gu, Zixun Fu, Yancheng YuanMixture-of-Experts PruningLLM Pruning

  10. Pruning and Distilling Mixture-of-Experts into Dense Language Models

    May 27, 2026Junhyuck Kim, Jihun Yun, Haechan Kim +3Mixture-of-Experts PruningLLM Pruning

  11. Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts

    May 27, 2026Liu O. Martin, Lucas Bandarkar, Nanyun PengMixture-of-Experts PruningLLM Pruning

  12. Locality-Aware Redundancy Pruning for LLM Depth Compression

    May 27, 2026Vincent-Daniel Yun, Youngrae Kim, Woosang Lim +3LLM PruningLLM Compression

  13. Dense2MoE: Pushing the Pareto Frontier of On-Device LLMs via Unified Pruning and Upcycling

    May 26, 2026Fengfa Li, Hongjin Ji, Yifeng Ding +2Mixture-of-Experts PruningLLM Pruning

  14. Prune, Update and Trim: Robust Structured Pruning for Large Language Models

    May 18, 2026Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidt-ThiemeLLM PruningLLM Inference Acceleration

  15. LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models

    May 17, 2026Mohammad Mozaffari, Younes Hourri, Mohammad Rastegari +1LLM PruningNeural Network Pruning

  16. Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

    May 15, 2026Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy +1LLM PruningLLM Compression

  17. Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning

    May 14, 2026Jingjing Wang, Xiwen Chen, Wenhui Zhu +6LLM PruningAI Coding Agents

  18. TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability

    May 14, 2026Krish Sharma, Omar Naim, Soumadeep Saha +3Representation GeometryLLM Pruning

  19. Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

    May 13, 2026Cristian Hinostroza, Rodrigo Toro Icarte, Christ Devia +4LLM PruningLLM Interpretability

  20. STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes

    May 13, 2026Chenjun Xu, Zhennan Zhou, Zhan Su +3LLM PruningLLM Fine-Tuning

  21. Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs

    May 12, 2026Jingzhou Jiang, Yi Yang, Kar Yan TamText EmbeddingsLLM Pruning

  22. Compute Where it Counts: Self Optimizing Language Models

    May 11, 2026Yash Akhauri, Mohamed S. AbdelfattahLLM Inference EfficiencyRL for Language Models

  23. A Game Theoretic Free Energy Analysis of Higher Order Synergy in Attention Heads of Large Language Models

    May 10, 2026Djamel BouchaffraLLM PruningAttention Head Analysis

  24. Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning

    May 9, 2026Tianhao Qian, Guilin Qi, Jiayu ChenLLM PruningStructured Sparsity

  25. ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

    May 9, 2026Junjie Li, Jiong Lou, Jie LiLLM PruningKV Caching

  26. SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

    May 9, 2026Shengkun Tang, Zekun Wang, Bo Zheng +7Mixture-of-Experts PruningLLM Pruning