LLM Pruning

LLM: Large Language Model

Momentum

12 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 113

All topics
CardsList
  1. From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression

    Jun 1, 2026Elia Cunegatti, Marcus Vukojevic, Erik Nielsen +1LLM PruningLLM Compression

  2. Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

    Jun 1, 2026Yujia Tong, Yuxi Wang, Yunyang Wan +3LLM QuantizationLLM Evaluation

  3. Before Parc Fermé: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving

    May 29, 2026Luca Benfenati, Ali Azimi, Matteo Risso +3LLM PruningLanguage Model-Based Control

  4. AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference

    May 28, 2026Yilin Feng, Ahmed Burak Gulhan, Mahmut Taylan KandemirLLM PruningVision-Language Models

  5. Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

    May 28, 2026Junlin He, Yihong Tang, Tong Nie +5LLM PruningLLM Compression

  6. Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

    May 27, 2026Karim Galliamov, Rochelle Choenni, Ivan TitovLLM PruningLLM Compression

  7. PrunePath: Towards Highly Structured Sparse Language Models

    May 27, 2026Zhexuan Gu, Zixun Fu, Yancheng YuanMixture-of-Experts PruningLLM Pruning

  8. Pruning and Distilling Mixture-of-Experts into Dense Language Models

    May 27, 2026Junhyuck Kim, Jihun Yun, Haechan Kim +3Mixture-of-Experts PruningLLM Pruning

  9. Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts

    May 27, 2026Liu O. Martin, Lucas Bandarkar, Nanyun PengMixture-of-Experts PruningLLM Pruning

  10. Locality-Aware Redundancy Pruning for LLM Depth Compression

    May 27, 2026Vincent-Daniel Yun, Youngrae Kim, Woosang Lim +3LLM PruningLLM Compression

  11. Dense2MoE: Pushing the Pareto Frontier of On-Device LLMs via Unified Pruning and Upcycling

    May 26, 2026Fengfa Li, Hongjin Ji, Yifeng Ding +2Mixture-of-Experts PruningLLM Pruning

  12. Prune, Update and Trim: Robust Structured Pruning for Large Language Models

    May 18, 2026Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidt-ThiemeLLM PruningLLM Inference Acceleration

  13. LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models

    May 17, 2026Mohammad Mozaffari, Younes Hourri, Mohammad Rastegari +1LLM PruningNeural Network Pruning

  14. Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs

    May 15, 2026Vincent-Daniel Yun, Junhyuk Jo, Sai Praneeth Karimireddy +1LLM PruningLLM Compression

  15. Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning

    May 14, 2026Jingjing Wang, Xiwen Chen, Wenhui Zhu +6LLM PruningAI Coding Agents

  16. TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability

    May 14, 2026Krish Sharma, Omar Naim, Soumadeep Saha +3Representation GeometryLLM Pruning

  17. Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

    May 13, 2026Cristian Hinostroza, Rodrigo Toro Icarte, Christ Devia +4LLM PruningLLM Interpretability

  18. STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes

    May 13, 2026Chenjun Xu, Zhennan Zhou, Zhan Su +3LLM PruningLLM Fine-Tuning

  19. Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs

    May 12, 2026Jingzhou Jiang, Yi Yang, Kar Yan TamText EmbeddingsLLM Pruning

  20. Compute Where it Counts: Self Optimizing Language Models

    May 11, 2026Yash Akhauri, Mohamed S. AbdelfattahLLM Inference EfficiencyRL for Language Models

  21. A Game Theoretic Free Energy Analysis of Higher Order Synergy in Attention Heads of Large Language Models

    May 10, 2026Djamel BouchaffraLLM PruningAttention Head Analysis

  22. Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning

    May 9, 2026Tianhao Qian, Guilin Qi, Jiayu ChenLLM PruningStructured Sparsity

  23. ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference

    May 9, 2026Junjie Li, Jiong Lou, Jie LiLLM PruningKV Caching

  24. SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

    May 9, 2026Shengkun Tang, Zekun Wang, Bo Zheng +7Mixture-of-Experts PruningLLM Pruning

  25. Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

    May 8, 2026Boyu Shi, Chang Liu, ChuanBao Gao +2LLM PruningLLM Interpretability

  26. SparseForge: Efficient Semi-Structured LLM Sparsification via Annealing of Hessian-Guided Soft-Mask

    May 7, 2026Liu Hanzuo, Chaofan Lin, Weixuan Sun +4LLM PruningSparse Recovery

  27. Importance-Guided Basis Selection for Low-Rank Decomposition of Large Language Models

    May 2, 2026Daniel Agyei Asante, Ernie Chang, Yang LiLLM PruningLLM Compression