Attention Head Pruning

Momentum

1 paper in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 13

All topics
CardsList
  1. CASS: Contribution-Aware Structured Sparsity for Model Merging

    Sep 28, 2026Yan Li, Guiping Cao, Meng Xu +5Multi-Task LearningStructured Sparsity

  2. Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

    Jul 31, 2026Kamil Książek, Piotr Suszyński, Michał Jan Włodarczyk +2Transformer InterpretabilityAttention Head Analysis

  3. WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

    Jul 30, 2026Haozhe Hu, Hao Wu, Peiran Yin +3LLM PruningToken Pruning

  4. Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

    Jul 21, 2026Maohua Li, Qirui Li, Yanke Zhou +10Transformer InterpretabilityDiffusion Transformer

  5. CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning

    Jul 20, 2026Zhiren Gong, Zihao Zeng, Zijie Wang +3LLM PruningLLM Compression

  6. MobileWan: Closing the Quality Gap for Mobile Video Diffusion

    Jul 7, 2026Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv +9Video Diffusion ModelsVideo Generation

  7. Complementary Attention Head Pruning for Efficient Transformers

    Jun 17, 2026Yaniv Livertovsky, Shahar Somin, Gonen SingerModel CompressionTransformer Attention

  8. From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

    May 20, 2026Hao Chen, Qi Zhang, Liyao Li +7LLM AlignmentLLM Fine-Tuning

  9. Prune, Update and Trim: Robust Structured Pruning for Large Language Models

    May 18, 2026Diego Coello de Portugal Mecke, Tom Hanika, Lars Schmidt-ThiemeLLM PruningLLM Inference Acceleration

  10. Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models

    Apr 27, 2026Amogh Sheth, Biruk Assefa, Yi Wen Huang +2LLM PruningSelf-Attention

  11. HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers

    Mar 12, 2026Andy Li, Aiden Durrant, Milan Markovic +1Efficient ViTsAttention Head Pruning

  12. Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning

    Jan 11, 2026Jaewon Sok, Jewon Yeom, Seonghyeon Park +2LLM PruningAttention Head Analysis

  13. Diffract: Spectral View of LLM Domain Adaptation

    Date pendingNikita Borodin, Maria Krylova, Artem Zabolotnyi +6Domain AdaptationDomain-Adaptive Pretraining