IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs
Organizations: AImageLab, University of Modena and Reggio Emilia, Italy
Abstract
We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at https://github.com/aimagelab/IrekoGPT
Figures & tables
| Sparsity | Dense | Sparsity | 25% | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Original | SliceGPT † | ||||||||
| PCA | PCA | ||||||||
| ECD | ECD | ||||||||
| IrekoGPT | IrekoGPT | ||||||||
| Sparsity | 10% | Sparsity | 30% | ||||||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.