cs.AIOct 8, 2026

Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention

Authors: Zhiqiang Pang, Zihong Sun, Qi Xie, Jun Shu, Deyu Meng, Zongben Xu

Organizations: School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an, Shaanxi, China

Abstract

Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining. Existing approaches primarily mitigate this trade-off through data replay or regularization, relying on additional data or explicit optimization constraints. We instead focus on a different question: where should adaptation be applied? We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation. To characterize this difference, we use layer-wise empirical Fisher information to measure target-task sensitivity. However, computing Fisher scores requires backward computation and becomes increasingly expensive for large models. We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity. Across models and tasks, layers with lower input--output similarity consistently exhibit higher empirical Fisher scores. Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity. Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Fora: From Weight-Space to Function-Space Protection in Capability-Preserving Fine-Tuning

    Jun 30, 2026Rui Zhou, Tianci XieFine-TuningLLM Fine-Tuning

  2. Strategic Over-Parameterization for Generalizable Low-Rank Adaptation

    May 15, 2026Jing Gao, Zhong-Yi Lu, Pan Zhang +1LLM Fine-TuningLow-Rank Adaptation