cs.AISep 28, 2026

TULIP: Targeted LLM Unlearning at Layers Identified Per-Input

Authors: Yejin Kim, William F. Shen, Seokwon Jung, Daeun Park, Seong Joon Oh

Organizations: Korea Advanced Institute of Science & Technology (KAIST) · University of Cambridge · Sookmyung Women’s University

Abstract

Representation-level unlearning intervenes on the intermediate hidden states of LLMs. Although knowledge is distributed across layers, existing methods operate at a single fixed layer for the entire forget set. We ask whether such a fixed layer is sufficient. To answer this, we design a hijacking experiment that grafts hidden states of the target model into an oracle trained only on the retain set. The oracle cannot produce the forget answer on its own, yet it produces the answer from the grafted state. Thus, the answer is formed at an intermediate layer and merely read out afterward, so unlearning should focus on formation, not readout. Moreover, the layer where formation ends varies widely across inputs. Motivated by these findings, we propose Targeted Unlearning at Layers Identified Per-input (TULIP). For each input, TULIP uses the logit lens to locate the formation-readout boundary and removes the hidden state's alignment with the forget answer's unembedding vector there. TULIP consistently outperforms output- and representation-level baselines on TOFU, PISTOL, and WMDP across Llama, Qwen, and Zephyr models. It also remains robust to paraphrase and quantization attacks. Beyond standalone use, its per-input layer selection serves as a plug-and-play component that further improves existing methods.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

    Sep 9, 2026Ravi Ranjan, Olivera Kotevska, Agoritsa PolyzouLarge Language Model UnlearningLarge Language Model Quantization

  2. RepSelect: Robust LLM Unlearning via Representation Selectivity

    Jun 15, 2026Filip Sondej, Yushi Yang, Adam MahdiLarge Language Model UnlearningSelectivity

  3. LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

    Jul 2, 2026Matteo Boglioni, Thibault Rousset, Siva Reddy +2Large Language Model UnlearningAttacker Large Language Model