Knowledge editing changes what a model knows by modifying parameters so that a requested fact updates while unrelated behavior is preserved. This is usually treated as a write problem, but editing also involves an address problem: deciding which hidden states should receive the new residual. An update that activates too narrowly memorizes one prompt, while one that activates too broadly disrupts neighboring knowledge. Parametric editors encode this scope implicitly, whereas memory-based editors make the selection explicit but keep it outside the edited model. We propose ALOE (Addressed Low-rank Operator for Editing), which learns semantic addresses from paraphrases and hard same-subject negatives, aligns them with autoregressive hidden states through rollout refinement and gate calibration, and embeds the resulting gated low-rank operator within one MLP layer, so that the deployed model runs in a single forward pass with no external retriever or auxiliary router. Evaluated on CounterFact, ZSRE, and KnowEdit across three 7--8B model families, ALOE achieves efficacy between 0.955 and 0.999 and locality between 0.981 and 1.000; mechanistic analyses confirm that the learned geometry separates competing edits and that calibration suppresses out-of-scope activation. The remaining errors concentrate in paraphrase coverage and write fitting.
Figures & tables
Figure 1: The address problem. The same write can miss paraphrases or affect neighboring facts. Filled cells denote an active write; empty cells denote no update (schematic).
Figure 2: ALOE learns semantic addresses, refines and calibrates continuous edit gates, and inserts the resulting low-dimensional residual into one MLP projection. The edited LLM executes the operator in a standard forward pass.
ALOE
ROME (2022)
MEMIT (2023)
MEND (2022)
GRACE (2023)
AlphaEdit (2025)
Benchmark
Eff.
Gen.
Loc.
Eff.
Gen.
Loc.
Eff.
Gen.
Loc.
Eff.
Gen.
Loc.
Eff.
Gen.
Loc.
Eff.
Gen.
Loc.
Llama-2-7B
CounterFact
0.956
0.289
0.982
0.093
0.083
0.014
0.056
0.188
0.030
0.011
0.012
0.005
0.949
0.003
0.972
0.821
0.129
0.838
KnowEdit
0.999
0.472
1.000
0.121
0.106
0.028
0.162
0.150
0.195
0.002
0.002
0.002
0.980
0.086
1.000
0.969
0.801
1.000
ZSRE
0.997
0.464
1.000
0.204
0.183
0.019
0.144
0.138
0.137
0.003
0.003
0.003
0.974
0.009
1.000
0.977
0.891
1.000
Llama-3.1-8B-Instruct
Table 1: Efficacy (Eff.), generalization (Gen.), and locality (Loc.) across base models and benchmarks. ALOE fits one operator to the complete benchmark stream (839 CounterFact, 1,266 KnowEdit, and 1,301 ZSRE edits); all results are means over seeds 0, 42, and 99. Best and second-best values are bolded and underlined , respectively.
Figure 3: Runtime behavior after 839 CounterFact edits on Llama-3.1-8B-Instruct. (a) Matched-slot gates for edits and rephrases, and the maximum gate for locality queries. (b) Responses of the first 32 sampled queries to their associated slots. (c) Paired final-block locality states under shared PCA, with empirical marginals and a full-range inset. PC1/PC2 explain 45.1%/10.7%; mean relative state drift is 11.54% and next-token agreement is 96.5%.
Knowledge editing aims to efficiently update factual information in Large Language Models (LLMs) without full retraining. However, existing methods still suffer from performance degradation in batch knowledge editing. We identify that semantic representation entanglement, such as overlapping concepts and shared syntactic patterns, accumulates interference in the representation space and reduces editing precision. To bridge this gap, in this paper, we propose Orthogonal Representation Editing (ORE), which performs edits in the hidden representation space of LLMs by constructing a general semantic subspace and enforcing orthogonal constraints on edit vectors, effectively decoupling semantic entanglement. Furthermore, we introduce a gated non-linear representation head to enable adaptive learning of editing locations and precise control over knowledge injection. Extensive experiments show that ORE outperforms existing methods and achieves superior performance in cross-lingual knowledge editing scenarios. We release our code at https://github.com/YVVH/ORE.
Wenhao Yu, Zhicong Lu, Bo Lv +4
School of Computer Science and Technology, Tianjin University · Kexin Technology · University of Chinese Academy of Sciences +2
Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited knowledge is often not fully erased and continues to surface, with consistent failures observed across diverse model architectures. To explain this behavior, we conduct a mechanistic analysis of popular KE methods. We show that low-rank updates do not overwrite existing knowledge but instead redistribute it within the model's representation space. Furthermore, we find that these methods act as targeted suppression mechanisms that reduce the likelihood of expressing original facts, rather than removing them from the model. Analysis of the loss landscape reveals that edited knowledge lies in narrow, anisotropic regions that are highly sensitive to perturbations, making them highly vulnerable to indirect prompting and adversarial attacks. By exposing these profound architectural vulnerabilities, our work proves that KE algorithms are inherently bypassable and motivates a fundamental reevaluation of how we deploy post-hoc updates in several LLM applications.
Advik Raj Basani, Anshuman Chhabra
Birla Institute of Technology and Science, Goa · University of South Florida
Knowledge editing systems must update selected facts while preserving nearby but irrelevant behavior. This paper studies this problem in a memory-assisted setting where an edit memory is retrieved at inference time and a parameter-efficient adapter corrects the model's object preference. We argue that the central design question is not only how to write an edit, but also when to suppress it. We introduce RRDA, a route-specialized dual-adapter editor. A relevance router first decides whether a prompt should receive an edit memory. Routed prompts use an edit adapter trained to prefer the new object over the original object; unrouted non-direct prompts use a separate locality adapter trained to preserve or restore the original-object preference. We evaluate RRDA on three 1,000-case protocols, CounterFact, ZsRE, and MQuAKE-CF, under the same memory protocol and two 7B/8B base models. On Llama-3.1-8B-Instruct, RRDA obtains the best overall probability-preference accuracy on all three benchmarks: 0.8180 on CounterFact, 0.8946 on ZsRE, and 0.9922 on MQuAKE-CF. The same trend holds on Qwen3-8B. Router ablations show that the relevant memory boundary differs across datasets: a lexical neural router is safest on CounterFact, while BGE embedding routing is better on ZsRE and MQuAKE-CF. Memory, component, and module ablations show that explicit memory supplies the largest edit gain, while route-specialized adapters improve the final reliability-locality balance rather than simply increasing LoRA capacity.