Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In such settings, existing editing methods often overfit to target prompts, significantly degrading both the generalization of the edited knowledge and the model's general capabilities. To address this issue, we propose GLIME (Generalizable Lifelong Model Editing), which combines knowledge editing with preference optimization over generation behavior. GLIME further incorporates replay-based editing and a gradient constraint to preserve previously edited knowledge. Experimental results show that GLIME significantly improves knowledge generalization in lifelong editing settings while maintaining both editing performance and general capabilities.
Figures & tables
Figure 1: Overall architecture of GLIME. Left : The model is updated over a stream of continual knowledge edits using new knowledge injection, preference optimization for behavior-level regularization, and experience replay to prevent catastrophic forgetting. The gradient is further projected orthogonally to previous update directions to mitigate long-term collapse under continual edits. Bottom right : The final edited model successfully answers paraphrased queries under autoregressive decoding, indicating improved generalization.
MQuAKE
ZSRE
Reliability
Generalization
Locality
Portability
Reliability
Generalization
Locality
Method
TF
AD
TF
AD
TF
AD
MC
MR
AVG
TF
AD
TF
AD
TF
AD
AVG
LLaMA-3.1-8B-Instruct
FT-L
0.021
0.035
0.012
0.038
0.004
0.000
0.000
0.003
0.016
0.134
0.106
0.114
0.094
0.023
0.000
0.078
R-ROME
0.026
0.074
0.009
0.087
0.005
0.010
0.000
0.008
0.030
0.033
0.004
0.028
0.087
0.000
0.000
0.025
GRACE
0.310
0.035
0.203
0.046
0.571
0.714
0.010
0.060
0.270
0.378
0.021
0.314
0.014
0.388
0.203
0.220
Table 1: Lifelong knowledge editing results. We compare different editing methods on the MQuAKE and ZSRE benchmarks. For each dataset, we report Reliability, Generalization, and Locality. We additionally evaluate both teacher forcing (TF) and autoregressive decoding (AD), which reflects real generation settings. On MQuAKE, we further measure Portability using Multiple-Choice QA (MC) and Multi-hop Reasoning QA (MR), and report the overall average (AVG) across the presented metrics. Best results per model group are in bold .
CPO
Replay
GC
Reliability
Generalization
Locality
Portability MC
Portability MR
✓
✓
✓
0.830
0.800
0.379
0.633
0.184
✓
✓
0.826 (-0.004)
0.745 (-0.055)
0.187 (-0.192)
0.014 (-0.619)
0.112 (-0.072)
✓
✓
0.657 (-0.173)
0.633 (-0.167)
0.326 (-0.053)
0.531 (-0.102)
0.096 (-0.088)
✓
✓
0.800 (-0.030)
0.786 (-0.014)
0.385 (+0.006)
0.412 (-0.221)
0.098 (-0.086)
✓
0.597 (-0.233)
0.601 (-0.199)
0.391 (+0.012)
0.530 (-0.103)
0.168 (-0.016)
✓
0.624 (-0.206)
0.615 (-0.185)
0.137 (-0.242)
0.063 (-0.570)
0.109 (-0.075)
Table 2: Ablation results for GLIME across all combinations of its core components. Values in parentheses indicate performance changes relative to the full GLIME. CPO denotes Continual Preference Optimization, Replay denotes the replay loss, and GC denotes the gradient-space constraint.
Method
Winogrande
ARC
MathQA
Squadv2
IFEval
Base
0.748
0.802
0.390
0.503
0.532
FT-L
0.248
0.270
0.188
0.002
0.171
R-ROME
0.514
0.234
0.226
0.000
0.229
GRACE
0.748
0.802
0.390
0.503
0.532
MEMIT
0.500
0.244
0.178
0.000
0.184
WISE
0.738
0.778
0.394
0.129
0.209
Table 3: General capability results. We evaluate each editing method on representative benchmarks unrelated to the edited knowledge. Base denotes the performance of the model before editing, while the other results are measured on the final model after accumulated edits.
Method
FT-L
R-ROME
GRACE
MEMIT
WISE
AlphaEdit
GLIME
Seconds
2.574
8.827
5.975
27.826
7.494
16.431
8.936
Table 4: Computational cost comparison. We report the time required for a single edit for each method.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Task Category
Input Prompt ( x )
Target Output ( y )
MQuAKE
Base Edit (Rel.)
jazz was created in the country of
Indonesia
Generalization
Which country was jazz created in?
Indonesia
Locality
What type of music does Hamid Drake play?
jazz
Portability (MC)
jazz was created in the country of Options: (A) Indonesia, (B) uncertain, (C) United States of America, (D) jazz
(A)
Portability (MR)
In which country was the music genre played by Hamid Drake created?
Indonesia
ZSRE
Base Edit (Rel.)
What programming language was used to write OpenCV?
Java
Appendix
Table 5: Example knowledge editing instances from the MQuAKE and ZSRE datasets.
MQuAKE
ZSRE
Reliability
Generalization
Locality
Reliability
Generalization
Locality
Target True
0.646
0.418
0.571
0.438
0.428
0.388
Target New
0.200
0.201
0.571
0.331
0.326
0.388
Appendix
Table 6: Pre-edit performance of LLaMA-3.1-8B-Instruct before editing.
Role
Content
User Query ( u )
What’s the most reliable medical treatment for premature graying?
Chosen ( y+ )
At present, there is no scientifically proven, consistently reliable medical treatment for premature graying, as the causes are often unknown and multifactorial. Graying hair is a natural part of the aging process when the pigment-producing cells in hair follicles slowly reduce their production of melanin, which gives hair its natural color. However, in some cases, premature hair graying can be a result of genetics, stress, certain medical conditions, or other lifestyle factors. If you are concerned about premature graying, it’s essential to consult with a healthcare professional, who can help determine the underlying causes and potential treatment options based on your individual case and your specific needs. They may recommend addressing any imbalances in your healthcare plan, improving overall well-being, and monitoring any related medical conditions.
Rejected ( y− )
There is currently no definitive medical treatment to stop or reverse premature graying. The graying of hair is primarily influenced by genetics. However, if there is an underlying medical condition causing premature graying, treating that condition may help slow down the process. It’s always best to consult with a healthcare provider or a dermatologist for personalized advice. In some cases, premature graying can be a sign of vitamin B12 deficiency, so addressing this deficiency might help. Certain hair care products and hair dyes can also help to mask the appearance of gray hair. It’s important to note that while there are many supplements and home remedies advertised to prevent or reverse graying, there is little scientific evidence to support these claims.
Appendix
Table 7: Example of preference pairs from the OpenHermesPreferences dataset.
Phase
Template Type
Input Prompt ( x )
Target Output ( y )
Editing
Context-free cloze
jazz was created in the country of
Indonesia
Evaluation
QA prompt with official chat template
Question: Which country was jazz created in? Answer:
Indonesia
Appendix
Table 8: Example prompt templates used in the editing and evaluation phases.
L
=Ledit(θ;et)+Lpref(θ;u,y+,y−)
+λrep∣R∣1e∈R∑Ledit(θ;e)
Appendix
Algorithm 1 GLIME
You are a strict grader.
Given:
- Question
- Gold target
- Predicted answer
Return:
A if the predicted answer semantically matches the gold target.
Appendix
Table 9: LLM-as-Judge prompt template.
Reliability
Generalization
Locality
Method
TF
AD
LLM
TF
AD
LLM
TF
AD
LLM
AlphaEdit
0.912
0.604
0.594
0.497
0.514
0.500
0.324
0.209
0.173
GLIME
0.937
0.830
0.775
0.806
0.800
0.746
0.517
0.379
0.328
Appendix
Table 10: Comparative analysis across evaluation metrics, including LLM-as-a-Judge.
Reliability
Generalization
Locality
Portability
Default
0.830
0.800
0.379
0.633
Training Epoch
1
0.653
0.665
0.436
0.507
2
0.770
0.744
0.350
0.555
4
0.796
0.798
0.308
0.608
5
0.829
0.811
0.288
0.631
Appendix
Table 11: Sensitivity analysis results of GLIME under different hyperparameter settings.
Method
Reliability
Generalization
Locality
Portability MC
Portability MR
Sketch
0.830
0.800
0.379
0.633
0.184
SVD
0.871
0.838
0.394
0.682
0.181
Appendix
Table 12: Comparison between the randomized sketch and exact SVD for the gradient-space constraint (GC).
Method
Reliability
Generalization
Locality
Portability MC
Portability MR
AlphaEdit
0.604
0.514
0.209
0.434
0.079
+ Replay
0.530
0.545
0.166
0.320
0.085
Appendix
Table 13: Effect of applying memory replay to AlphaEdit.
Epoch
Ledit
Lpref
Lreplay
1
2.8943
0.6979
0.5638
2
1.1538
0.5074
0.5545
3
0.5844
0.4319
0.5439
Appendix
Table 14: Average magnitude of each training objective across epochs.
Method
Predicted Answers
[Reliability]
Prompt: "Frank R. Strayer is a citizen of" Target: Canada
Large language models (LLMs) require frequent knowledge updates to reflect changing facts and mitigate hallucinations. To meet this demand, lifelong knowledge editing has emerged as a continual approach to modify specific pieces of knowledge without retraining the entire model. Existing parameter editing methods struggle with stability during sequential edits due to catastrophic forgetting. While retrieval-based approaches are proposed to alleviate this issue, their applicability remains limited across various datasets because of high training costs. To address these limitations and enhance scalability in lifelong settings, we propose LightEdit. Our framework first selects relevant knowledge from retrieved information to modify the query effectively. It then incorporates a decoding strategy to suppress the model's original knowledge probabilities, thereby enabling efficient edits based on the selected information. Extensive experiments on ZSRE, Counterfact, and RIPE benchmarks demonstrate that LightEdit outperforms existing lifelong knowledge editing methods. Furthermore, by minimizing training costs, LightEdit achieves cost-effective scalability, enabling easy adaptation to various datasets.
Dahyun Jung, Jaewook Lee, Heuiseok Lim
Department of Computer Science and Engineering, Korea University
Lifelong knowledge editing aims to efficiently and sequentially update language models over time, as new knowledge becomes available or when the model makes mistakes, while preserving acceptable performance on past knowledge. One unresolved challenge is that existing methods modify a fixed set of layers for all new knowledge samples, reducing flexibility and increasing catastrophic forgetting. Another is requiring access to previous knowledge and extensive pre-processing to obtain data statistics. To address these challenges, we introduce LOKI, a novel approach that uses dynamic layer selection based on the Hilbert-Schmidt Independence Criterion and projects gradient updates onto the null-space of the model weights, bypassing the requirement for previous knowledge access. We show that LOKI achieves superior performance to existing approaches across a wide variety of experiments, achieving up to a 14% improvement in average accuracy.
Masih Eskandar, Miquel Sirera Perelló, Stratis Ioannidis +1
Department of Electrical and Computer Engineering Northeastern University Boston, MA, 02115
Large language models (LLMs) often produce incorrect or outdated content after being employed. Efficient and accurate knowledge updates without costly retraining are a major challenge. This problem is particularly challenging in lifelong settings, where complex, unstructured knowledge must coexist without interference. We introduce RILKE (Representation Intervention for Lifelong KnowledgE Control), a robust and scalable method that treats knowledge control as interventions within the model's representation space. Leveraging representation-space expressiveness, we identify two key properties enabling RILKE to achieve fine-grained control over complex, unstructured knowledge while maintaining general utility with frozen base weights. During training, RILKE learns paraphrase-robust and edit-localized modules that limit each update to a low-dimensional subspace to minimize cross-edit interference. At inference, a query-adaptive router selects the appropriate module to guide the model's generation. Across LLaMA and Qwen models, RILKE scales effectively to large-scale benchmarks, demonstrating high edit success and strong paraphrase generalization while preserving general utility with modest memory overhead. These results show RILKE is an effective and scalable solution for lifelong knowledge control in LLMs.
Xuyuan Liu, Shengyu Chen, Xinshuai Dong +6
Dartmouth College · NEC Laboratories America · Carnegie Mellon University