ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective
Authors: Rui Liu, Chenheng Zhang, Haoxuan Li, Zhouchen Lin
Organizations: State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Institute for Artificial Intelligence, Peking University
Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forgetting and degradation of general capabilities. To address these challenges, we reframe knowledge editing from a manifold perspective, viewing it as a localized displacement of an edit sub-manifold within the global knowledge manifold. Under this formulation, the problem can be decomposed into two key questions: (i) how to identify representative edit points that effectively anchor the edit sub-manifold, and (ii) how to preserve the remaining manifold structure during the sub-manifold displacement process. Based on this perspective, we propose ManiEdit, a novel manifold-aware autoregressive editing framework consisting of two core components. Pivot Localization addresses the mediocre-point dilemma by identifying high-leverage pivots to anchor the edit sub-manifold. Manifold-Aware Preservation preserves different knowledge types through an energy-weighted penalty combined with recursive null-space alignment. Experiments on two base LLMs and four unstructured editing benchmarks demonstrate that ManiEdit achieves state-of-the-art performance, outperforming the strongest baseline by up to +27.81 BERTScore and +8.50 ROUGE-L, while maintaining near-original general capabilities across six representative downstream tasks. Our code is available at: https://github.com/Areyliu/ManiEdit
Figures & tables
Figure 1: Motivation for ManiEdit. Existing methods struggle with sequential unstructured knowledge editing, suffering from edit forgetting and general-capability collapse (left). In contrast, ManiEdit enables effective and reliable knowledge editing, preserving both edit knowledge and pre-trained general capabilities (right).
Figure 2: A manifold perspective of model editing. Left : The input representation manifold Min . Retain keys (blue) distribute along the principal axes {ψj} , forming a global sub-manifold Minretain ; edit-related keys (gold) form a local sub-manifold Minedit . Right : The output representation manifold Mout=W(Min) . The edit operator E displaces the edit sub-manifold Moutedit (gold) to its target Mout∗,edit (green), while the retain sub-manifold Moutretain (blue) remains unchanged.
Figure 3: Overview of the ManiEdit paradigm. Pivot Localization identifies high-leverage semantic pivots to guide the segmentation of input samples, then the Manifold-Aware Preservation mechanism applies energy-weighted penalty and recursive null-space alignment to preserve general capabilities and historical edits.
Method
UnKEBench
AKEW (Counterfact)
AKEW (MQUAKE)
Ori
Para
Ori
Para
Ori
Para
BERTScore
ROUGE-L
BERTScore
ROUGE-L
BERTScore
ROUGE-L
BERTScore
ROUGE-L
BERTScore
ROUGE-L
BERTScore
ROUGE-L
Based on Qwen2.5-7B-Instruct
MEMIT
8.36
16.89
8.95
16.73
9.07
19.23
9.43
19.22
10.47
18.71
9.17
18.67
RECT
12.28
17.57
12.35
17.07
6.21
17.08
5.90
16.76
13.34
21.90
13.20
21.37
EvoEdit
25.79
16.08
21.60
15.16
35.89
24.56
19.85
20.95
52.28
26.35
44.51
21.76
Table 1: Comparison of ManiEdit with existing methods on the sequential unstructured model editing tasks. The best results are highlighted in bold, while the second-best results are underlined.
Figure 4: General capability accuracy on three GLUE tasks (SST-2, CoLA, MRPC) as the number of sequential unstructured edits increases. ManiEdit maintains near-original accuracy across all tasks, while baselines collapse rapidly.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Ori
Para
Sub
BERTScore
ROUGE-L
ROUGE-1
BERTScore
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
Llama3-8B-Instruct
SERAC
70.15
27.89
29.73
70.44
26.91
28.88
22.03
22.71
MELO
69.58
25.19
27.13
70.04
25.26
27.26
21.98
22.66
GRACE
73.02
34.34
36.17
70.11
25.42
27.43
22.41
23.09
ManiEdit
76.22
37.30
39.66
75.80
35.61
38.11
32.26
33.28
Appendix
Table 2: Comparison of ManiEdit with memory-based methods under the sequential unstructured editing setting on UnKEBench.
Method
Ori
Para
Sub
BERTScore
ROUGE-L
ROUGE-1
BERTScore
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
Llama3-8B-Instruct
SERAC
69.22
23.20
24.88
42.97
15.25
16.58
37.36
38.35
MELO
69.69
20.30
22.22
52.09
19.11
21.04
34.42
35.42
GRACE
69.05
19.60
21.25
41.90
12.97
14.18
33.92
34.98
ManiEdit
78.50
37.24
39.98
53.88
28.73
30.80
40.05
41.18
Appendix
Table 3: Comparison of ManiEdit with memory-based methods under the sequential unstructured editing setting on the AKEW (Counterfact) benchmark.
Method
Ori
Para
Sub
BERTScore
ROUGE-L
ROUGE-1
BERTScore
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
Llama3-8B-Instruct
SERAC
69.97
19.94
21.50
69.75
21.20
23.54
42.36
43.64
MELO
69.52
17.00
18.54
69.78
21.14
23.48
39.95
41.35
GRACE
81.82
52.57
53.37
69.82
21.13
23.45
39.97
41.38
ManiEdit
71.51
30.03
32.57
71.24
35.34
38.35
32.70
34.05
Appendix
Table 4: Comparison of ManiEdit with memory-based methods under the sequential unstructured editing setting on the AKEW (MQUAKE) benchmark.
Method
Qwen2.5-7B-Instruct
Llama3-8B-Instruct
UnKEBench
AKEW (MQUAKE)
AKEW (Counterfact)
UnKEBench
AKEW (MQUAKE)
AKEW (Counterfact)
ROUGE-1
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
ROUGE-L
ROUGE-1
ROUGE-L
MEMIT
16.54
16.35
23.06
22.82
26.57
25.80
1.39
1.35
10.33
10.00
13.09
12.80
RECT
15.52
15.31
26.10
25.58
24.39
24.10
15.00
14.63
17.52
17.38
16.19
16.13
UnKE
6.72
6.38
25.03
26.62
11.26
11.09
0.69
0.68
5.51
5.32
4.06
4.03
AnyEdit
15.63
15.29
21.63
21.14
21.82
21.37
15.84
15.62
21.01
20.43
13.22
12.98
Appendix
Table 5: Comparison of ManiEdit with existing methods on the sequential unstructured model editing tasks under the Sub-question split. The best results are highlighted in bold, while the second-best results are underlined.
Backbone
Benchmark
General-Capability Preservation
Edit-Knowledge Preservation
Original
Para
BERTScore
ROUGE-L
BERTScore
ROUGE-L
Qwen2.5-7B-Instruct
UnKEBench
Null-space Projection
Recursive Null-space Alignment
43.69
24.12
41.85
23.64
Energy-weighted Penalty
Regularization
46.73
29.14
46.20
28.46
Null-space Projection
Regularization
37.08
30.90
37.99
31.45
Energy-weighted Penalty
Recursive Null-space Alignment
76.63
39.43
74.97
38.07
AKEW (Counterfact)
Null-space Projection
Recursive Null-space Alignment
45.76
27.54
30.49
24.60
Appendix
Table 6: Ablation of Manifold-Aware Preservation components on UnKEBench, AKEW (Counterfact), and AKEW (MQUAKE). The left component in each pair protects general pre-trained capabilities, and the right component protects sequentially edited knowledge.
Setting
MOSE
ManiEdit
Ori
Para
Ori
Para
BERTScore
ROUGE-L
BERTScore
ROUGE-L
BERTScore
ROUGE-L
BERTScore
ROUGE-L
@500
62.38
31.76
61.11
30.95
76.22
37.30
75.80
35.61
@1,000
54.65
20.57
55.56
19.80
77.00
34.99
75.66
34.61
Δ
−7.73
−11.19
−5.55
−11.15
+0.78
−2.31
−0.14
−1.00
Appendix
Table 7: Long-horizon stress test on Llama3-8B-Instruct with UnKEBench (scores in %). Δ denotes the change from 500 to 1,000 sequential edits.
Update
leakt↓
cumt↓
Median
Mean
Max
Median
Without alignment ( Pt−1=I )
9.02
8.62
16.0
8.72
ManiEdit
0.305
0.343
2.09
0.297
Reduction (median)
29.6×
29.4×
Appendix
Table 8: leakt and cumt over 500 sequential edits, with and without recursive null-space alignment.
Sequential editing of structured knowledge in large language models allows targeted factual updates without retraining, yet existing methods often rely on complex regularization or constraint mechanisms whose necessity remains unclear. In this work, we systematically investigate the mechanisms underlying effective and stable sequential editing. Specifically, we first analyze the empirical success of AlphaEdit and establish, via a rigorous optimization analysis, the formal equivalence between one-time and sequential editing. Building on this insight, we generalize the equivalence to a broader class of editing objectives, demonstrating that stability emerges naturally from properly accounting for accumulated editing constraints, rather than from specialized regularization or null-space operations. We empirically confirm that many commonly used regularization strategies are unnecessary for reliable sequential updates. Furthermore, we extend our framework to handle conflicting edits, ensuring robust and consistent behavior under contradictory updates. Ultimately, our work provides Ariadne's thread through the labyrinth of sequential editing, charting a path toward simpler, more interpretable, and dependable knowledge updates. Our code is available at https://github.com/Wangzzzzzzzz/OTE-SE-Alignment.
Zheng Wang, Kaixuan Zhang, Wanfang Chen +2
Bosch Center for Artificial Intelligence (BCAI) · Bosch (China) Investment Ltd. · School of Statistics, East China Normal University.
Knowledge editing enables rapid updates of specific factual knowledge in large language models (LLMs) without full retraining. However, more realistic scenarios call for a lifelong framework that handles continual updates rather than one-off modifications. In such settings, existing editing methods often overfit to target prompts, significantly degrading both the generalization of the edited knowledge and the model's general capabilities. To address this issue, we propose GLIME (Generalizable Lifelong Model Editing), which combines knowledge editing with preference optimization over generation behavior. GLIME further incorporates replay-based editing and a gradient constraint to preserve previously edited knowledge. Experimental results show that GLIME significantly improves knowledge generalization in lifelong editing settings while maintaining both editing performance and general capabilities.
Dahyun Jung, Suhyune Son, Heuiseok Lim
Department of Computer Science and Engineering, Korea University
Null-space-based methods have garnered considerable attention in model editing by constraining updates to the null space of the pre-existing knowledge representation, thereby preserving the model's original behavior. However, in practice these methods rely on an approximate null space--leading to knowledge leakage--and further suffer from severe performance degradation during sequential editing. Recent work shows that history-aware editing strategies can empirically mitigate this decline, yet the underlying reason remains unclear. In this paper, we first expose the knowledge leakage inherent in existing null-space approaches and then analyze why history-aware updates effectively preserve both editing performance and general capabilities during long-horizon editing. Building on these insights, we propose BetaEdit, a refined framework that effectively controls the knowledge leakage and integrates history-aware updates into the null-space paradigm. Extensive experiments on three large language models across two standard benchmarks show that BetaEdit consistently outperforms prior methods in the challenging regime of massive-scale sequential editing. Code is available at: https://github.com/lbq8942/BetaEdit.
Bingqing Liu, Wei Liu, Yuhua Li
Huazhong University of Science and Technology, Wuhan, China