Lifelong editing of LLMs requires storing thousands of edits after acquisition. A widely used family of approaches attaches one LoRA adapter per edit, which preserves behavior but grows linearly in storage. To address this challenge, we propose LadderEdit, a method that compresses each LoRA adapter after it is acquired. Each edit is first stored at low rank as a cheap sketch. We then check whether this sketch still satisfies the rewrite, generalization, and locality contract on probe prompts. Edits that pass keep the sketch; those that fail are promoted to a higher rank along a ladder until the contract is met. Because every edit retains some representation, coverage is maintained, and only hard edits consume more rank. Across ZsRE, CounterFact, and WikiBigEdit benchmarks on LLaMA-3-8B, Mistral-7B, and Qwen2.5-7B, LadderEdit tracks exact LoRA storage at 5.2x less memory and remains effective at 50,000 sequential edits.
Figures & tables
Figure 1: Performance comparison of LadderEdit, exact LoRA, and baseline methods across sequential editing tasks. LadderEdit usually tracks exact LoRA closely while reducing persistent memory, with small losses in some compressed settings.
Figure 2: LadderEdit insertion pipeline. We first fit an exact LoRA update for the incoming edit, then a candidate rank r^i from the update spectrum and exact-edit margin (one SVD, no behavioral cost). We further decompose the edit into ladder rungs: low-rank sketches at increasing rank. And finally we audit the predicted sketch against rewrite, generalization, and locality. If the audit passes, store at r^i ; otherwise promote along the spectral ladder until a passing rung is found.
Method
T=100
T=500
T=1,000
T=2,000
Rel.
Gen.
Loc.
Avg.
Rel.
Gen.
Loc.
Avg.
Rel.
Gen.
Loc.
Avg.
Rel.
Gen.
Loc.
Avg.
LLaMA-3-8B
FT
0.13 ± .004
0.11 ± .005
0.02 ± .001
0.09 ± .002
0.14 ± .005
0.12 ± .006
0.02 ± .001
0.09 ± .003
0.13 ± .005
0.12 ± .007
0.01 ± .001
0.09 ± .003
0.09 ± .006
0.08 ± .007
0.01 ± .001
0.06 ± .003
ROME Meng et al. (2022)
0.08 ± .003
0.08 ± .003
0.02 ± .001
0.06 ± .001
0.04 ± .002
0.04 ± .002
0.02 ± .001
0.03 ± .001
0.03 ± .002
0.03 ± .002
0.02 ± .001
0.03 ± .001
0.01 ± .001
0.01 ± .001
0.01 ± .001
0.01 ± .001
MEMIT Meng et al. (2023)
0.03 ± .002
0.03 ± .002
0.01 ± .001
0.02 ± .001
0.01 ± .001
0.01 ± .001
0.00 ± .000
0.01 ± .001
0.00 ± .000
0.00 ± .000
0.00 ± .000
0.00 ± .001
0.00 ± .000
0.00 ± .000
0.00 ± .000
0.00 ± .001
GRACE Hartvigsen et al. (2023)
1.00 ± .001
0.39 ± .009
1.00 ± .001
0.80 ± .003
1.00 ± .001
0.38 ± .011
1.00 ± .001
0.79 ± .004
1.00 ± .001
0.37 ± .012
1.00 ± .001
0.79 ± .004
1.00 ± .001
0.34 ± .013
1.00 ± .001
0.78 ± .004
Table 1: Q&A task results on the ZsRE dataset. T denotes the number of edits. Mean ± standard deviation over three random seeds. Avg. is the arithmetic mean of Rel., Gen., and Loc. The best average performance per column is marked in bold .
Method
LLaMA-3-8B
Mistral-7B
T=100
T=500
T=1,000
T=2,000
T=100
T=500
T=1,000
T=2,000
Rel.
Loc.
Rel.
Loc.
Rel.
Loc.
Rel.
Loc.
Rel.
Loc.
Rel.
Loc.
Rel.
Loc.
Rel.
Loc.
FT
0.42 ± .004
0.55 ± .001
0.25 ± .005
0.30 ± .001
0.12 ± .005
0.12 ± .001
0.05 ± .002
0.03 ± .002
0.39 ± .005
0.52 ± .001
0.22 ± .006
0.27 ± .001
0.10 ± .006
0.10 ± .002
0.03 ± .002
0.02 ± .001
ROME Meng et al. (2022)
0.28 ± .003
0.45 ± .002
0.08 ± .004
0.15 ± .002
0.02 ± .001
0.05 ± .002
0.00 ± .000
0.02 ± .001
0.25 ± .003
0.42 ± .002
0.06 ± .004
0.12 ± .003
0.01 ± .001
0.04 ± .002
0.00 ± .000
0.01 ± .001
MEMIT Meng et al. (2023)
0.35 ± .002
0.52 ± .001
0.12 ± .002
0.20 ± .001
0.03 ± .002
0.08 ± .001
0.01 ± .001
0.03 ± .002
0.31 ± .002
0.48 ± .001
0.09 ± .003
0.16 ± .001
0.02 ± .001
0.06 ± .002
0.00 ± .000
0.02 ± .001
GRACE Hartvigsen et al. (2023)
0.75 ± .001
0.96 ± .001
0.71 ± .001
0.96 ± .001
0.68 ± .001
0.96 ± .001
0.62 ± .001
0.95 ± .001
0.72 ± .001
0.95 ± .001
0.68 ± .001
0.95 ± .001
0.65 ± .001
0.95 ± .001
0.59 ± .001
0.94 ± .001
Table 2: Sequential editing results on CF. T denotes the number of edits. Mean ± standard deviation over three random seeds. The best performance per column is marked in bold .
Method
T=10,000
T=50,000
Rel.
Gen.
Loc.
Avg.
Rel.
Gen.
Loc.
Avg.
MEMIT Meng et al. (2023)
0.00 ± .000
0.00 ± .000
0.00 ± .000
0.00 ± .001
N/A
N/A
N/A
N/A
GRACE Hartvigsen et al. (2023)
1.00 ± .001
0.31 ± .017
1.00 ± .001
0.77 ± .006
0.99 ± .002
0.28 ± .019
1.00 ± .001
0.76 ± .006
WISE Wang et al. (2024a)
0.59 ± .027
0.57 ± .029
1.00 ± .001
0.72 ± .013
0.43 ± .030
0.41 ± .032
1.00 ± .001
0.61 ± .015
AlphaEdit Fang et al. (2025)
0.72 ± .023
0.77 ± .025
0.74 ± .015
0.74 ± .012
0.62 ± .026
0.60 ± .028
0.71 ± .017
0.64 ± .014
MEMOIR Wang et al. (2025)
0.72 ± .015
0.75 ± .019
0.83 ± .012
0.77 ± .009
0.66 ± .017
0.68 ± .022
0.73 ± .013
0.69 ± .010
Table 3: Editing performance on Qwen2.5-7B at long-horizon scale on WikiBigEdit. Mean ± standard deviation over three random seeds.
Figure 3: Scaling reference on WikiBigEdit: average performance against persistent edit memory.
Diagnostic case
ZsRE
CF
Sketch-sufficient (rank-1 passes)
81.2%
76.4%
Detail-needed (rank ≥2 required)
13.4%
16.8%
Residual-needed (rank Rmax required)
3.1%
4.7%
Compression-regularized (sketch > exact)
4.3%
6.7%
Table 4: Population structure of edit streams at N=2,000 . Categories are determined by the behavioral audit on audit-split probes. Most edits are over-resolved by the full LoRA update.
Dataset
Policy
Memory
Comp.
Utility
CF
exact-cache (matched)
0.180
5.57
0.434
CF
rank-1 sketch only
0.125
8.00
0.753
CF
static mixed-rank
0.180
5.57
0.815
CF
dual-budget LadderEdit
0.180
5.57
0.862
ZsRE
exact-cache (matched)
0.178
5.63
0.420
ZsRE
rank-1 sketch only
0.124
8.09
0.765
Table 5: Budget-matched strict-memory frontier at N=512 . LadderEdit preserves coverage and allocates residual fidelity through behavioral rung selection.
Figure 4: Spectral-filter frontier analysis on ZsRE and CF at N=512 edits. The plots show the tradeoff between memory compression and behavioral contract utility for different rank tiers and allocation strategies.
Method
Time
Peak Mem.
Persistent
Serving path
(ms)
(GB)
(GB)
LoRA (Exact) Hu et al. (2022)
86
18.4
100.93
exact-adapter lookup
MELO Yu et al. (2023)
91
18.6
100.93
retrieval
ELDER Li et al. (2024a)
95
19.1
0.08
MoE routing
MEMOIR Wang et al. (2025)
104
17.9
0.14
mask gate
LadderEdit
124
18.8
19.44
ladder lookup
Table 6: Cost decomposition at T=10,000 on LLaMA-3-8B. LadderEdit reduces this storage by 5.19× relative to exact LoRA. Serving path summarizes the additional edit mechanism used at inference.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Rel.
Gen.
Loc.
Avg.
Δ vs. matched
Matched association (Exact LoRA)
0.95
0.91
1.00
0.95
—
Matched association (LadderEdit)
0.95
0.91
1.00
0.95
—
Contriever retriever (Exact LoRA)
0.91
0.86
0.99
0.92
−0.03
Contriever retriever (LadderEdit)
0.90
0.86
0.99
0.92
−0.03
E5-large retriever (LadderEdit)
0.88
0.85
0.99
0.91
−0.04
MPNet (frozen) (LadderEdit)
0.85
0.81
0.98
0.88
−0.07
Appendix
Table 7: Retrieval sensitivity on ZsRE / LLaMA-3-8B at T=1,000 . The persistent-memory ratio is preserved under all retrieval settings in the table because the retriever is shared between exact LoRA and LadderEdit. Locality is barely affected (0.99 → 0.98 under MPNet) because retrieved adapters still land on real edit directions rather than random directions.
Policy
Memory (norm.)
Pooled Avg.
Worst-edit Avg.
Fixed rank-1
0.125
0.85 (0.008)
—
Fixed rank-2
0.250
0.89 (0.007)
—
Fixed rank-4
0.500
0.90 (0.008)
—
Fixed rank-8
1.000
0.90 (0.007)
0.62
Uniform random rank
0.562
0.87 (0.010)
—
LadderEdit (audited)
0.19
0.95 (0.005)
0.74
Appendix
Table 8: Audit contribution against fixed-rank storage at T=2,000 on ZsRE / LLaMA-3-8B across three seeds. Memory is normalized to exact LoRA. Pooled Avg. is the pooled evaluation Avg. across edits; worst-edit Avg. is the per-edit Avg. averaged over only the bottom 10% of edits. Standard deviations over three seeds are reported in parentheses when available.
Figure 5: Sensitivity of LadderEdit to the locality weight wL on ZsRE / LLaMA-3-8B at T=1,000 , with wR and wG held fixed. Left: average performance, peaking at the default wL=1.2 with a stable region across [1.0,1.5] . Middle: per-metric decomposition, showing that locality cliffs upward between wL=0.8 and wL=1.0 while reliability and generalization degrade for wL≥1.5 . Right: failure-mode decomposition, with locality failures dominating at low wL and reliability or generalization rejections dominating at high wL ; the default minimizes the total failure rate.
Policy
Memory (norm.)
Avg.
Exact LoRA (FP16)
1.00
0.94
Exact LoRA (INT8)
0.25
0.93
LadderEdit (FP16)
0.19
0.91
LadderEdit (INT8)
0.05
0.90
Appendix
Table 9: Composition of LadderEdit with INT8 quantization at T=2,000 on ZsRE / LLaMA-3-8B. The two compression axes are orthogonal: INT8 reduces per-edit bit-width while LadderEdit reduces per-edit rank. Composing them yields a 20× reduction over FP16 exact LoRA. The composed configuration (LadderEdit + INT8) reaches 0.90 Avg. at 0.05 normalized memory, dominating exact LoRA + INT8 alone (0.93 Avg. at 0.25 memory) on the cost axis.
Stored representation
Fraction of edits
Rank-1 sketch
79.4%
Rank-2 sketch
16.1%
Rank-3 sketch
1.8%
Rank-4 sketch
1.2%
Rank-5 sketch
0.7%
Rank-6 sketch
0.4%
Appendix
Table 10: Distribution of stored representations (N=1,000 edits at T=1,000 on LLaMA-3-8B / ZsRE; sums to 100.0%) across the full r=1 to r=8 ladder. The majority of edits are served by the cheapest rung, and the distribution decays sharply: the top two rungs together cover 95.5% of edits, while only the remaining ∼4.5% requires ranks 3 or above. The audit promotes edits to higher rungs only when the low-rank sketch fails to satisfy the behavioral contract, which corresponds approximately to the detail-needed population of Table 4 . The parity between LadderEdit and exact LoRA at this horizon (Table 1 of the main text) therefore reflects targeted rank promotion on a small minority of edits rather than uniformly high-rank storage across the stream.
Dataset
Audit probes/edit
Val. probes/edit
Test probes/edit
ZsRE
2
1
2
CounterFact
3
1
3
WikiBigEdit
2
1
2
Appendix
Table 11: Per-edit probe counts after the three-way split. Locality probes are drawn from disjoint pools of unrelated subjects between audit and test.
Configuration
Rel.
Gen.
Loc.
Avg.
LadderEdit, strict split (main)
0.95
0.91
1.00
0.95
LadderEdit, leaky (audit=test)
0.99
0.97
1.00
0.99
Inflation gap
+0.04
+0.06
0.00
+0.04
Appendix
Table 12: Leaky-versus-clean comparison at T=1,000 on ZsRE / LLaMA-3-8B. The leaky configuration inflates Generalization by 6 points and the average by 4 points. The numbers reported in the main paper correspond to the strict split.
Dataset
Case
Edit id
Energy
ΔU
Gap
Interpretation
CF
Sketch sufficient
0
0.517
0.000
0.000
Rank-1 preserves the exact edit’s measured behavior, so higher resolution is unnecessary.
CF
Detail needed
9
0.590
-0.333
0.267
The sketch drops needed behavior; the audit should promote to a higher rung.
CF
Compression helps
14
0.474
0.333
0.000
Removing tail directions improves measured behavior, so exact LoRA is not always a behavioral ceiling.
CF
Exact limited
4
0.524
0.000
0.267
Both exact and sketch leave a contract gap; this is an edit-writer limitation, not a compression error.
ZsRE
Sketch sufficient
0
0.496
0.000
0.000
The cheap sketch is behaviorally indistinguishable from exact on the probes.
ZsRE
Detail needed
3
0.576
-0.167
0.100
Residual detail is needed for a hard edit that rank-1 cannot safely represent.
Appendix
Table 13: Representative edit-level diagnostic cases at N=512 , seed 0, rank 1. ΔU is the measured utility of the rank-1 sketch minus exact LoRA. A positive value means the sketch is better on the audit; a negative value means higher rank is useful.
Audit threshold τ with r=8
LoRA rank r with τ=0.05
τ=0.01
τ=0.025
τ=0.05
τ=0.10
r=4
r=8
r=16
r=32
Rel. ( ↑ )
0.73
0.86
0.95
0.81
0.78
0.95
0.88
0.74
Gen. ( ↑ )
0.82
0.88
0.91
0.85
0.81
0.91
0.89
0.83
Loc. ( ↑ )
0.99
1.00
1.00
0.94
0.99
1.00
0.97
0.92
Avg. ( ↑ )
0.85
0.91
0.95
0.87
0.86
0.95
0.91
0.83
Memory (GB)
4.21
3.97
3.84
3.62
2.18
3.84
7.32
14.18
Appendix
Table 14: Sensitivity of LadderEdit to the audit threshold τ and LoRA rank r on ZsRE / LLaMA-3-8B at T=1,000 . Avg. is the arithmetic mean of Rel., Gen., and Loc. Bold marks the default configuration. The default (τ=0.05,r=8) sits at the peak of both sweeps: smaller τ or r underfits the audit, larger τ or r admits noisy edits that erode locality.
Probe bank size ∣P∣
Coverage
Probed edits
Unprobed edits
Overall
Avg. ( ↑ )
Failures (%)
Avg. ( ↑ )
Failures (%)
Avg. ( ↑ )
256
38%
0.95
1.2
0.78
14.6
0.84
1,024
64%
0.95
0.9
0.85
8.3
0.91
4,096
87%
0.95
0.8
0.89
4.1
0.94
16,384
96%
0.95
0.7
0.91
2.4
0.94
65,536
99%
0.95
0.7
0.92
1.8
0.95
Appendix
Table 15: Probe-coverage stress test on ZsRE / LLaMA-3-8B at T=2,000 . We vary the audit probe-bank size ∣P∣ and report performance on probe-covered and unprobed edits. Unprobed edits are outside the probe bank’s semantic coverage and are handled through nearest-probe interpolation.
Case
Prompt and target
Exact LoRA
Sketch rung
Ladder decision
ZsRE 29: sketch sufficient
Rewrite: Which is the position of Jules Basile Onambele? Target: winger. Generalization: What is Jules Basile Onambele’s position?
Jules Basile Onambele primarily plays as a winger or attacking midfielder.
Rank 1 gives the same rewrite and generalization output as exact LoRA.
Store rank 1. The full residual is unnecessary.
CF 11: sketch sufficient
Rewrite: Andreas Ivanschitz professionally plays the sport. Target: football. Generalization: After work Walther attended the Y.M.C.A. Andreas Ivanschitz, the
Andreas Ivanschitz is known for playing football (soccer) professionally. The generalization output remains incomplete.
Rank 1 preserves the same rewrite behavior and the same incomplete generalization response.
Store rank 1. A hard dataset still contains cheap compressible edits.
ZsRE 27: detail needed
Rewrite: What sports team was Stanislav Romanov a member of? Target: Spartak Myjava. Generalization: Which sports team was Stanislav Romanov’s?
Stanislav Romanov was a member of FC Spartak Myjava.
Rank 1 changes the value to FC Spartak Mytishchi, while rank 2 recovers FC Spartak Myjava.
Promote to rank 2. The edit needs a little more detail, but not exact storage.
ZsRE 24: compression helps
Rewrite: What kind of occupation does Karim-Mohamed Maamoun have? Target: architect. Generalization: What kind of occupation is Karim-Mohamed Maamoun?
Exact LoRA answers archaeologist on the rewrite and generalization prompts.
Rank 1 and rank 2 include archaeologist and architect in the rewrite output.
Store the sketch. Low-rank projection can filter harmful detail from the exact edit.
Appendix
Table 16: Prompt-level qualitative examples illustrating the main residual-ladder decisions. Outputs are shortened only for formatting. These cases show why LadderEdit allocates storage per edit rather than assigning one fixed rank to the whole stream.
Case
Exact LoRA
Sketches
Interpretation
CF 27, target Russian. Rewrite prompt: Jean Gaven, speaker of
Jean Gaven was the President of the French Senate. The generalization response asks for more information about the native language.
Rank 1 and rank 2 produce the same incorrect or uninformative behavior.
Exact-limited. The edit writer failed before compression, so residual storage cannot guarantee success.
Appendix
Table 17: Caveat example: exact acquisition can fail under free generation. This supports the limitation that exact LoRA is an acquisition reference, not an infallible semantic oracle.
Figure 6: Additional audit diagnostics. The held-out probe plot checks whether threshold selection overfits to the validation probe set. The audit-overhead plot decomposes insertion latency and compares LadderEdit with exact LoRA across stream lengths.
Quantity
Value
Crossover horizon
T≈2,000
Exact LoRA @ T=100
86 ms/edit
LadderEdit @ T=100
93 ms/edit
Exact LoRA @ T=50,000
850 ms/edit
LadderEdit @ T=50,000
150 ms/edit
Speedup @ T=50,000
5.7×
Appendix
Table 18: Insertion latency summary (A100-40GB).
Predictor
Utility
Audits/edit
Memory
Random rank proposal
0.81
2.31
0.22
Always rank-1, then promote
0.89
1.85
0.20
Spectral tail Ti(r) only
0.88
1.32
0.19
Margin μiE only
0.86
1.42
0.21
Spectral × margin (ours)
0.91
1.21
0.19
Oracle (exhaustive audit)
0.93
8.00
0.19
Appendix
Table 19: Predictor ablation at T=1,000 on ZsRE / LLaMA-3-8B. Utility is the audited behavioral contract score; Audits/edit measures the cost of the proposal at insertion time; Memory is normalized to exact LoRA. Bold marks the configuration used in the main paper. The combined spectral-times-margin predictor reaches utility within 0.02 of oracle at 4.6× fewer audits per edit, while each factor alone trails by 0.03–0.05. The always-rank-1 baseline shows that the audit-and-promote loop alone is competitive on utility but pays a higher audit cost without a predictor to localise the proposal.
Dataset
N
Rank
Memory
Energy
Contract
False Ret.
CF
32
1
0.125
0.560
0.625
1.67
CF
32
2
0.250
0.747
0.656
0.00
CF
32
4
0.500
0.899
0.646
0.00
CF
32
8
1.000
1.000
0.646
0.00
CF
128
1
0.125
0.559
0.641
7.00
CF
128
2
0.250
0.744
0.662
1.67
Appendix
Table 20: Fixed-rank tier diagnostics for CF and ZsRE at N=32 , 128 , 512 . Memory is normalized to exact LoRA. Retained energy is 1−∥ρ∥F2/∥ΔE∥F2 . False retirements count exact-pass edits that fail after compression.
Figure 7: Swap-count diagnostics for ZsRE and CF at N=512 . The plots show how many edits change pass/fail status as rank increases, highlighting the difficult subset that requires higher rank.
Backbone
Method
Acc.
Gen.
Loc.
Avg.
Qwen2.5-14B
Exact LoRA
0.795
0.905
0.999
0.900
LadderEdit
0.790
0.900
0.999
0.896
Qwen2.5-32B
Exact LoRA
0.821
0.923
0.999
0.914
LadderEdit
0.816
0.918
0.999
0.911
Appendix
Table 21: LadderEdit on larger Qwen2.5 backbones. Accuracy (Acc.), generalization (Gen.), locality (Loc.), and their average are reported. LadderEdit remains close to Exact LoRA as the backbone increases from 14B to 32B parameters.
Cache policy
Full-stream utility
Random retention
0.18
FIFO
0.20
LRU
0.22
Learned retention
0.35
Learned retention + base fallback
0.41
Appendix
Table 22: Full-stream utility of exact-cache policies under the same capacity of 91 retained adapters out of 512 edits. Learned retention with base-model fallback is the strongest exact-cache baseline.
Dataset
Policy
Norm. memory
Utility
CF
Exact-cache
0.180
0.434
CF
Rank-1 for all edits
0.125
0.753
CF
Static mixed-rank
0.180
0.815
CF
LadderEdit
0.180
0.862
ZsRE
Exact-cache
0.178
0.420
ZsRE
Rank-1 for all edits
0.124
0.765
Appendix
Table 23: Strict-memory comparison on CounterFact (CF) and ZsRE. Utility is measured under approximately matched normalized persistent memory.
Method
Insert / edit
Peak GPU
Persistent memory
Exact LoRA
86 ms
18.4 GB
100.93 GB
LadderEdit
124 ms
18.8 GB
19.44 GB
Difference
+38 ms
+0.4 GB
−81.49 GB
Appendix
Table 24: Measured insertion and memory cost at T=10,000 edits. Persistent memory refers to the retained edit bank, whereas peak GPU memory measures transient working memory during insertion.
Diagnostic property
ZsRE
CounterFact
Rank-1 sufficient
81.2%
76.4%
Rank 2 or higher required
13.4%
16.8%
Full acquisition rank required
3.1%
4.7%
Sketch passes while exact fails
4.3%
6.7%
Appendix
Table 25: Diagnostic rates describing edit-level rank requirements and compression behavior. The rows are diagnostic properties and are not assumed to form a mutually exclusive partition.
Selection mechanism
Exact LoRA Avg.
LadderEdit Avg.
Matched association
0.95
0.95
Contriever ( Izacard et al., 2021 )
0.92
0.92
E5-large ( Wang et al., 2022 )
–
0.91
Frozen MPNet ( Song et al., 2020 )
–
0.88
Appendix
Table 26: Performance under matched association and learned retrieval. Contriever is evaluated directly for both Exact LoRA and LadderEdit. Additional LadderEdit results with E5-large and frozen MPNet characterize sensitivity to the retriever.
Rank proposal
Utility
Audits / edit
Random
0.81
2.31
Always rank-1, then promote
0.89
1.85
Spectral tail only
0.88
1.32
Behavioral margin only
0.86
1.42
Spectral tail + margin
0.91
1.21
Exhaustive oracle
0.93
8.00
Appendix
Table 27: Rank-proposal ablation. The combined spectral-tail and behavioral-margin proposal approaches exhaustive rank search while requiring substantially fewer audits.
Probe coverage
Overall Avg.
38%
0.84
64%
0.91
87%
0.94
96%
0.94
99%
0.95
Appendix
Table 28: Sensitivity to semantic coverage of the behavioral audit probes.
Large language models encode vast factual knowledge that can become outdated or incorrect after deployment, yet retraining is prohibitively costly. This motivates lifelong model editing, which updates targeted behavior while preserving the rest of the model. Existing editors, both parameter-modifying and parameter-preserving, degrade severely as edits accumulate and struggle to generalize across paraphrases. We propose HoReN, a codebook-based parameter-preserving editor that wraps a single MLP layer with a discrete key-value memory. HoReN treats each codebook entry as both a knowledge key and a Hopfield stored pattern, retrieves edits by angular similarity on the unit hypersphere, and refines queries through damped Hopfield dynamics so paraphrases converge to the correct memory basin while unrelated inputs remain stable. HoReN achieves strong editing performance with consistent gains across diverse benchmarks spanning standard ZsRE, structured WikiBigEdit, and unstructured UnKE evaluations. Moreover, HoReN scales to 50K sequential edits on ZsRE with stable overall performance above 0.93, while prior editors collapse or degrade severely before reaching 10K. Our code is available at https://github.com/ha11ucin8/HoReN.
Yuan Fang, Yi Xie, Xuming Ran
1IXL Learning, Inc · 2Technical University of Munich · 3National University of Singapore
Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forgetting and degradation of general capabilities. To address these challenges, we reframe knowledge editing from a manifold perspective, viewing it as a localized displacement of an edit sub-manifold within the global knowledge manifold. Under this formulation, the problem can be decomposed into two key questions: (i) how to identify representative edit points that effectively anchor the edit sub-manifold, and (ii) how to preserve the remaining manifold structure during the sub-manifold displacement process. Based on this perspective, we propose ManiEdit, a novel manifold-aware autoregressive editing framework consisting of two core components. Pivot Localization addresses the mediocre-point dilemma by identifying high-leverage pivots to anchor the edit sub-manifold. Manifold-Aware Preservation preserves different knowledge types through an energy-weighted penalty combined with recursive null-space alignment. Experiments on two base LLMs and four unstructured editing benchmarks demonstrate that ManiEdit achieves state-of-the-art performance, outperforming the strongest baseline by up to +27.81 BERTScore and +8.50 ROUGE-L, while maintaining near-original general capabilities across six representative downstream tasks. Our code is available at: https://github.com/Areyliu/ManiEdit
Rui Liu, Chenheng Zhang, Haoxuan Li +1
State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Institute for Artificial Intelligence, Peking University
Null-space-based methods have garnered considerable attention in model editing by constraining updates to the null space of the pre-existing knowledge representation, thereby preserving the model's original behavior. However, in practice these methods rely on an approximate null space--leading to knowledge leakage--and further suffer from severe performance degradation during sequential editing. Recent work shows that history-aware editing strategies can empirically mitigate this decline, yet the underlying reason remains unclear. In this paper, we first expose the knowledge leakage inherent in existing null-space approaches and then analyze why history-aware updates effectively preserve both editing performance and general capabilities during long-horizon editing. Building on these insights, we propose BetaEdit, a refined framework that effectively controls the knowledge leakage and integrates history-aware updates into the null-space paradigm. Extensive experiments on three large language models across two standard benchmarks show that BetaEdit consistently outperforms prior methods in the challenging regime of massive-scale sequential editing. Code is available at: https://github.com/lbq8942/BetaEdit.
Bingqing Liu, Wei Liu, Yuhua Li
Huazhong University of Science and Technology, Wuhan, China