LadderEdit: Edit-Level Residual Compression for Memory-Efficient Lifelong Editing of LLMs
Organizations: Yale University · Washington University in St. Louis · Icahn School of Medicine at Mount Sinai · Arizona State University
Abstract
Lifelong editing of LLMs requires storing thousands of edits after acquisition. A widely used family of approaches attaches one LoRA adapter per edit, which preserves behavior but grows linearly in storage. To address this challenge, we propose LadderEdit, a method that compresses each LoRA adapter after it is acquired. Each edit is first stored at low rank as a cheap sketch. We then check whether this sketch still satisfies the rewrite, generalization, and locality contract on probe prompts. Edits that pass keep the sketch; those that fail are promoted to a higher rank along a ladder until the contract is met. Because every edit retains some representation, coverage is maintained, and only hard edits consume more rank. Across ZsRE, CounterFact, and WikiBigEdit benchmarks on LLaMA-3-8B, Mistral-7B, and Qwen2.5-7B, LadderEdit tracks exact LoRA storage at 5.2x less memory and remains effective at 50,000 sequential edits.
Figures & tables
| Method | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rel. | Gen. | Loc. | Avg. | Rel. | Gen. | Loc. | Avg. | Rel. | Gen. | Loc. | Avg. | Rel. | Gen. | Loc. | Avg. | |
| LLaMA-3-8B | ||||||||||||||||
| FT | 0.13 .004 | 0.11 .005 | 0.02 .001 | 0.09 .002 | 0.14 .005 | 0.12 .006 | 0.02 .001 | 0.09 .003 | 0.13 .005 | 0.12 .007 | 0.01 .001 | 0.09 .003 | 0.09 .006 | 0.08 .007 | 0.01 .001 | 0.06 .003 |
| ROME Meng et al. (2022) | 0.08 .003 | 0.08 .003 | 0.02 .001 | 0.06 .001 | 0.04 .002 | 0.04 .002 | 0.02 .001 | 0.03 .001 | 0.03 .002 | 0.03 .002 | 0.02 .001 | 0.03 .001 | 0.01 .001 | 0.01 .001 | 0.01 .001 | 0.01 .001 |
| MEMIT Meng et al. (2023) | 0.03 .002 | 0.03 .002 | 0.01 .001 | 0.02 .001 | 0.01 .001 | 0.01 .001 | 0.00 .000 | 0.01 .001 | 0.00 .000 | 0.00 .000 | 0.00 .000 | 0.00 .001 | 0.00 .000 | 0.00 .000 | 0.00 .000 | 0.00 .001 |
| GRACE Hartvigsen et al. (2023) | 1.00 .001 | 0.39 .009 | 1.00 .001 | 0.80 .003 | 1.00 .001 | 0.38 .011 | 1.00 .001 | 0.79 .004 | 1.00 .001 | 0.37 .012 | 1.00 .001 | 0.79 .004 | 1.00 .001 | 0.34 .013 | 1.00 .001 | 0.78 .004 |
| Method | LLaMA-3-8B | Mistral-7B | ||||||||||||||
| Rel. | Loc. | Rel. | Loc. | Rel. | Loc. | Rel. | Loc. | Rel. | Loc. | Rel. | Loc. | Rel. | Loc. | Rel. | Loc. | |
| FT | 0.42 .004 | 0.55 .001 | 0.25 .005 | 0.30 .001 | 0.12 .005 | 0.12 .001 | 0.05 .002 | 0.03 .002 | 0.39 .005 | 0.52 .001 | 0.22 .006 | 0.27 .001 | 0.10 .006 | 0.10 .002 | 0.03 .002 | 0.02 .001 |
| ROME Meng et al. (2022) | 0.28 .003 | 0.45 .002 | 0.08 .004 | 0.15 .002 | 0.02 .001 | 0.05 .002 | 0.00 .000 | 0.02 .001 | 0.25 .003 | 0.42 .002 | 0.06 .004 | 0.12 .003 | 0.01 .001 | 0.04 .002 | 0.00 .000 | 0.01 .001 |
| MEMIT Meng et al. (2023) | 0.35 .002 | 0.52 .001 | 0.12 .002 | 0.20 .001 | 0.03 .002 | 0.08 .001 | 0.01 .001 | 0.03 .002 | 0.31 .002 | 0.48 .001 | 0.09 .003 | 0.16 .001 | 0.02 .001 | 0.06 .002 | 0.00 .000 | 0.02 .001 |
| GRACE Hartvigsen et al. (2023) | 0.75 .001 | 0.96 .001 | 0.71 .001 | 0.96 .001 | 0.68 .001 | 0.96 .001 | 0.62 .001 | 0.95 .001 | 0.72 .001 | 0.95 .001 | 0.68 .001 | 0.95 .001 | 0.65 .001 | 0.95 .001 | 0.59 .001 | 0.94 .001 |
| Method | ||||||||
|---|---|---|---|---|---|---|---|---|
| Rel. | Gen. | Loc. | Avg. | Rel. | Gen. | Loc. | Avg. | |
| MEMIT Meng et al. (2023) | 0.00 .000 | 0.00 .000 | 0.00 .000 | 0.00 .001 | N/A | N/A | N/A | N/A |
| GRACE Hartvigsen et al. (2023) | 1.00 .001 | 0.31 .017 | 1.00 .001 | 0.77 .006 | 0.99 .002 | 0.28 .019 | 1.00 .001 | 0.76 .006 |
| WISE Wang et al. (2024a) | 0.59 .027 | 0.57 .029 | 1.00 .001 | 0.72 .013 | 0.43 .030 | 0.41 .032 | 1.00 .001 | 0.61 .015 |
| AlphaEdit Fang et al. (2025) | 0.72 .023 | 0.77 .025 | 0.74 .015 | 0.74 .012 | 0.62 .026 | 0.60 .028 | 0.71 .017 | 0.64 .014 |
| MEMOIR Wang et al. (2025) | 0.72 .015 | 0.75 .019 | 0.83 .012 | 0.77 .009 | 0.66 .017 | 0.68 .022 | 0.73 .013 | 0.69 .010 |
| Diagnostic case | ZsRE | CF |
|---|---|---|
| Sketch-sufficient (rank-1 passes) | 81.2% | 76.4% |
| Detail-needed (rank required) | 13.4% | 16.8% |
| Residual-needed (rank required) | 3.1% | 4.7% |
| Compression-regularized (sketch exact) | 4.3% | 6.7% |
| Dataset | Policy | Memory | Comp. | Utility |
|---|---|---|---|---|
| CF | exact-cache (matched) | 0.180 | 5.57 | 0.434 |
| CF | rank-1 sketch only | 0.125 | 8.00 | 0.753 |
| CF | static mixed-rank | 0.180 | 5.57 | 0.815 |
| CF | dual-budget LadderEdit | 0.180 | 5.57 | 0.862 |
| ZsRE | exact-cache (matched) | 0.178 | 5.63 | 0.420 |
| ZsRE | rank-1 sketch only | 0.124 | 8.09 | 0.765 |
| Method | Time | Peak Mem. | Persistent | Serving path |
|---|---|---|---|---|
| (ms) | (GB) | (GB) | ||
| LoRA (Exact) Hu et al. (2022) | 86 | 18.4 | 100.93 | exact-adapter lookup |
| MELO Yu et al. (2023) | 91 | 18.6 | 100.93 | retrieval |
| ELDER Li et al. (2024a) | 95 | 19.1 | 0.08 | MoE routing |
| MEMOIR Wang et al. (2025) | 104 | 17.9 | 0.14 | mask gate |
| LadderEdit | 124 | 18.8 | 19.44 | ladder lookup |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Rel. | Gen. | Loc. | Avg. | vs. matched |
|---|---|---|---|---|---|
| Matched association (Exact LoRA) | 0.95 | 0.91 | 1.00 | 0.95 | — |
| Matched association (LadderEdit) | 0.95 | 0.91 | 1.00 | 0.95 | — |
| Contriever retriever (Exact LoRA) | 0.91 | 0.86 | 0.99 | 0.92 | |
| Contriever retriever (LadderEdit) | 0.90 | 0.86 | 0.99 | 0.92 | |
| E5-large retriever (LadderEdit) | 0.88 | 0.85 | 0.99 | 0.91 | |
| MPNet (frozen) (LadderEdit) | 0.85 | 0.81 | 0.98 | 0.88 |
| Policy | Memory (norm.) | Pooled Avg. | Worst-edit Avg. |
|---|---|---|---|
| Fixed rank-1 | 0.125 | 0.85 (0.008) | — |
| Fixed rank-2 | 0.250 | 0.89 (0.007) | — |
| Fixed rank-4 | 0.500 | 0.90 (0.008) | — |
| Fixed rank-8 | 1.000 | 0.90 (0.007) | 0.62 |
| Uniform random rank | 0.562 | 0.87 (0.010) | — |
| LadderEdit (audited) | 0.19 | 0.95 (0.005) | 0.74 |
| Policy | Memory (norm.) | Avg. |
|---|---|---|
| Exact LoRA (FP16) | 1.00 | 0.94 |
| Exact LoRA (INT8) | 0.25 | 0.93 |
| LadderEdit (FP16) | 0.19 | 0.91 |
| LadderEdit (INT8) | 0.05 | 0.90 |
| Stored representation | Fraction of edits |
|---|---|
| Rank-1 sketch | 79.4% |
| Rank-2 sketch | 16.1% |
| Rank-3 sketch | 1.8% |
| Rank-4 sketch | 1.2% |
| Rank-5 sketch | 0.7% |
| Rank-6 sketch | 0.4% |
| Dataset | Audit probes/edit | Val. probes/edit | Test probes/edit |
|---|---|---|---|
| ZsRE | 2 | 1 | 2 |
| CounterFact | 3 | 1 | 3 |
| WikiBigEdit | 2 | 1 | 2 |
| Configuration | Rel. | Gen. | Loc. | Avg. |
|---|---|---|---|---|
| LadderEdit, strict split (main) | 0.95 | 0.91 | 1.00 | 0.95 |
| LadderEdit, leaky (audit=test) | 0.99 | 0.97 | 1.00 | 0.99 |
| Inflation gap | +0.04 | +0.06 | 0.00 | +0.04 |
| Dataset | Case | Edit id | Energy | Gap | Interpretation | |
|---|---|---|---|---|---|---|
| CF | Sketch sufficient | 0 | 0.517 | 0.000 | 0.000 | Rank-1 preserves the exact edit’s measured behavior, so higher resolution is unnecessary. |
| CF | Detail needed | 9 | 0.590 | -0.333 | 0.267 | The sketch drops needed behavior; the audit should promote to a higher rung. |
| CF | Compression helps | 14 | 0.474 | 0.333 | 0.000 | Removing tail directions improves measured behavior, so exact LoRA is not always a behavioral ceiling. |
| CF | Exact limited | 4 | 0.524 | 0.000 | 0.267 | Both exact and sketch leave a contract gap; this is an edit-writer limitation, not a compression error. |
| ZsRE | Sketch sufficient | 0 | 0.496 | 0.000 | 0.000 | The cheap sketch is behaviorally indistinguishable from exact on the probes. |
| ZsRE | Detail needed | 3 | 0.576 | -0.167 | 0.100 | Residual detail is needed for a hard edit that rank-1 cannot safely represent. |
| Audit threshold with | LoRA rank with | |||||||
| Rel. ( ) | 0.73 | 0.86 | 0.95 | 0.81 | 0.78 | 0.95 | 0.88 | 0.74 |
| Gen. ( ) | 0.82 | 0.88 | 0.91 | 0.85 | 0.81 | 0.91 | 0.89 | 0.83 |
| Loc. ( ) | 0.99 | 1.00 | 1.00 | 0.94 | 0.99 | 1.00 | 0.97 | 0.92 |
| Avg. ( ) | 0.85 | 0.91 | 0.95 | 0.87 | 0.86 | 0.95 | 0.91 | 0.83 |
| Memory (GB) | 4.21 | 3.97 | 3.84 | 3.62 | 2.18 | 3.84 | 7.32 | 14.18 |
| Probe bank size | Coverage | Probed edits | Unprobed edits | Overall | ||
|---|---|---|---|---|---|---|
| Avg. ( ) | Failures (%) | Avg. ( ) | Failures (%) | Avg. ( ) | ||
| 256 | 38% | 0.95 | 1.2 | 0.78 | 14.6 | 0.84 |
| 1,024 | 64% | 0.95 | 0.9 | 0.85 | 8.3 | 0.91 |
| 4,096 | 87% | 0.95 | 0.8 | 0.89 | 4.1 | 0.94 |
| 16,384 | 96% | 0.95 | 0.7 | 0.91 | 2.4 | 0.94 |
| 65,536 | 99% | 0.95 | 0.7 | 0.92 | 1.8 | 0.95 |
| Case | Prompt and target | Exact LoRA | Sketch rung | Ladder decision |
|---|---|---|---|---|
| ZsRE 29: sketch sufficient | Rewrite: Which is the position of Jules Basile Onambele? Target: winger. Generalization: What is Jules Basile Onambele’s position? | Jules Basile Onambele primarily plays as a winger or attacking midfielder. | Rank 1 gives the same rewrite and generalization output as exact LoRA. | Store rank 1. The full residual is unnecessary. |
| CF 11: sketch sufficient | Rewrite: Andreas Ivanschitz professionally plays the sport. Target: football. Generalization: After work Walther attended the Y.M.C.A. Andreas Ivanschitz, the | Andreas Ivanschitz is known for playing football (soccer) professionally. The generalization output remains incomplete. | Rank 1 preserves the same rewrite behavior and the same incomplete generalization response. | Store rank 1. A hard dataset still contains cheap compressible edits. |
| ZsRE 27: detail needed | Rewrite: What sports team was Stanislav Romanov a member of? Target: Spartak Myjava. Generalization: Which sports team was Stanislav Romanov’s? | Stanislav Romanov was a member of FC Spartak Myjava. | Rank 1 changes the value to FC Spartak Mytishchi, while rank 2 recovers FC Spartak Myjava. | Promote to rank 2. The edit needs a little more detail, but not exact storage. |
| ZsRE 24: compression helps | Rewrite: What kind of occupation does Karim-Mohamed Maamoun have? Target: architect. Generalization: What kind of occupation is Karim-Mohamed Maamoun? | Exact LoRA answers archaeologist on the rewrite and generalization prompts. | Rank 1 and rank 2 include archaeologist and architect in the rewrite output. | Store the sketch. Low-rank projection can filter harmful detail from the exact edit. |
| Case | Exact LoRA | Sketches | Interpretation |
| CF 27, target Russian. Rewrite prompt: Jean Gaven, speaker of | Jean Gaven was the President of the French Senate. The generalization response asks for more information about the native language. | Rank 1 and rank 2 produce the same incorrect or uninformative behavior. | Exact-limited. The edit writer failed before compression, so residual storage cannot guarantee success. |
| Quantity | Value |
|---|---|
| Crossover horizon | |
| Exact LoRA @ | 86 ms/edit |
| LadderEdit @ | 93 ms/edit |
| Exact LoRA @ | 850 ms/edit |
| LadderEdit @ | 150 ms/edit |
| Speedup @ |
| Predictor | Utility | Audits/edit | Memory |
|---|---|---|---|
| Random rank proposal | 0.81 | 2.31 | 0.22 |
| Always rank-1, then promote | 0.89 | 1.85 | 0.20 |
| Spectral tail only | 0.88 | 1.32 | 0.19 |
| Margin only | 0.86 | 1.42 | 0.21 |
| Spectral margin (ours) | 0.91 | 1.21 | 0.19 |
| Oracle (exhaustive audit) | 0.93 | 8.00 | 0.19 |
| Dataset | Rank | Memory | Energy | Contract | False Ret. | |
|---|---|---|---|---|---|---|
| CF | 32 | 1 | 0.125 | 0.560 | 0.625 | 1.67 |
| CF | 32 | 2 | 0.250 | 0.747 | 0.656 | 0.00 |
| CF | 32 | 4 | 0.500 | 0.899 | 0.646 | 0.00 |
| CF | 32 | 8 | 1.000 | 1.000 | 0.646 | 0.00 |
| CF | 128 | 1 | 0.125 | 0.559 | 0.641 | 7.00 |
| CF | 128 | 2 | 0.250 | 0.744 | 0.662 | 1.67 |
| Backbone | Method | Acc. | Gen. | Loc. | Avg. |
|---|---|---|---|---|---|
| Qwen2.5-14B | Exact LoRA | 0.795 | 0.905 | 0.999 | 0.900 |
| LadderEdit | 0.790 | 0.900 | 0.999 | 0.896 | |
| Qwen2.5-32B | Exact LoRA | 0.821 | 0.923 | 0.999 | 0.914 |
| LadderEdit | 0.816 | 0.918 | 0.999 | 0.911 |
| Cache policy | Full-stream utility |
|---|---|
| Random retention | 0.18 |
| FIFO | 0.20 |
| LRU | 0.22 |
| Learned retention | 0.35 |
| Learned retention + base fallback | 0.41 |
| Dataset | Policy | Norm. memory | Utility |
|---|---|---|---|
| CF | Exact-cache | 0.180 | 0.434 |
| CF | Rank-1 for all edits | 0.125 | 0.753 |
| CF | Static mixed-rank | 0.180 | 0.815 |
| CF | LadderEdit | 0.180 | 0.862 |
| ZsRE | Exact-cache | 0.178 | 0.420 |
| ZsRE | Rank-1 for all edits | 0.124 | 0.765 |
| Method | Insert / edit | Peak GPU | Persistent memory |
|---|---|---|---|
| Exact LoRA | 86 ms | 18.4 GB | 100.93 GB |
| LadderEdit | 124 ms | 18.8 GB | 19.44 GB |
| Difference | +38 ms | +0.4 GB | GB |
| Diagnostic property | ZsRE | CounterFact |
|---|---|---|
| Rank-1 sufficient | 81.2% | 76.4% |
| Rank 2 or higher required | 13.4% | 16.8% |
| Full acquisition rank required | 3.1% | 4.7% |
| Sketch passes while exact fails | 4.3% | 6.7% |
| Selection mechanism | Exact LoRA Avg. | LadderEdit Avg. |
|---|---|---|
| Matched association | 0.95 | 0.95 |
| Contriever ( Izacard et al., 2021 ) | 0.92 | 0.92 |
| E5-large ( Wang et al., 2022 ) | – | 0.91 |
| Frozen MPNet ( Song et al., 2020 ) | – | 0.88 |
| Rank proposal | Utility | Audits / edit |
|---|---|---|
| Random | 0.81 | 2.31 |
| Always rank-1, then promote | 0.89 | 1.85 |
| Spectral tail only | 0.88 | 1.32 |
| Behavioral margin only | 0.86 | 1.42 |
| Spectral tail + margin | 0.91 | 1.21 |
| Exhaustive oracle | 0.93 | 8.00 |
| Probe coverage | Overall Avg. |
|---|---|
| 38% | 0.84 |
| 64% | 0.91 |
| 87% | 0.94 |
| 96% | 0.94 |
| 99% | 0.95 |