FedAlphaEdit: Null-Space-Aligned Merging for Collaborative Knowledge Editing
Authors: Sota Sugawara, Yukihiko Okada
Organizations: Graduate School of Science and Technology, University of Tsukuba, Tsukuba, Japan · Center for Artificial Intelligence Research, Tsukuba Institute for Advanced Research, University of Tsukuba, Tsukuba, Japan · Institute of Systems and Information Engineering, University of Tsukuba, Tsukuba, Japan
Multiple institutions may each hold their own private knowledge edits and wish to integrate them into a single large language model without sharing raw edit requests. Null-space-constrained editing methods such as AlphaEdit mathematically guarantee that each update leaves unrelated knowledge intact, while collaborative frameworks such as CollabEdit aggregate edits from multiple clients without data sharing. Combining the two appears trivial. However, we show that this naive combination fails structurally, and we identify its cause. Guided by this analysis, we propose FedAlphaEdit. To our knowledge, this is the first collaborative knowledge editing framework that aligns both local editing and the server-side merging rule under a single null-space principle for preserving existing knowledge. FedAlphaEdit builds on null-space-aligned merging, in which clients share projected statistics and the server provably recovers the result of editing everything in one place under a one-shot idealization. Empirically, the proposed method repairs the collapse and brings edit success and preservation simultaneously close to the level of centralized editing across two architecture families. FedAlphaEdit thus lets institutions that cannot share raw edit data, such as hospitals and financial firms, jointly maintain a shared model that closely approximates editing all facts in one place.
Figures & tables
Centralized
Naive
Δ
ES
97.08
79.69±0.88
−17.39
NS
64.43
60.24±0.11
−4.19
oons
≈3×10−4% (floor)
14.0–14.1%
—
Table 1: Naive combination versus the centralized reference on GPT-2 XL (MCF, 5000 edits, 10×500 ). Centralized is a single deterministic run; Naive is three-seed mean ± SD. Metric definitions are in Sections 4.1 and 3.5 .
Figure 1: The naive combination and FedAlphaEdit as the same total workload ( 5000 edits) is partitioned across a growing number of clients N : neighborhood specificity (left), edit success (middle), and oons (right; the FedAlphaEdit series is at the measurement floor throughout). Dashed lines mark the centralized reference ( 64.43 and 97.08 ). All points are three-seed mean ± SD. The per-client edit count co-varies with N (Section 4.2 ).
ES
NS
oons
Centralized AlphaEdit
97.08
64.43
≈3×10−4% (floor)
Centralized MEMIT
90.80
64.42
26.6%
CollabEdit (MEMIT)
90.75±0.03
64.34±0.13
26.5±0.05%
FedLEKE (our reimplementation)
87.74±0.51
66.78±0.54
23.98±0.02%
Naive (AlphaEdit + CollabEdit)
79.69±0.88
60.24±0.11
14.0–14.1%
CollabEdit (MEMIT) + null-space projection
82.45±0.06
71.90±0.03
≈3×10−4% (floor)
Table 2: FedAlphaEdit against the centralized reference and the naive combination on GPT-2 XL (MCF). All rows edit the same 5,000 facts, single-round except the five-round variant ( 10×100×5 ) and the FedLEKE reimplementation (ten time slots). Centralized rows are single deterministic runs; all client-partitioned rows are three-seed mean ± SD. NS is bounded above by the unedited model, which by construction attains NS =100 , so a higher NS at a collapsed ES indicates weaker editing rather than better preservation, as the simple-average row illustrates. The simple-average and unnormalized-sum rows are the scalar merges γ∑iΔi with γ=1/N and γ=1 .
Qwen2.5-1.5B
ES
NS
Centralized
98.28
73.32
Naive
93.53±0.58
71.84±0.75
FedAlphaEdit
96.40±0.27
74.32±0.16
Table 3: Generality on Qwen2.5-1.5B (MCF, 10 clients × 500 edits). Centralized is a single deterministic run; Naive and FedAlphaEdit are three-seed mean ± SD.
zsRE
Centralized
Naive
FedAlphaEdit
GPT-2 XL
Rewrite acc
89.40
66.43±1.27
89.40±0.41
Paraphrase acc
81.20
62.22±0.39
81.60±0.16
Neighborhood acc
24.98
23.72±0.26
25.12±0.11
oons
3.17×10−4% (floor)
15.75±0.04%
3.17×10−4% (floor)
Qwen2.5-1.5B
Table 4: Dataset generality on zsRE (each 10 clients × 200 edits, single round). Naive and FedAlphaEdit are three-seed mean ± SD; Centralized is a single run; oons of order 10−4% is the measurement floor. zsRE accuracies are first-token exact-match ( ×100 ); oons is a weight-space five-layer mean. No numerical comparison with MCF or between the two model families is intended.
Per-client 500
ES
Centralized ES
NS
oons
N=2 ( 1000 )
99.83±0.12
99.80
68.37±0.07
4.05%
N=5 ( 2500 )
95.03±0.15
99.24
62.80±0.17
9.52%
N=10 ( 5000 )
79.69±0.88
97.08
60.24±0.11
14.04%
Table 5: Per-client-fixed control series (naive combination, GPT-2 XL, per-client 500 ); three-seed mean ± SD, with oons standard deviations 0.027 , 0.057 , and 0.022 percentage points. The N=10 row reuses the main naive runs. Centralized entries are deterministic single runs editing the same case sets as the corresponding rows; the N=10 centralized entry reuses the main centralized run. Total edits co-vary with N ; cross-column comparisons are within the same row (matched total edit count) only.
Figure 2: Sequential drift ∥Δproposed−ΔcentralAlpha∥F/∥ΔcentralAlpha∥F and operator divergence ∥Nall−NG∥F/∥NG∥F increase monotonically across the edited layers ( 13 – 17 ), peaking at the last; five-layer means 0.4547 and 0.1879 . Markers show the layer- 13 and layer- 17 endpoints.
GLUE avg. F1
base
0.4013
Centralized
0.3728
Naive
0.4087
FedAlphaEdit
0.3732±0.0161
Table 6: GLUE average F1 over six tasks ( n=100 per task), a supportive, non-binding observation. FedAlphaEdit is the three-seed mean ± SD; base and Centralized are single values, and Naive is a three-seed mean. The naive combination sits nearer the base level than centralized editing does, consistent with the distorted merge writing edits more weakly (Section 4.2 ) rather than an advantage.
λ
ES
NS
oons
10000
84.28
62.19
15.80%
20000
78.84
60.12
14.05%
40000
73.72
58.05
11.98%
Table 7: Sensitivity of the naive combination to the shared-covariance weight λ (GPT-2 XL, MCF, 10×500 , seed 1). Varying λ trades edit success and specificity against oons ; no setting recovers both axes toward the centralized level.
γ
ES
NS
oons
0.1 ( =1/N )
31.23±0.22
77.16±0.17
3.143×10−4%
0.5
55.16±0.47
54.98±0.42
3.143×10−4%
1.0
51.17±0.27
51.07±0.29
3.143×10−4%
Table 8: Scalar merges γ∑iΔi (GPT-2 XL, MCF, 10×500 ), three-seed mean ± SD. The per-client AlphaEdit updates of the simple-average arm are reused and only evaluated, with no new solves. The row γ=0.1 is the simple-average row of Table 2 . Only three levels of γ were tested, and finer tuning is untested. oons is at the measurement floor in every row by construction of a scalar merge.
Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guaranteeing that edits do not disrupt previously preserved knowledge, and reports substantial gains over existing editing methods on LLaMA3, GPT2-XL, and GPT-J. In this work, we present a reproducibility study of AlphaEdit, reproducing its reported results under the original experimental setup and extending the evaluation along three axes: new model architectures, additional downstream benchmarks, and substantially longer sequential editing horizons. We successfully reproduce AlphaEdit's reported metrics across the original models, though we identify a discrepancy in the reported fluency and consistency metric. Extending AlphaEdit to newer model families, we find that its advantage does not generalize uniformly, which we trace to architectural assumptions in the locate-then-edit paradigm that are violated by these newer models. We further stress-test AlphaEdit's central sequential-editing claim by extending the number of edits well beyond those evaluated in the original paper, and find that performance, which is stable at the originally reported scale, degrades as edits reach a much higher count, indicating that the null-space projection's protection against catastrophic forgetting is bounded rather than unconditional. Finally, we extend evaluation of edited models on three extra benchmarks, namely, BoolQ, HellaSwag, and XSTest, and we find that large-scale sequential editing degrades both general downstream task competence and safety-relevant refusal behavior. Our results confirm that AlphaEdit performs as reported within its original scope, while showing that its core theoretical guarantees are sensitive to model architecture and editing scale in ways that have practical implications for its deployment.
Single-edit updates in large language models can trigger ripple effects across local knowledge neighborhoods: desirable propagation to related facts and unintended perturbation of preserved ones. Existing methods address these two effects separately, without explicitly modeling their coupling. We challenge this separation through an analysis of ripple responses across typical baselines, identifying two coupled design pressures: editable-side coordination and preserved-side leakage. We propose Joint Neighborhood Optimization (JNO), a new knowledge-editing framework to formalize and jointly address both pressures at the target-planning stage. JNO instantiates this principle through Pressure-Aware Coordination (PAC), which jointly optimizes neighborhood target representations under coupled constraints, and a semantic pre-execution gate that rejects high-risk target plans before parameter execution. Experiments on RippleEdits show JNO improves propagation and preservation metrics by at least 7.0% while preserving cross-backbone editing stability.
Haoben Huang, Shuxin Liu, Ou Wu +1
Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou, China · College of Information Engineering, Zhejiang University of Technology, Hangzhou, China
Null-space-based methods have garnered considerable attention in model editing by constraining updates to the null space of the pre-existing knowledge representation, thereby preserving the model's original behavior. However, in practice these methods rely on an approximate null space--leading to knowledge leakage--and further suffer from severe performance degradation during sequential editing. Recent work shows that history-aware editing strategies can empirically mitigate this decline, yet the underlying reason remains unclear. In this paper, we first expose the knowledge leakage inherent in existing null-space approaches and then analyze why history-aware updates effectively preserve both editing performance and general capabilities during long-horizon editing. Building on these insights, we propose BetaEdit, a refined framework that effectively controls the knowledge leakage and integrates history-aware updates into the null-space paradigm. Extensive experiments on three large language models across two standard benchmarks show that BetaEdit consistently outperforms prior methods in the challenging regime of massive-scale sequential editing. Code is available at: https://github.com/lbq8942/BetaEdit.
Bingqing Liu, Wei Liu, Yuhua Li
Huazhong University of Science and Technology, Wuhan, China