Code language models must be maintained like the software around them: when a library evolves, a model keeps writing the interface that it saw during training. Repairing the model itself lets one correction reach all downstream uses. Existing repair methods attribute a failure to neurons, select the highest-ranked ones, and apply a generic update. This pipeline assumes that the attributed neurons are the ones to patch and that a generic update fits every failure, and neither assumption has been examined. We examine both on executable API evolution tasks in Python and Rust with three code models and identify two gaps. The targeting gap separates the neurons that failure attribution targets from the neurons that can carry a patch: their top sets have a Jaccard overlap of only 0.15 to 0.20. The tailoring gap separates a generic update from a patch built for the failure: the patches that different carrier neurons need are nearly orthogonal. To address both gaps, we propose ASTRA. It targets neurons by contrastive semantics, an attribution that scores a neuron by its contribution to the logit contrast between the target token and the produced token. It then tailors the patch by solving one small linear system in closed form, which corrects all failing tokens of a sample jointly and needs neither an optimizer nor a backward pass. Contrastive semantics selects significantly better carrier neurons than gradient-based attributions in 3 of 6 settings and comparable ones in the others. On average, ASTRA reaches 66.7 percent Pass@1, against 48.4 percent for the best of AlphaEdit, STAR and low-rank adaptation, and repairs a sample in 2.7 seconds. It is the best method in all 6 settings and for every type of API change, and this advantage persists under an unseen phrasing of the test prompt. Its side effects on unrelated code are small on the large models and larger on the small one.
Figures & tables
Figure 1. Overview of Astra on a running example. The three panels are the steps of one round, and the loop below them repeats the first and the third step.
Model
Benchmark
Tokens
Correlation ↑
Overlap ↑
Cosine ↑
Shift ↓
MiniCPM5-1B
PyEvo
293
0.50
0.17
+0.001
0.11
RustEvo
122
0.53
0.18
+0.009
0.14
OpenCoder-8B
PyEvo
247
0.50
0.18
− 0.001
0.13
RustEvo
92
0.49
0.15
− 0.001
0.17
OxCoder-9B
PyEvo
221
0.37
0.20
0.000
0.10
RustEvo
85
0.39
0.19
+0.002
0.12
Table 1. Gap statistics between failure attribution and the natural carriers of the patch, averaged over the shared subset of each setting. Tokens counts the failing tokens per sample, Correlation the rank correlation, Overlap the top-set overlap, Cosine the patch cosine, and Shift the attribution shift after the first round. Correlation and Overlap quantify the targeting gap, Cosine the tailoring gap, and Shift the coupling of the two gaps. The arrows point toward a smaller gap.
Figure 2. Distribution of the rank correlation, the top-set overlap and the shift over the samples of the shared subset of each setting. Boxes show the median and the quartiles, and whiskers the extremes.
Pass@1 ↑
Relative patch norm ↓
Model
Benchmark
Input × Grad.
Cont. grad.
Cont. sem.
Input × Grad.
Cont. grad.
Cont. sem.
MiniCPM5-1B
PyEvo
29.8
35.6
37.0
0.19
0.16
0.10
RustEvo
64.0
68.6
74.6
0.10
0.09
0.07
OpenCoder-8B
PyEvo
25.0
27.9
61.1
0.55
0.46
0.17
RustEvo
78.8
78.0
89.0
0.37
0.35
0.24
OxCoder-9B
PyEvo
40.4
42.3
46.2
0.30
0.26
0.09
Table 2. Pass@1 in percent and median relative patch norm of three attributions under the same closed-form patch and neuron budget, on the shared subset. Input × Grad., Cont. grad. and Cont. sem. denote Input × Gradient, contrastive gradient and contrastive semantics. The best Pass@1 of each setting is in bold.
Figure 3. Relative patch norm under the three attributions on the shared subset of all 6 settings. Bars show the median over the repairs and whiskers the interquartile range. A lower norm means that the targeted neurons carry the patch under a smaller change, which addresses the targeting gap.
Figure 4. Layers that Astra selects in the main comparison (1330 repairs per model). Each line shows the share of the repairs of a model that select a layer against its relative depth, the layer number divided by the number of layers. No repair selects a layer left of the plotted range.
Model
Benchmark
Unrepaired
AlphaEdit
STAR
LowRank
Astra
MiniCPM5-1B
PyEvo
7.4
17.0
25.1
7.7
40.2
RustEvo
6.2
26.4
45.2
11.6
77.1
OpenCoder-8B
PyEvo
17.4
31.2
39.7
23.5
61.9
RustEvo
37.1
57.8
64.0
52.4
90.8
OxCoder-9B
PyEvo
20.3
37.1
44.1
45.2
47.3
RustEvo
32.5
62.9
72.5
74.3
82.8
Table 3. Pass@1 in percent in the main comparison, on all 622 PyEvo and 708 RustEvo samples per model. Unrepaired is the model without a repair. The best value of each row is in bold, and higher values are better.
Signature change
Implicit change
Stabilized function
Deprecated API
Model
Benchmark
Best
Astra
Best
Astra
Best
Astra
Best
Astra
MiniCPM5-1B
PyEvo
23.3
31.6
27.9
40.3
20.0
40.5
36.8
55.9
RustEvo
52.9
80.6
47.0
79.7
37.8
71.6
22.2
70.4
OpenCoder-8B
PyEvo
32.3
48.1
41.8
62.7
36.4
66.4
58.8
72.1
RustEvo
66.9
92.1
76.0
93.1
52.7
87.8
33.3
85.2
OxCoder-9B
PyEvo
41.4
40.6
46.3
49.8
44.1
44.5
61.8
61.8
Table 4. Pass@1 in percent by type of API change in the main comparison, on all samples of the 6 settings. PyEvo has 133, 201, 220 and 68 samples of the four types and RustEvo has 242, 217, 222 and 27. Best is the best baseline for the setting and the type, the better value of each pair is in bold, and higher values are better.
PyEvo
RustEvo
Model
Method
Pass ↑
Test ↓
API ↓
Other ↓
Pass ↑
Comp ↓
API ↓
Test ↓
Other ↓
MiniCPM5-1B
Unrepaired
7.4
74.3
15.4
2.9
6.2
55.2
36.3
2.1
0.1
AlphaEdit
17.0
70.3
8.4
4.3
26.4
44.8
23.3
2.5
3.0
STAR
25.1
54.8
11.3
8.8
45.2
38.7
10.9
1.8
3.4
LowRank
7.7
80.4
9.6
2.3
11.6
41.5
45.2
1.1
0.6
Astra
40.2
52.3
4.5
3.1
77.1
15.1
3.4
1.4
3.0
Table 5. Outcomes of the main comparison in percent of all samples of the 6 settings. Pass is an answer that passes all tests, Test an answer that fails a test, API an answer that does not use the required interface, Comp an answer that does not compile, and Other an unusable answer, one without the required function signature, without extractable code or with a timeout. The arrows mark the better direction.
Figure 5. Mean repair time per sample in seconds in the main comparison, without the generation of the answer.
Model
Benchmark
Astra
Sequential
Margin 1
One view
Two views
MiniCPM5-1B
PyEvo
37.0
9.1
8.2
28.8
32.7
RustEvo
74.6
27.5
23.7
66.1
69.5
OpenCoder-8B
PyEvo
61.1
26.0
32.2
56.2
56.7
RustEvo
89.0
56.4
56.4
87.3
85.2
OxCoder-9B
PyEvo
46.2
18.3
27.9
37.0
43.3
RustEvo
80.5
47.5
53.8
73.7
75.0
Table 6. Ablation study of Astra : Pass@1 in percent of Astra and of variants that change one component, on the shared subset of all 6 settings. Sequential corrects the failing tokens one after another, Margin 1 replaces the margin of five logits by one logit, One view fits only the original repair prompt, and Two views uses two views and eight neurons per answer token. The best value of each row is in bold, and higher values are better.
Model
Method
Acc-APD ↑
EM-APD ↑
BLEU-RPD ↑
MiniCPM5-1B
AlphaEdit
− 3.26
+1.35
− 10.77
STAR
− 0.20
+0.36
− 1.66
LowRank
+0.54
+0.38
− 0.15
Astra
− 0.69
+0.71
− 4.11
OpenCoder-8B
AlphaEdit
− 0.36
+0.76
− 1.42
STAR
+0.43
+0.00
+0.16
Table 7. Side effects on HumanEval after the repair of one problem, averaged over 132 repairs per model and method. Acc-APD and EM-APD are the absolute changes of the teacher-forced token accuracy and of exact match in points, and BLEU-RPD is the relative change of BLEU in percent, each measured on the 32 held-out problems against the unrepaired model, so a negative value is a loss. The arrows mark the better direction.
Unrepaired
AlphaEdit
STAR
LowRank
Astra
Model
Benchmark
Orig
Spec
Orig
Spec
Orig
Spec
Orig
Spec
Orig
Spec
MiniCPM5-1B
PyEvo
7.2
8.7
15.9
22.1
22.1
27.4
5.3
13.9
37.0
55.8
RustEvo
5.1
3.8
25.4
33.1
41.9
46.6
12.3
16.9
74.6
85.2
OpenCoder-8B
PyEvo
13.5
12.0
28.8
27.4
35.6
37.0
19.2
18.3
61.1
59.6
RustEvo
34.7
23.7
53.8
40.3
62.3
51.3
47.5
44.9
89.0
85.6
OxCoder-9B
PyEvo
15.4
15.9
35.6
35.1
40.9
43.3
42.3
51.4
46.2
52.9
Table 8. Pass@1 in percent of all methods under the original test prompt (Orig) and under a specification prompt (Spec), a phrasing that resembles none of the views, on the shared subset of all 6 settings. The best value of each row and phrasing is in bold, and higher values are better.