High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models and memory hierarchies differ substantially from those of GPUs. To address this transfer gap, post-training methods adapt LLMs to NPU programming but depend on scarce expert data and substantial training compute. Memory-learning agents instead adapt through external memory, but their uniform credit assignment gives adopted and unused experiences the same reward target, potentially biasing subsequent retrieval rankings. Moreover, when learned values guide only retrieval, high-value experiences that generalize across operators must be retrieved repeatedly rather than retained in context, thereby increasing retrieval overhead and weakening cross-task guidance. We therefore present SAGE, a persistent self-improving agent for NPU kernel synthesis. Adoption-Traced Utility estimation (ATU) combines explicit adoption records with kernel evaluation outcomes for adoption-aware credit assignment. Utility-Gated Consolidation (UGC) uses positive utility and repeated adoption across operators to select and abstract reusable rules into a bounded resident context. On NPUKernelBench, SAGE achieves a 95.5% execution rate versus 84.1% for the strongest controlled baseline, with 86.9% of solved operators outperforming torch_npu. With GLM-5.3, SAGE achieves a 43.99x speedup over the torch_npu reference on sparse flash attention. These results show that adoption-aware credit assignment and selective consolidation enable agents to accumulate and reuse hardware-specific knowledge across tasks.
Figures & tables
Figure 1: Retrieval-oriented memory learning with uniform credit and repeated retrieval (left) versus SAGE’s adoption-aware credit assignment and selective consolidation into resident context (right).
Figure 2: Overview of SAGE. (a) A frozen agent iteratively refines kernels using retrieved cold experiences, resident hot rules, and execution feedback. (b) ATU combines adoption traces with terminal outcomes to estimate item-level utility. (c) UGC uses utility and repeated adoption across operators to select reusable rules for a bounded resident context and demote underused rules.
Model
Method
Execution Rate
Overall
Performance
L1
L2
L3
CR
ER
Fast1.0
Sself
Qwen3-Coder-Next (80B)
Refinement
5.0
0.0
0.0
35.2
1.1
0.0
1.00 ×
Static RAG
30.0
26.0
0.0
51.1
21.6
21.1
1.10 ×
Value Memory
50.0
24.0
5.6
65.9
26.1
17.4
1.16 ×
SAGE
55.0
32.0
5.6
75.0
31.8
28.6
1.27 ×
DeepSeek-V4-Flash (284B)
Refinement
35.0
18.0
0.0
55.7
18.2
31.3
1.20 ×
Table 1: Main results on the 88-operator NPUKernelBench subset. All rates are reported as percentage. The best result for each model is marked in bold.
Task schedule
L3 ER
Fast1.0
L3 Only
38.9
42.9
L2+L3 Mixed
61.1
63.6
L2 → L3
77.8
71.4
Table 2: Cross-difficulty transfer to the 18 Level-3 operators. All rates are percentages.
TilingData field types must match their kernel-side consumers.
32.47
No
Appendix
Table 8: Retrieved experiences for the GatherV3 correctness-stage runtime failure. Relevance is the retriever score; adoption is determined from the subsequent diagnosis and code revision.
Iteration
Diagnosis
Revision
Outcome
0–2
A shared Matmul object changed its K dimension despite fixed tiling; the default task type also incorrectly assigned Cube work on AIV core.
Separate the QK and PV configurations and declare KERNEL_TYPE_MIX_AIC_1_0 .
Sun Yat-sen University, Zhuhai, China · Greater Bay Area National Technology Innovation Center, Guangzhou, China · Peng Cheng Laboratory, Shenzhen, China