Recent progress in LLM agents has advanced the prospect of autonomous research. Yet whether AI can complete difficult long-horizon tasks, especially those that advance AI research itself, remains largely unexplored. We present ASI-Arch, a system for AI-driven AI research that autonomously conducts neural architecture research through a closed-loop research-experiment-analyze-update process. Applied to linear attention, ASI-Arch ran 1,773 iterative experiments and discovered 105 state-of-the-art architectures. Its best architecture improves over DeltaNet by nearly three times the gain achieved by Mamba2. Beyond the final performance gains, we analyze the contributions of different parts of the framework in this hard research setting, shedding light on what enables autonomous progress in complex AI research tasks.
Figures & tables
Figure 1: Overview of the ASI-Arch framework. The Researcher proposes new architectures and implements their code using knowledge retrieved from the Cognition Base and experience accumulated from previous iterations. The Engineer then trains and evaluates the implementations, invoking a debugging agent when needed. The Analyzer interprets the resulting multifaceted experimental evidence and distills it into insights that guide the Researcher in the next iteration. This feedback loop supports iterative improvement.
Figure 2: Cumulative count of discovered SOTA architectures vs. number of exploration experiments. The near-linear trend confirms that breakthroughs grow steadily with compute.
Figure 3
Benchmarks
DeltaNet
Gated- DeltaNet
Mamba2
PG
C
FG
H
AM
Development
Wiki ppl ↓
17.00
16.84
16.66
16.18
16.05
15.77
16.65
16.26
LMB ppl ↓
13.63
13.31
13.33
12.62
13.45
12.34
13.06
13.75
LMB
45.47
46.26
46.24
47.60
46.13
47.53
46.56
45.04
PIQA
73.12
74.10
73.78
72.91
74.37
72.91
74.37
74.10
Hella
56.29
57.55
58.58
56.99
57.00
58.47
56.85
57.17
Table 1: Top block: 10 development benchmarks, used in our exploration stage; bottom block: 6 generalization benchmarks for out-of-distribution testing. Dev., Gen., and Overall averages cover the 8 development, 6 generalization, and all 14 accuracy metrics, respectively; perplexity is reported separately. Bold indicates the best results, and underlining indicates the second-best results. Model abbreviations are as follows: PG = PathGateFusionNet, C = ContentSharpRouter, FG = FusionGatedFIRNet, H = HierGateNet, and AM = AdaMultiPathGateNet.
Figure 4: (a) Average raw benchmark score and loss of top-50 candidates vs. cumulative samples. (b) Composite fitness and its components vs. cumulative samples. The fitness plateau is a designed effect of the sigmoid transformation, not a performance ceiling.
Figure 6
Figure 6: Component-mention distributions: model gallery (top 105) vs. remaining architectures.
Family
Baseline avg.
Variant avg.
Gain
Mamba2
47.84
48.19
+0.35
HGRN2
47.67
47.81
+0.14
GLA
45.55
47.60
+2.05
Table 3: Summary of cross-family transfer at 340M parameters. Each row reports the baseline and the selected variant; full benchmark results are reported in Appendix Table 8 .
Model
Train tok/s ↑
Prefill tok/s ↑
Peak mem. ↓
DeltaNet
5,063
23,162
11.44G
Gated DeltaNet
4,959
22,413
11.54G
Mamba2
4,476
12,725
11.29G
ASI-Arch -1
4,037
13,589
18.32G
ASI-Arch -2
4,168
12,237
16.28G
ASI-Arch -3
3,426
11,932
18.39G
Table 4: End-to-end efficiency at 1.3B scale under the same implementation and hardware setting.
Strategy
Rate
Loss ↓
Benchmark ↑
ASI-Arch
37.4%
4.1043 (6.4%)
0.3500 (9.2%)
OpenEvolve
8.3%
4.3454 (0.9%)
0.3405 (6.2%)
Table 5: Strategy-level comparison under the same proposal model and attempt budget.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Run 1
Run 2
Difference
Final best score
2.5161
2.4311
0.0850 (3.44%)
AUC
281.3209
280.0682
1.2527 (0.45%)
Std. dev. from average curve
0.0798
0.0798
–
Appendix
Table 6: Two independent full-pipeline reruns.
Metric
Human Curated
AI Curated
Final best score
2.5161
2.5500
AUC
281.3209
285.6003
Best score at 20 results
2.3518
2.3526
Best score at 40 results
2.3518
2.3526
Best score at 60 results
2.3518
2.4510
Best score at 80 results
2.3518
2.5500
Appendix
Table 7: Comparison of human-curated and AI-curated cognition sources over the first 120 results.
Model
Wiki ↓
LMB ↓
LMB ↑
PIQA ↑
Hella ↑
Wino ↑
ARC-e ↑
ARC-c ↑
SIQA ↑
BoolQ ↑
Avg.
Mamba2
27.08
40.09
31.32
67.90
42.25
51.46
62.04
29.27
39.25
59.24
47.84
Mamba2 variant
26.79
38.46
32.49
67.52
42.40
51.30
63.22
30.29
38.49
59.79
48.19
HGRN2
28.09
36.02
32.39
68.82
41.42
51.30
61.03
29.01
37.92
59.48
47.67
HGRN2 variant
27.56
38.31
32.45
68.28
41.33
51.07
61.03
29.14
38.95
60.21
47.81
GLA
30.83
53.44
28.55
65.61
38.04
51.14
57.87
26.54
37.51
59.14
45.55
GLA variant
27.48
36.08
32.84
67.52
41.51
50.67
62.21
28.16
38.59
59.33
47.60
Appendix
Table 8: Complete cross-family transfer results at 340M parameters. For each family, we report the baseline and the selected variant. Metrics follow the main setup: Wiki ppl, LAMBADA ppl/acc, PIQA, HellaSwag (normalized), WinoGrande, ARC-Easy, ARC-Challenge (normalized), SIQA, BoolQ, and the average across accuracy metrics. Bold indicates the better result within each family pair.
Model Name
20M params / 1B tokens
340M params / 1B tokens
Train Loss ↓
Test Score ↑
Train Loss ↓
Test Score ↑
DeltaNet (Baseline)
4.5749
36.23
3.5055
41.16
Gated DeltaNet (Baseline)
4.5678
36.60
3.4768
42.10
AdaptiveContextFusionNet
4.4973 -0.0705
37.03 +0.43
3.4624 -0.0144
42.74 +0.64
AdaptiveEntropyGateNet
4.4423 -0.1255
36.91 +0.31
3.4558 -0.0210
42.37 +0.27
AdaptiveEntropyRouter
4.3547 -0.2131
39.26 +2.66
3.4066 -0.0702
44.31 +2.21
Appendix
Table 9: Model performance comparison. Train Loss represents the loss at the final training step. Test Score is the average performance across 7 tasks: ARC-Challenge, ARC-Easy, BoolQ, HellaSwag, PIQA, Social IQA, and WinoGrande. Green subscripts indicate improvements over the Gated DeltaNet baseline.