Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each of 15 iterations the model reads the measured outcome of the previous one and edits the Chisel RTL, the memory configuration and the sparse-kernel schedule of a Gemmini accelerator through MCP tool servers, and no candidate counts until it has been checked for legality, elaborated, simulated cycle-accurately, checked bit-for-bit on every output against a golden reference, and synthesised. The harness turns each measurement into the next work order, a diagnosed bottleneck with matching strategy guidance, the history of tried designs and a score of the model's own prediction, and a second model repairs changes that fail a gate. On a 512×512 GraphChallenge sparse-DNN layer the loop reaches 2.1x fewer cycles, 9.8x less off-chip traffic and 22.8% less area than the block-sparse Gemmini baseline, with 5.61x higher modelled perf/W and 11.8x lower EDP. The levers span three layers: a schedule that keeps the dense operand resident removes 9.8x of the traffic, a zero-gated MAC and a zero-row skip unit that the model wrote in Chisel cut energy, and resizing the memories cuts area.
Figures & tables
Fig. 1: The SparseCraft loop on CHIA, with the LLM inside it. Purple: LLM agents; blue: EDA tool runs in containers; white: harness checks. The proposer ( N10 ) edits the tree through an MCP tool, and the harness carries the change through the gate ladder ( N13 – N53 ), with synthesis ( N52 ) beside simulation. A failed gate (red) goes to the repairer ( N73 ), whose fixed tree re-enters at N13 . N61 turns the measurement into the next work order (thick purple): diagnosis, strategy modules, counters, tried designs and prediction score. Stage labels match Algorithm 1 .
Fig. 2: Every iteration, along the design lineage. Admitted designs (solid green) form the serpentine, each the parent of the next; rejected proposals (dotted red) hang from the parent they were proposed from. Blocks give perf/W relative to iteration 1 and area.
Fig. 3: Metric evolution , normalised to iteration 1; titles give the absolute values at iterations 1 and 15. The step line is the design held after each iteration.
Fig. 4: Wall-clock profile of the loop , from the CHIA profile log.
metric
it. 1
it. 15
change
cycles
M
106,650
50,862
2.1 × fewer
off-chip bytes
M
3,211,264
327,680
9.8 × less
area (mm 2 )
M+e
4.035
3.115
− 22.8%
energy ( μ J)
m
95.67
17.04
5.61 × less
power (W)
m
0.449
0.168
2.68 × less
performance (GOPS)
M
4.92
10.31
2.1 ×
TABLE I: Baseline (iteration 1) against the best design (iteration 15); neither produces a wrong output. M: measured; M+e: measured logic plus estimated SRAM area; m: modelled.