SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing
Organizations: International Institute of Information Technology, Hyderabad, India
Abstract
Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each of 15 iterations the model reads the measured outcome of the previous one and edits the Chisel RTL, the memory configuration and the sparse-kernel schedule of a Gemmini accelerator through MCP tool servers, and no candidate counts until it has been checked for legality, elaborated, simulated cycle-accurately, checked bit-for-bit on every output against a golden reference, and synthesised. The harness turns each measurement into the next work order, a diagnosed bottleneck with matching strategy guidance, the history of tried designs and a score of the model's own prediction, and a second model repairs changes that fail a gate. On a GraphChallenge sparse-DNN layer the loop reaches 2.1x fewer cycles, 9.8x less off-chip traffic and 22.8% less area than the block-sparse Gemmini baseline, with 5.61x higher modelled perf/W and 11.8x lower EDP. The levers span three layers: a schedule that keeps the dense operand resident removes 9.8x of the traffic, a zero-gated MAC and a zero-row skip unit that the model wrote in Chisel cut energy, and resizing the memories cuts area.
Figures & tables
| metric | it. 1 | it. 15 | change | |
|---|---|---|---|---|
| cycles | M | 106,650 | 50,862 | 2.1 fewer |
| off-chip bytes | M | 3,211,264 | 327,680 | 9.8 less |
| area (mm 2 ) | M+e | 4.035 | 3.115 | 22.8% |
| energy ( J) | m | 95.67 | 17.04 | 5.61 less |
| power (W) | m | 0.449 | 0.168 | 2.68 less |
| performance (GOPS) | M | 4.92 | 10.31 | 2.1 |