cs.AROct 4, 2026

SparseCraft: Agentic Hardware-Software Co-Optimization for Sparse Computing

Authors: Rajatabha Chakraborty, M P Samartha, Vedant Pahariya, Priyesh Shukla

Organizations: International Institute of Information Technology, Hyderabad, India

Abstract

Sparse-accelerator design spaces are usually searched against analytical models, so a design point is admitted on what a model predicts rather than on what the hardware does. SparseCraft closes that gap with a language model inside a closed CHIA loop. In each of 15 iterations the model reads the measured outcome of the previous one and edits the Chisel RTL, the memory configuration and the sparse-kernel schedule of a Gemmini accelerator through MCP tool servers, and no candidate counts until it has been checked for legality, elaborated, simulated cycle-accurately, checked bit-for-bit on every output against a golden reference, and synthesised. The harness turns each measurement into the next work order, a diagnosed bottleneck with matching strategy guidance, the history of tried designs and a score of the model's own prediction, and a second model repairs changes that fail a gate. On a 512×512512 \times 512 GraphChallenge sparse-DNN layer the loop reaches 2.1x fewer cycles, 9.8x less off-chip traffic and 22.8% less area than the block-sparse Gemmini baseline, with 5.61x higher modelled perf/W and 11.8x lower EDP. The levers span three layers: a schedule that keeps the dense operand resident removes 9.8x of the traffic, a zero-gated MAC and a zero-row skip unit that the model wrote in Chisel cut energy, and resizing the memories cuts area.

Figures & tables

Explore similar work

CardsList
  1. SparseDitto: An Agentic Sparse Compilation Framework through Architecture-Aware Synthesis on GPUs

    Aug 5, 2026Shiyang Li, Guangyan Sun, Jinwei Tang +3Graphics Processing Unit KernelsCuda

  2. Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators

    Jul 24, 2026Afzal Ahmad, Gaoyu Mao, Shoubo Hu +4Fast InferenceSparsity