Learning Exact NVIDIA SASS Encoders with Linear Algebra
Organizations: Independent Researcher.
Abstract
NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. To our knowledge, F2Asm is the first system to learn SASS instruction encoders as vector-valued affine maps over and the first open-source NVIDIA SASS assembler to support Rubin SM107. F2Asm uses Gaussian elimination over to incrementally build a compact basis, detect inconsistencies, and reject inputs outside the learned span. F2Asm separates target-specific control bits, relocation rules, and CUBIN metadata from its learning algorithm. We train encoders for Hopper SM90/SM90a, Blackwell SM100, and Rubin SM107 using 3,225 CUBINs from pinned NVIDIA and third-party production libraries, CUDA 13.3 packages, and CUDA 13.4 Developer Preview archives. In round-trip tests, F2Asm reassembles each CUBIN's disassembled SASS, and all compared executable text sections match the originals exactly. Joint training with F2Asm yields one shared encoder for five Blackwell SM targets and another for three Rubin SM targets, providing strong evidence of a common SASS encoding scheme for instructions shared within each family. Continual training extends the Rubin encoder to all 17,159 previously unsupported cuTile and GROMACS queries with 1,504 additional basis rows, matching the derived lower bound.
Figures & tables
| Target | GPU architecture | Corpus sources (CUBINs) | Total | Raw rows | Exact-context basis rows | Exact contexts | Skips a |
|---|---|---|---|---|---|---|---|
| SM90/90a | Hopper | NVIDIA TensorRT-LLM (1,108); NVIDIA Transformer Engine (60); vLLM (239); FlashAttention (72); SGLang (60); xFormers (8); bitsandbytes (3); rVLLM (1) | 1,551 | 169.87M | 440,308 | 60,412 | 200 |
| SM100 | Blackwell | NVIDIA cuBLAS (378); cuDNN (283); cuSOLVER (194); NPP (182); cuSPARSE (100); nvJPEG (10); cuRAND (7); cuSPARSELt (3) | 1,157 | 25.09M | 385,141 | 59,894 | 0 |
| SM107 | Rubin | NVIDIA cuSOLVER (203); NPP (197); cuSPARSE (100); nvJPEG (10); cuRAND (7) | 517 | 17.98M | 238,190 | 37,277 | 13,524 |
| Total | 3,225 | 212,937,948 | 1,063,639 | 157,583 | 13,724 | ||
| Target | groups | Affine maps | Final basis rows |
|---|---|---|---|
| SM90/90a | 2,865 | 2,936 | 50,099 |
| SM100 | 3,669 | 3,673 | 60,159 |
| SM107 | 2,896 | 2,896 | 44,882 |
| Total | 9,430 | 9,505 | 155,140 |
| Target | Forms | CUBINs | Sections | Text MiB | P/F/E |
|---|---|---|---|---|---|
| SM90/90a | 389 | 1,551 | 87,141 | 2,592.0 | 1,551/0/0 |
| SM100 | 326 | 1,157 | 30,491 | 382.8 | 1,157/0/0 |
| SM107 | 279 | 517 | 32,676 | 274.5 | 517/0/0 |
| Partition | CUBINs | Instructions | Unique queries | Text sections | P/F/E |
|---|---|---|---|---|---|
| Training | 14,970 | 71,927,724 | 4,412,203 | 59,921 | 14,970/0/0 |
| Held-out | 3,608 | 12,788,800 | 740,247 | 9,017 | 3,608/0/0 |
| Source | Obs. | Queries | Forms | Maps | |
|---|---|---|---|---|---|
| cuTile | 3,702 | 3,421 | 21 | 33 | 104 |
| GROMACS | 53,134 | 13,738 | 63 | 173 | 1,400 |
| Combined | 56,836 | 17,159 | 75 | 206 | 1,504 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Target(s) | Architecture | Status | Opcodes | Forms |
|---|---|---|---|---|
| sm_90/90a | Hopper | Supported | 140 | 389 |
| sm_100/100f/100a sm_103/103a | Blackwell | Supported | 153 | 388 |
| sm_107/107f/107a | Rubin | Supported | 153 | 369 |
| Active features in | (128-bit hexadecimal) | |
|---|---|---|
| 259 | 00000000000000040000000004040000 | |
| 199 | 00000000000000000000004000000000 | |
| 196 | 00000000000000000000000800000000 | |
| 135 | 00000000000000000000000040000000 | |
| 3 | 0400000000041800000000000000723c |