cs.LGAug 20, 2026

Learning Exact NVIDIA SASS Encoders with F2\mathbb{F}_2 Linear Algebra

Authors: Jiading Gai

Organizations: Independent Researcher.

Abstract

NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. To our knowledge, F2Asm is the first system to learn SASS instruction encoders as vector-valued affine maps over F2\mathbb{F}_2 and the first open-source NVIDIA SASS assembler to support Rubin SM107. F2Asm uses Gaussian elimination over F2\mathbb{F}_2 to incrementally build a compact basis, detect inconsistencies, and reject inputs outside the learned span. F2Asm separates target-specific control bits, relocation rules, and CUBIN metadata from its learning algorithm. We train encoders for Hopper SM90/SM90a, Blackwell SM100, and Rubin SM107 using 3,225 CUBINs from pinned NVIDIA and third-party production libraries, CUDA 13.3 packages, and CUDA 13.4 Developer Preview archives. In round-trip tests, F2Asm reassembles each CUBIN's disassembled SASS, and all compared executable text sections match the originals exactly. Joint training with F2Asm yields one shared encoder for five Blackwell SM targets and another for three Rubin SM targets, providing strong evidence of a common SASS encoding scheme for instructions shared within each family. Continual training extends the Rubin encoder to all 17,159 previously unsupported cuTile and GROMACS queries with 1,504 additional basis rows, matching the derived lower bound.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching

    May 7, 2026Gabriel Bathie, Baptiste Mouillon, Nathanaël FijalkowSynthesisSynthesizer

  2. FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing

    Jun 8, 2026Yuhan Ma, Yong Li, Stefan SchmidLLM Inference OptimizationFixed-Point Iteration

  3. ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants

    Apr 16, 2026Haohui Mai, Xiaoyan Guo, Xiangyun Ding +7Graphics Processing Unit KernelsAgentic Optimization