Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lowering achieves 0.83x-3.34x the performance of autotuned Triton. The largest gains come from transformations that Triton's lowering pipeline does not perform, such as decoding packed binary weights directly into Tensor Core operands (3.34x on BitDelta), assigning each thread a complete softmax row in tensor memory (1.37x on FlashAttention), and reusing overlapping convolution windows (up to 2.23x). These results rely on a robust evaluation harness with comprehensive verification support. We build on Volta, an existing PTX verifier, and substantially extend it to support modern GPU architectures by introducing support for Blackwell's tcgen05 Tensor Core interface. This requires modeling three architectural features: managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated through commits, waits, memory barriers, and proxy fences. We discuss the challenges involved in formalizing them, as well as the current limitations. Our results suggest an emerging future in which AI compilers replace custom-written intermediate representations and checkers, reducing the time and engineering effort required to bring up software for new general-purpose and custom chips.
Figures & tables
Figure 1: Traditional compilation, superoptimization, and AI lowering. (a) A traditional compiler applies hand-written lowering and optimization passes across intermediate representations. (b) A superoptimizer uses conventionally lowered target code as a seed, then iteratively mutates and scores candidate programs. (c) AI lowering bypasses the conventional optimizing/lowering backend: an LLM generates PTX directly and revises it using numerical and performance feedback, and optionally symbolic verification. In the search-based approaches, acceptance is guarded by numerical or formal checks.
Figure 2: System overview and detailed AI-compilation loop. Top: conventional compilation and AI lowering start from the same compilation contract, produce an interface-compatible PTX module, and share the NVIDIA execution path. The orange block expands TAIC into the detailed loop below, where TAIC proposes and revises PTX while TCEnv assembles, tests, benchmarks, profiles, and selects candidates.
Ada
used?
Hopper
used?
Blackwell
used?
MMA
mma.sync + ldmatrix
✓
wgmma + TMA loads
✓
tcgen05 + TMEM accum.
✓
SMEM
XOR bank swizzle
✓
128-B XOR swizzle
✓
64-B offset swizzle
✓
Copies
cp.async pipeline
✓
TMA + mbarrier
✓
TMA 2-D + phase parity
✓
Figure 3: Architecture-specific features and their observed use in LLM-generated PTX. A check mark denotes that the generated kernels use the feature on the corresponding microarchitecture. The Hopper and Blackwell entries ( wgmma , TMA, mbarrier , tcgen05 , TMEM) fall outside Volta’s originally validated semantics; verifying kernels that use them relies on the emulator extension of Section 4.5 .
Figure 4: Per-kernel speedup/cost Pareto frontiers relative to Triton for FP16 and FP8 on three GPU architectures. Each point is a validated candidate that improves the best speedup reached at its cumulative API cost; the dashed line marks parity with the corresponding autotuned Triton baseline. Axes are scaled independently across panels.
Kernel
Budget
Trial 1
Trial 2
Trial 3
Trial 4
Trial 5
Mean ± SD
Best
Convolution (FP8)
$15
2.081×
1.706×
2.229×
1.933×
2.070×
2.004±0.197
2.229×
Matrix multiplication (FP16)
$15
1.002×
1.026×
1.032×
1.029×
1.024×
1.023±0.012
1.032×
SwiGLU (FP16)
$1
1.000×
1.002×
1.000×
1.000×
1.000×
1.000±0.001
1.002×
FlashAttention forward (FP16)
$15
1.261×
1.075×
1.214×
1.263×
1.366×
1.236±0.106
1.366×
Table 1: Repeatability on NVIDIA B200 across five independent Triton AI Compiler runs. Values are speedups over the same autotuned Triton baseline. Budget is the configured per-run API-cost gate, checked before each new model request; an in-flight request can finish above it.
Figure 5: Best numerically validated LLM-generated kernels on NVIDIA B200. Each bar reports the best accepted candidate for one kernel; the dashed line marks parity with Triton.
Kernel
Provably Verified
Numerically Verified
ReLU
1.00×
1.00×
GEMV
1.15×
1.15×
GEMM
0.24×
1.03×
FlashAttention
0.78×
1.37×
Convolution 2D
0.77×
1.35×
Table 2: Performance of the best provably verified and best numerically verified implementations on an NVIDIA B200 in FP16 using the original Volta.