We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.
Figures & tables
Figure 1: Overview of the PTXBench benchmark and adaptation workflow.
GPU
Model
GEMM
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
H100
Gemini 3.1 Pro
33.3 / 60.4 / 56.2 (33.3 / 62.5 / 59.4)
33.3 / 39.6 / 45.8 (33.3 / 39.6 / 45.8)
25.0 / 52.1 / 55.2 (25.0 / 52.1 / 55.2)
8.3 / 33.3 / 38.5 (8.3 / 33.3 / 38.5)
– / 22.9 / 25.0 (8.3 / 25.0 / 26.0)
Claude Opus 4.8
91.7 / 95.8 / 94.8 (91.7 / 95.8 / 94.8)
50.0 / 81.2 / 90.6 (66.7 / 89.6 / 94.8)
50.0 / 77.1 / 75.0 (83.3 / 89.6 / 82.3)
8.3 / 62.5 / 79.2 (50.0 / 83.3 / 89.6)
– / 22.9 / 44.8 (25.0 / 41.7 / 59.4)
GLM-5.2
33.3 / 60.4 / 61.5 (33.3 / 62.5 / 62.5)
16.7 / 8.3 / 8.3 (16.7 / 14.6 / 15.6)
– / 8.3 / 10.4 (8.3 / 16.7 / 17.7)
8.3 / 6.2 / 5.2 (8.3 / 18.8 / 16.7)
– (– / 22.9 / 15.6)
Qwen3.6-27B
–
–
–
–
–
GPT-5.6 xhigh
91.7 / 81.2 / 77.1 (91.7 / 83.3 / 79.2)
8.3 / 39.6 / 60.4 (16.7 / 58.3 / 69.8)
– / 50.0 / 56.2 (– / 64.6 / 63.5)
8.3 / 39.6 / 55.2 (58.3 / 60.4 / 65.6)
– / 31.2 / 52.1 (33.3 / 56.2 / 65.6)
B200
Gemini 3.1 Pro
8.3 / 47.9 / 45.8 (8.3 / 47.9 / 45.8)
– / 12.5 / 15.6 (– / 12.5 / 15.6)
– / 8.3 / 10.4 (8.3 / 12.5 / 12.5)
– / 2.1 / 2.1 (– / 2.1 / 2.1)
– (– / 2.1 / 1.0)
Table 1: Target-instruction and unrestricted turn correctness rates (%) on H100 and B200.
Figure 2: FastpInst. on H100 (top) and B200 (bottom).
Model
SWE-bench Pro (%)
Model release
Knowledge cutoff
Lag after PTX release (months)
Hopper
Blackwell
Gemini 3.1 Pro
54.2
Feb. 2026
Jan. 2025
25
≈0
Claude Opus 4.8
69.2
May 2026
Jan. 2026
37
12
GPT-5.6 Sol
64.6
July 2026
Feb. 2026
38
13
GLM-5.2
62.1
June 2026
Not disclosed
≤42
≤17
Qwen3.6-27B
53.5
Apr. 2026
Not disclosed
≤40
≤15
Table 2: Model release dates, knowledge cutoffs, vendor-reported SWE-bench Pro scores, and estimated calendar lag from the release of Hopper PTX ISA 8.0 (Dec. 2022) and Blackwell PTX ISA 8.7 (Jan. 2025) [ 25 , 28 ] . Lags use monthly granularity, from PTX release to disclosed cutoff; otherwise, model release dates provide upper bounds.
Figure 3: Gemini 3.1 Pro Fastp distributions for Triton and CUDA-PTX on H100 and B200.
Method or ratio
GEMM
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
Multi-turn refinement
0.505×
0.584×
0.560×
0.240×
0.166×
Antigravity
0.583×
0.425×
0.413×
0.053×
0.000×
Agent / refinement total tokens
3.32×
5.06×
3.90×
3.73×
3.52×
Table 3: Antigravity versus multi-turn refinement after eight H100 evaluation requests per trajectory. Speedups are medians of four per-trajectory best instruction-qualified values. Token ratios divide the agent’s median cumulative total tokens by the refinement median within each workload.
Table 4: Ablation of architecture-specific prompt knowledge for Gemini 3.1 Pro.
Figure 4: Base versus s1 generalization on the s1 training problem classes and held-out workloads. Top: Qwen3.6-27B before adaptation. Bottom: Qwen3.6-27B-s1, trained with shuffle seed 0. Runtime-instruction rows require dynamically observed Hopper SASS execution.
Figure 5: Six-recipe data ablation. Pie area is proportional to record count, slices show the fraction drawn from each problem, and the bottom line identifies the reasoning synthesizer.
Figure 6: Six-recipe SFT data ablation at eight turns (complete results in Appendix Tables 10 and 11 ).
Figure 7: How Fixit SFT changes reasoning length and error types across turns.
Figure 8: Cross-language transfer of Fixit SFT from CUDA-PTX to Triton on Hopper.
Figure 9: Turn-level error-state transitions for GEMM.
Condition
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
Qwen3.6-27B w/o expert guidance
–
–
–
–
Qwen3.6-27B
–
–
–
–
Qwen3.6-27B-s1 w/o expert guidance
–
– / – / 4.2 (– / – / 4.2)
– / – / 4.2 (– / – / 4.2)
– / 4.2 / 3.1 (– / 4.2 / 3.1)
Qwen3.6-27B-s1
16.7 / 16.7 / 19.8 (16.7 / 16.7 / 19.8)
– / – / 3.1 (– / – / 3.1)
– / 2.1 / 4.2 (– / 2.1 / 4.2)
– / 2.1 / 4.2 (– / 2.1 / 5.2)
Qwen3.6-27B + retrieved repair notes
–
–
–
–
Qwen3.6-27B + retrieved repair notes and fixed kernel
– / 29.2 / 29.2 (– / 29.2 / 29.2)
– / 20.8 / 27.1 (– / 20.8 / 27.1)
– / 14.6 / 18.8 (– / 14.6 / 20.8)
– / 20.8 / 33.3 (– / 25.0 / 37.5)
Table 5: Target-instruction and unrestricted turn correctness rates (%) under SFT and prompt-time supervision.
Figure 10: FastpInst. under SFT and prompt-time supervision.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Mechanism
Indicator
PTX interface
SASS execution criterion
Tensor compute
CGen4
wgmma.mma_async
NHGMMA>0
CGen5
tcgen05.mma , tcgen05.ld
(NUTCHMMA>0)∧(NLDTM>0∨NLDT>0)
Data movement
DTMA
cp.async.bulk.tensor
NUTMALDG>0∨NUTMASTG>0
DTMEM
tcgen05.cp , tcgen05.st
NUTCCP>0∨NSTTM>0∨NSTT>0
Appendix
Table 6: PTX interfaces and SASS execution criteria for the selected BF16 mechanisms.
Figure 11: Detailed execution infrastructure supporting PTXBench. MiniPTXAgent compiles generated CUDA–PTX locally and sends successfully compiled candidates to an isolated GPU profiling service for sanitization, evaluation, and optional diagnostics. A reliability monitor pauses dispatch, restarts unhealthy service state, and excludes affected turns from trajectories. After a restart, it also excludes thermally abnormal GPUs because throttling distorted latency by over 15% in our measurements [ NVIDIA Corporation, 2026f ] .
Figure 12: Workload-state caching (a), rollout concurrency with generation–profiling pipelining (b), and profiling-GPU scaling (c).
GPU
Component
Tokens
H100
Architecture parameter
259
Template functions
18,601
Architecture contract
3,454
Total
22,314
B200
Architecture parameter
413
Template functions
18,871
Appendix
Table 7: Token counts for architecture-specific prompt components.
Figure 13: GPT-5.6 Sol Fastp distributions for Triton and CUDA-PTX on H100 and B200.
Method or ratio
GEMM
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
Multi-turn refinement
0.641×
0.682×
0.805×
0.383×
0.214×
Codex
0.725×
0.565×
0.581×
0.274×
0.116×
Agent / refinement total tokens
5.78×
4.50×
3.47×
5.60×
4.20×
Appendix
Table 8: Codex versus multi-turn refinement after eight H100 evaluation requests per trajectory. Speedups are medians of four per-trajectory best instruction-qualified values; token ratios compare workload-level medians of cumulative total tokens.
Figure 14: Median best instruction-qualified speedup at shared per-trajectory total-token budgets on H100 for Antigravity/Gemini 3.1 Pro (top) and Codex/GPT-5.6 Sol (bottom), each compared with multi-turn refinement using the same model. Each curve includes four trajectories per method and workload; bands show pointwise 95% bootstrap intervals. The horizontal axis is logarithmic.
Figure 15: H100 Fastp and FastpInst. .
Figure 16: B200 Fastp and FastpInst. .
GPU
Model
GEMM
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
H100
Gemini 3.1 Pro
0.687 / 0.934 / 0.962 (0.687 / 0.934 / 0.962)
0.555 / 0.730 / 0.730 (0.555 / 0.730 / 0.730)
0.614 / 0.651 / 0.768 (0.614 / 0.651 / 0.768)
0.206 / 0.375 / 0.515 (0.206 / 0.375 / 0.515)
– / 0.634 / 0.639 (0.065 / 0.634 / 0.639)
Claude Opus 4.8
0.770 / 0.968 / 0.976 (0.770 / 0.968 / 0.976)
0.759 / 0.770 / 0.839 (0.759 / 0.770 / 0.839)
0.758 / 0.806 / 0.806 (0.758 / 0.806 / 0.806)
0.300 / 0.440 / 0.489 (0.300 / 0.440 / 0.489)
– / 0.499 / 0.499 (0.058 / 0.499 / 0.499)
GLM-5.2
0.447 / 0.692 / 0.692 (0.447 / 0.692 / 0.692)
0.407 / 0.470 / 0.607 (0.407 / 0.470 / 0.607)
– / 0.471 / 0.533 (0.015 / 0.471 / 0.533)
0.316 / 0.437 / 0.437 (0.316 / 0.437 / 0.437)
– (– / 0.101 / 0.101)
Qwen3.6-27B
–
–
–
–
–
GPT-5.6 xhigh
0.788 / 0.870 / 0.870 (0.788 / 0.870 / 0.870)
0.282 / 0.774 / 0.872 (0.282 / 0.774 / 0.872)
– / 0.699 / 0.865 (– / 0.699 / 0.865)
0.230 / 0.699 / 0.703 (0.230 / 0.699 / 0.703)
– / 0.650 / 0.746 (0.094 / 0.650 / 0.746)
B200
Gemini 3.1 Pro
0.273 / 0.680 / 0.892 (0.273 / 0.680 / 0.892)
– / 0.280 / 0.280 (– / 0.280 / 0.280)
– / 0.206 / 0.248 (0.013 / 0.206 / 0.248)
– / 0.087 / 0.133 (– / 0.087 / 0.133)
– (– / 0.015 / 0.015)
Appendix
Table 9: Best speedup with target instruction execution on H100 and B200.
Figure 17: Generalization of the 158-record 4ops s1 recipe under additional SFT shuffle seeds. Top: seed 1. Bottom: seed 2. The seed-1 GEMM cell contains 90 of 96 expected turn rows; all other plotted cells contain the complete 96 turns. Runtime-instruction rows require dynamically observed Hopper SASS execution.
SFT-ed Model Label
GEMM
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
Qwen3.6-27B-s1
25.0 / 14.6 / 13.5 (25.0 / 14.6 / 13.5)
16.7 / 16.7 / 19.8 (16.7 / 16.7 / 19.8)
– / – / 3.1 (– / – / 3.1)
– / 2.1 / 4.2 (– / 2.1 / 4.2)
– / 2.1 / 4.2 (– / 2.1 / 5.2)
Qwen3.6-27B-s2
– / 2.1 / 2.1 (– / 2.1 / 2.1)
– / 2.1 / 5.2 (– / 2.1 / 5.2)
– / – / 1.0 (– / – / 1.0)
– / – / 1.0 (– / – / 1.0)
–
Qwen3.6-27B-s3
– / 2.1 / 3.1 (– / 2.1 / 3.1)
– / 4.2 / 2.1 (– / 4.2 / 2.1)
– / 6.2 / 4.2 (– / 6.2 / 4.2)
– / 4.2 / 4.2 (– / 4.2 / 4.2)
–
Qwen3.6-27B-s4
8.3 / 8.3 / 10.4 (8.3 / 8.3 / 10.4)
– / – / 4.2 (– / – / 4.2)
– / 4.2 / 2.1 (– / 4.2 / 2.1)
– / – / 1.0 (– / – / 1.0)
–
Qwen3.6-27B-s5
16.7 / 14.6 / 12.5 (16.7 / 14.6 / 12.5)
– / 2.1 / 4.2 (– / 2.1 / 4.2)
– / 2.1 / 3.1 (– / 2.1 / 3.1)
– / – / 5.2 (– / – / 5.2)
– / – / 4.2 (– / – / 4.2)
Qwen3.6-27B-s6
8.3 / 2.1 / 1.0 (8.3 / 2.1 / 1.0)
–
–
–
–
Appendix
Table 10: Turn correctness rates for the five-problem SFT evaluation.
SFT-ed Model Label
GEMM
MHA-Fwd
MHA-Fwd-Causal
MHA-Bwd
MHA-Bwd-Causal
Qwen3.6-27B-s1
0.303 / 0.303 / 0.340 (0.303 / 0.303 / 0.340)
0.556 / 0.556 / 0.565 (0.556 / 0.556 / 0.565)
– / – / 0.395 (– / – / 0.395)
– / 0.380 / 0.389 (– / 0.380 / 0.389)
– / 0.192 / 0.199 (– / 0.192 / 0.199)
Qwen3.6-27B-s2
– / 0.209 / 0.211 (– / 0.209 / 0.211)
– / 0.548 / 0.548 (– / 0.548 / 0.548)
– / – / 0.315 (– / – / 0.315)
– / – / 0.296 (– / – / 0.296)
–
Qwen3.6-27B-s3
– / 0.274 / 0.280 (– / 0.274 / 0.280)
– / 0.465 / 0.465 (– / 0.465 / 0.465)
– / 0.388 / 0.388 (– / 0.388 / 0.388)
– / 0.494 / 0.494 (– / 0.494 / 0.494)
–
Qwen3.6-27B-s4
0.276 / 0.373 / 0.373 (0.276 / 0.373 / 0.373)
– / – / 0.452 (– / – / 0.452)
– / 0.383 / 0.383 (– / 0.383 / 0.383)
– / – / 0.199 (– / – / 0.199)
–
Qwen3.6-27B-s5
0.325 / 0.446 / 0.446 (0.325 / 0.446 / 0.446)
– / 0.425 / 0.573 (– / 0.425 / 0.573)
– / 0.241 / 0.246 (– / 0.241 / 0.246)
– / – / 0.295 (– / – / 0.295)
– / – / 0.246 (– / – / 0.246)
Qwen3.6-27B-s6
0.073 / 0.073 / 0.073 (0.073 / 0.073 / 0.073)
–
–
–
–
Appendix
Table 11: Best speedups for the five-problem SFT evaluation.
Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present Atrex-Bench, a benchmark whose 30 operators and 440 shapes are sampled directly from full-cluster production inference traces of compute-limited, memory-rich GPUs. Each problem carries an importance weight derived from its share of observed GPU time, weighted by application card-hours and computed separately for the serving phases in which it runs, together with a per-problem roofline ceiling, so the aggregate score emphasizes the kernels that consume the most serving time. Evaluating six frontier coding agents on Atrex-Bench shows that even the best vanilla model reaches only ∼10% of the hardware roofline on production operators; and correctness alone overstates capability, since much of the apparent pass rate comes from PyTorch fallbacks rather than kernels the model wrote. To close this gap, we co-release Atrex-Kernel-Agent (AKA), a profile-driven kernel-optimization agent that combines iterative measure-revise search, optimization dropout for escaping stalled search contexts, and a layered GPU-optimization knowledge base (298 reference-kernel files and 244 optimization-knowledge documents, plus external upstream reference projects for API/ISA lookup). In a controlled case study, the agent converts zero-FlyDSL fallbacks into real kernels that match or exceed hand-tuned production baselines.
LLM-based Triton kernel generation has attracted significant interest, yet a fundamental empirical question remains unanswered: where does this capability break down, and why? We present KernelBenchX, a benchmark designed to answer this question through category-aware evaluation of correctness and hardware efficiency across 176 tasks in 15 categories. Our systematic comparison of five representative methods yields three main findings. First, task structure determines correctness more than method design. Category explains nearly three times more variance in semantic correctness than method (9.4% vs 3.3% explained deviance), and 72% of Fusion tasks fail across all five methods while Math tasks are solved consistently. Second, iterative refinement improves correctness, but not performance. Across GEAK iterations, compile rate rises from 52.3% to 68.8% while average speedup declines from 1.58× to 1.44×; newly rescued kernels consistently underperform persistently correct ones (1.16× vs 1.58× speedup in round~0→1). Third, correctness does not imply efficiency. 46.6% of correct kernels are slower than the PyTorch eager baseline, and cross-hardware speedup variance reaches 21.4×. Besides, quantization remains completely unsolved (0/30 successes) despite non-trivial compilation rates, revealing systematic misunderstanding of numerical computation contracts rather than surface-level syntax errors. These findings suggest that future progress depends on handling global coordination, explicitly modeling numerical precision, and incorporating hardware efficiency into generation. The code is available at https://github.com/BonnieW05/KernelBenchX
Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization, memory access coalescing, tile size tuning, and architecture-specific workarounds -- to every Triton kernel in their code-base. This manual, repetitive effort is a major bottleneck: each kernel demands the same cycle of trial-and-error profiling against hardware constraints that vary across devices, yet the underlying optimization patterns remain largely consistent. We present Xe-Forge, a multi-stage LLM-powered pipeline that automates this process for Intel GPU. Given a functionally correct Triton kernel, the system applies up to nine optimization stages -- from algorithmic restructuring and operator fusion through block pointer modernization, GPU-specific tuning, and open-ended discovery -- each driven by a Chain-of-Verification-and-Refinement (CoVeR) agent that generates candidates, validates them on real hardware, and iterates on failures. A curated knowledge base encodes Intel GPU constraints (power-of-two warp counts, GRF modes, SLM sizing) that are absent from LLM training data, keeping the model within architecturally valid bounds. We evaluate Xe-Forge on 97 Level-2 KernelBench kernels and Flash Attention on the Intel Arc Pro B70, achieving a 1.17x geometric mean speedup over PyTorch eager with 67% of kernels improving, nine kernels exceeding 5x (up to 82x), and 2--13.3x speedups on Flash Attention across all tested configurations without regression -- demonstrating that structured domain knowledge with hardware-in-the-loop verification can systematically eliminate the repetitive porting effort that currently gates algorithm deployment on new accelerators.
Marcin Spoczynski, Daniel Fleischer, Moshe Berchansky +5