Hasty Briefsbeta

Bilingual

Compiling Triton kernels without the Triton compiler

2 days ago
  • Compiler backends are costly to build and maintain due to evolving programming models, workloads, and accelerators.
  • The paper investigates replacing conventional compiler lowering with large language models (LLMs), a process called AI lowering.
  • AI lowering is demonstrated by translating Triton kernels directly into NVIDIA PTX using an LLM agent and an evaluation harness.
  • Tested on 12 common kernels across Ada, Hopper, and Blackwell GPUs, and 10 kernels from recent ML papers, AI lowering achieves 0.83x to 3.34x the performance of autotuned Triton.
  • Performance gains come from transformations not in Triton's pipeline, such as decoding packed binary weights into Tensor Core operands (3.34x on BitDelta) and reusing overlapping convolution windows (up to 2.23x).
  • The approach extends the Volta PTX verifier to support modern GPUs, including Blackwell's Tensor Core interface with managed memory and asynchronous execution.
  • The findings suggest a future where AI compilers replace custom intermediate representations and checkers, reducing engineering effort for new hardware.