Compiling Triton kernels without the Triton compiler
2 days ago
- Compiler backends are costly to build and maintain due to evolving programming models, workloads, and accelerators.
- The paper investigates replacing conventional compiler lowering with large language models (LLMs), a process called AI lowering.
- AI lowering is demonstrated by translating Triton kernels directly into NVIDIA PTX using an LLM agent and an evaluation harness.
- Tested on 12 common kernels across Ada, Hopper, and Blackwell GPUs, and 10 kernels from recent ML papers, AI lowering achieves 0.83x to 3.34x the performance of autotuned Triton.
- Performance gains come from transformations not in Triton's pipeline, such as decoding packed binary weights into Tensor Core operands (3.34x on BitDelta) and reusing overlapping convolution windows (up to 2.23x).
- The approach extends the Volta PTX verifier to support modern GPUs, including Blackwell's Tensor Core interface with managed memory and asynchronous execution.
- The findings suggest a future where AI compilers replace custom intermediate representations and checkers, reducing engineering effort for new hardware.