TIRx Harness: An Open Compiler Harness for Agentic GPU Programming
6 hours ago
- TIRx Harness is a compiler harness that combines a minimal, stable compiler foundation, a knowledge base, tools, and a benchmark server to help AI agents develop correct and fast GPU kernels.
- On evaluated workloads, Kimi Delta Attention (KDA) kernels achieved geometric-mean speedups of 2.94× over FlashKDA (forward) and 6.84× over Flash Linear Attention (FLA) (backward).
- Key components include a hardware-close TIRx Foundation for predictable compilation, synchronization and data-race analysis tools, a kernel zoo with over 60 reusable implementations, and a remote benchmark server (KCoral) for controlled evaluation.
- The harness enables agents to discover new optimization directions from PTX documentation, transfer strategies across kernels via the zoo, and diagnose unsafe candidates with detailed feedback beyond pass/fail.
- Evaluation across KDA, MSA, MLA, and VSA families showed family-level geometric-mean speedups ranging from 1.33× to 6.84×, with kernels competitive against optimized reference implementations.
- The system features a self-improvement loop where successful kernels and tool fixes become part of the environment for future runs, and future work includes adapting to new architectures, pushing further toward hardware, and extending correctness tooling for megakernels.