Rebuilding Triton and Helion from Scratch in 4,000 Lines of Python
Triton and Helion are how most custom GPU kernels get written today. Triton, from OpenAI, lets you write a kernel as a Python function over a block of data and handles the thread mapping, memory coale