the.ai

Compilers / Methods

verified

Custom Kernel

Sometimes the compiler will not produce what you need and you write the kernel yourself. Languages like Triton put this within reach of someone who is not a CUDA specialist — you write in terms of blocks and the compiler handles the thread-level detail — which is why the number of hand-written kernels in open models rose rather than fell.

Viz primitive · budget-splithand-written = 4

hand-written holds 13% of the budget; rest holds the remaining 87%.

Operations covered by hand-written kernels against those left to the compiler, in operations. Drag the hand-written share up to watch performance come under your control — and the maintenance with it.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:1802.04799 — TVM: An Automated End-to-End Optimizing Compiler for Deep Learning

Origin · not linkable

  • Tillet et al. 2019 — Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations · MAPL 2019 · doi:10.1145/3315508.3329973