August 12, 2025 · 9 min read

Searching the Triton tile space for attention variants

How we frame kernel tuning as a constrained search problem, and why most of the wins come from memory layout rather than math.

Attention variants used in production rarely match the shapes that public kernels were tuned for. Sequence lengths, head dimensions, and masking patterns all shift the optimal tile configuration.

Framing the search

We treat each kernel as a parameterized template: block sizes, number of warps, pipeline stages, and load ordering. The search space is large, but most points are invalid on a given GPU because of shared-memory or register limits, so we prune aggressively before benchmarking anything.

Where time actually goes

On the workloads we profiled, the largest improvements came from reducing redundant global-memory reads, not from changing arithmetic. A tile that keeps K and V resident across more query blocks often beats a tile with higher theoretical occupancy.

Verification first

Every candidate is checked for numerical parity against the reference implementation across multiple shapes before its timing is recorded. A fast kernel that diverges is discarded, not reported.

Numbers in this post are illustrative. We publish methodology alongside any benchmark we share with customers.

← All engineering notes