Attention variants used in production rarely match the shapes that public kernels were tuned for. Sequence lengths, head dimensions, and masking patterns all shift the optimal tile configuration.
Framing the search
We treat each kernel as a parameterized template: block sizes, number of warps, pipeline stages, and load ordering. The search space is large, but most points are invalid on a given GPU because of shared-memory or register limits, so we prune aggressively before benchmarking anything.
Where time actually goes
On the workloads we profiled, the largest improvements came from reducing redundant global-memory reads, not from changing arithmetic. A tile that keeps K and V resident across more query blocks often beats a tile with higher theoretical occupancy.
Verification first
Every candidate is checked for numerical parity against the reference implementation across multiple shapes before its timing is recorded. A fast kernel that diverges is discarded, not reported.
Numbers in this post are illustrative. We publish methodology alongside any benchmark we share with customers.