Engineering notes
What we're learning about GPU performance.
Technical writing from the Kernova AI team on GPU kernels, profiling, and numerical validation.
August 12, 2025 · 9 min read
Searching the Triton tile space for attention variants
How we frame kernel tuning as a constrained search problem, and why most of the wins come from memory layout rather than math.
June 3, 2025 · 7 min read
Why memory bandwidth, not compute, bounds long-context MoE inference
A roofline walk-through of mixture-of-experts decoding, and what it implies for which kernels are worth optimizing.
March 18, 2025 · 6 min read
Checking numerical equivalence without golden outputs
How we validate optimized kernels against references when floating-point reordering makes bit-exact comparison impossible.