Intelligence at the kernel level

More intelligence.
Less compute.

AI-powered GPU optimization. Turn complex workloads into faster, leaner, verified code.

Profile. Optimize. Validate.

Optimization preview

Validated

Baseline
Optimized

Illustrative workload — demo data

Built around your stack

PyTorchCUDATritonNVIDIA GPUs

The optimization engine

From bottleneck to breakthrough.

01

Profile the workload

Automatically analyze your PyTorch workloads, find bottlenecks, and understand what matters.

Workload trace (PyTorch)

ComputeMemoryOther

aten.matmul

aten.layer_norm

aten.softmax

aten.add

Demo data

02

Generate better kernels

Use AI to design and generate optimized GPU kernels, tailored to your hardware and workload patterns.

Generated kernel (Triton)

Auto-optimized
1@triton.jit
2def fused_attention(Q, K, V, stride_q, stride_k, stride_v):
3 pid = tl.program_id(0)
4 offs = pid * BLOCK_M + tl.arange(0, BLOCK_M)
5 q = tl.load(Q + offs[:, None] * stride_q)
6 k = tl.load(K + offs[None, :] * stride_k)
7 v = tl.load(V + offs[:, None] * stride_v)
8 # ... optimized attention kernel ...

Illustrative example

03

Verify every change

Automatically validate correctness and measure performance across real workloads and hardware.

Validation results

Validated
  • Numerical correctnessPass
  • Output parity (tolerance)Pass
  • Performance regression checkPass
  • Multi-shape validationPass
  • Multi-GPU consistencyPass

Demo data

Technology

Kernels written for your hardware, not against it.

Kernova analyzes the execution trace of your model, identifies the operators that dominate runtime, and generates fused GPU kernels tuned to your exact shapes, dtypes, and memory layout. Every generated kernel is checked against the reference implementation before it ever reaches your training loop.

Trace-driven
Decisions come from measured workload behavior, not heuristics.
Hardware-aware
Tuned for the specific GPU architecture you deploy on.
Verified output
Numerical parity checks gate every optimization.

How a run works

  1. 1

    Connect

    Point Kernova at your PyTorch training or inference script.

  2. 2

    Profile

    Capture operator-level traces across representative batches.

  3. 3

    Optimize

    Generate and benchmark candidate kernels for hot operators.

  4. 4

    Validate

    Verify numerical parity and ship only what passes.

Benchmarks

Measured, not promised.

Speedups below come from an illustrative internal workload and are shown to explain the methodology — they are demo data, not published results.

Demo benchmark results for illustrative kernels
KernelBaselineOptimizedMemory
fused_attention1.00×2.41×-38%
layernorm_gemm1.00×1.87×-22%
softmax_reduce1.00×1.64×-17%

Demo data — illustrative workload, not a published benchmark.

Early access feedback

What early teams are saying.

Illustrative quotes shown for layout purposes — attributed by role only, pending permission to publish named testimonials.

“The trace view alone changed how we think about our training loop. We finally saw which operators actually mattered instead of guessing.”

ML Infrastructure Lead

Series C AI startup

“What sold us was the validation step. Generated kernels we couldn't trust would have been useless — these came with proof.”

Principal Engineer

Inference platform team

“We expected a research prototype. Instead it slotted into our PyTorch workflow in an afternoon and the reports went straight into code review.”

Head of Compute

Foundation model lab

Demo content — not real customer statements.

Developers

Fits into the workflow you already have.

Kernova works alongside your existing PyTorch codebase. No rewrites, no new framework — profiling and optimization run as a step in your development loop, and validated kernels drop in as standard Triton modules.

  • Works with stock PyTorch — no fork required
  • Generated kernels are readable, reviewable Triton code
  • Validation reports your team can audit before merging

Integration sketch

Conceptual
# conceptual example — final API may differ
import kernova as kn

trace = kn.profile(model, sample_batch)
plan = kn.optimize(trace, target="h100")
report = kn.validate(plan, tolerance=1e-4)

assert report.all_passed

Platform updates

A year of steady progress.

  1. Sep 2025v0.9

    Multi-GPU consistency checks

    Validation now covers multi-GPU runs, catching numerical drift across devices before kernels ship.

  2. Jun 2025v0.8

    Hardware-aware autotuning for H100

    Kernel generation now tunes block sizes and memory layout against the target GPU architecture.

  3. Mar 2025v0.7

    Operator-level PyTorch tracing

    Profiling moved from whole-model summaries to per-operator traces with compute/memory breakdown.

  4. Nov 2024v0.6

    First early-access cohort

    Opened the platform to a small group of teams running production PyTorch workloads.

Illustrative release history — demo content.

Early access

Put your workloads on a diet.

We're onboarding teams with heavy GPU workloads in small batches. Tell us about your stack and we'll be in touch.