Company

We make GPU compute go further.

Kernova AI is building an optimization engine that profiles AI workloads, generates faster GPU kernels, and verifies they produce the same results — so teams can ship more capability on the hardware they already have.

Our story

Kernova started with a frustration every AI infrastructure engineer knows: a training run that should take days takes weeks, and the profiler shows the GPU idle half the time — waiting on memory, on kernels that were compiled for a generic shape, on operators nobody had time to tune.

The usual answer is to hire a kernel specialist and wait. Hand-tuned CUDA or Triton kernels can unlock dramatic speedups, but the expertise is scarce, the work is slow, and every model change starts the cycle over. Most teams simply accept the waste and pay for more hardware instead.

We asked a different question: what if the search for a faster kernel could be automated — and, just as importantly, verified? A faster kernel is only useful if it produces the same results. So Kernova pairs AI-driven kernel generation with rigorous numerical validation, so every optimization ships with evidence, not hope.

Today we work with early-access teams running training and inference at scale, and we hold ourselves to a simple standard: every number we report comes with its methodology, its hardware, and its baseline. That is what "more intelligence, less compute" means to us.

Why we exist

GPU capacity is the most expensive line item for most AI teams, yet a large share of it is lost to generic kernels that were never tuned for a specific model, shape, or chip. Hand-tuning closes that gap, but it takes specialist engineers weeks per operator. Our goal is to make that level of optimization routine, repeatable, and safe.

How we work

Measure before you optimize

Every change starts from an operator-level profile of a real workload, not from assumptions about where time is spent.

Correctness is non-negotiable

A faster kernel that changes model outputs is a regression. Numerical validation runs before any result is reported.

Show the evidence

We report methodology, hardware, and baselines alongside every number, and we label illustrative data as such.

Fit into existing stacks

Teams already run PyTorch, CUDA, and Triton. Kernova works with that toolchain instead of asking you to replace it.

Our journey so far

  1. 2024

    Research prototype for automated kernel search on NVIDIA GPUs.

  2. Nov 2024

    First early-access cohort begins profiling real workloads.

  3. 2025

    Validation pipeline and hardware-aware autotuning added to the platform.

  4. Today

    Expanding early access to teams running training and inference at scale.

Advisors & backers

Kernova is supported by angel investors and advisors with backgrounds in GPU systems, compilers, and large-scale ML infrastructure.

Placeholder

Advisor — GPU systems

Former GPU architecture engineer; advises on kernel performance modeling.

Placeholder

Advisor — ML compilers

Compiler engineer with deep experience in deep-learning graph optimization.

Placeholder

Angel — AI infrastructure

Operator who scaled inference platforms serving production LLM traffic.

Get in touch

Product

Early access

Company

Press & partnerships

The story, timeline, and advisor entries are illustrative. Replace them with your verified company history and real people before launch or any external submission.