Documentation

How Kernova works

A high-level guide to the profile → optimize → verify workflow. Full setup guides are provided to early-access teams.

Workflow

01 · Profile

Capture a representative workload

Run a short training or inference loop under the Kernova profiler. It records operator-level timings, tensor shapes, and memory traffic.

02 · Optimize

Generate candidate kernels

The engine targets the hottest operators, proposes fused or retuned kernels in Triton or CUDA, and autotunes them for your GPU.

03 · Verify

Validate before you ship

Every candidate is checked against the original outputs within configurable numerical tolerances, then benchmarked against the baseline.

04 · Deploy

Drop in the result

Accepted kernels are delivered as source with a validation report, so your team can review, version, and deploy them like any other code.

Supported environments

Frameworks
PyTorch 2.x
Kernel languages
Triton, CUDA C++
Hardware
NVIDIA Ampere and Hopper (A100, H100)
Precision
FP32, BF16, FP16

FAQ

Do I need to change my model code?

No. Profiling attaches to existing workloads, and optimized kernels are integrated where the original operators run.

What happens if a kernel changes my outputs?

It is rejected. Only candidates that pass numerical validation are reported.

Where is the full API reference?

Detailed setup guides and API references are shared with early-access teams during onboarding.

Want the full guides? Request early access.

Supported environments reflect the planned scope. Update them to match what the product supports today.