giveen/KernelOPT

Dispatch-aware agentic GPU kernel optimization

Rust

0

83 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

GitHub - giveen/KernelOPT: Dispatch-aware agentic GPU kernel optimization (r/LocalLLaMA)

I want to share something I've been working on, and research paper from Redhat really helped. This is my agentic GPU kernel optimizer for inference engine development Its whole goal is to look at inference engine kernels, and using cloud models, plan, test and find improvements. Nothing is changed…

2

Oct 2, 2026

README

KernelOPT

Dispatch-aware, agentic GPU-kernel optimization. Point KernelOPT at a single kernel, a directory of kernels, or a whole model — Triton kernels in a compiled PyTorch model, a CUDA op in ninfer, a ggml-cuda kernel in llama.cpp, or any CMake-based GPU repo (CUDA or HIP/ROCm) that declares its build/test/bench in kernelopt.toml. It profiles them and drives a profile → plan → edit → verify → gate loop until each is faster — or a budget runs out. Your working tree is never touched: edits happen in an isolated git worktree, every accepted candidate is committed and tagged (so it can be read or reverted instantly), and the winner is handed back as a unified diff.

An implementation of KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization (arXiv:2609.30059), extended from its Triton/Inductor setting to CUDA kernel trees (ninfer, llama.cpp) and a generic descriptor-driven backend (CUDA and HIP/ROCm), with engine-share ranking via the NVIDIA suite (nsys/ncu) or Graphsignal (CUDA or ROCm).

ModeTarget treeOptimization unitCorrectnessTiming
tritoncompiled PyTorch modelone Inductor @triton.jit kerneleager PyTorch + allclosetriton.testing.do_bench
ninferninfer checkoutone src/ops/<family>/<impl>/*.cuthe Op's own ctest suiteninfer_<op>_bench
llamacppllama.cpp checkoutone ggml/src/ggml-cuda/*.cutest-backend-ops test -o <OP>test-backend-ops perf -o <OP>
customany CMake GPU repo (CUDA or HIP/ROCm) with kernelopt.tomla file you declareyour test_cmdyour bench_cmd

How it works

discover ─▶ baseline ─▶ ┌ plan ─▶ edit ─▶ compile ─▶ correctness ─▶ bench ┐ ─▶ γ/roofline gate ─▶ shape gate ─▶ diff
                        └────────── iterate until patience/target/budget ──┘

LLM agents (Planner → Executor → Summarizer) work with an NCU/Graphsignal-guided UCB-guided beam search and shared experience memory. Five gates filter candidates: static compile, multi-seed correctness, model-level (engine) verification, a performance gate (pinned shape, interleaved fresh baseline/candidate rounds, a noise floor and a sign test), and measured-shape correctness — the op's own suite may not cover the launch path the perf gate measures. A perf win must also be physically plausible: a candidate that appears to beat the device's memory roofline is rejected as "doing less work", not accepted as a win. If nothing passes, the baseline is preserved.

The Planner can also pull context on demand (search_repo/read_file, ripgrep) instead of receiving a full file dump, and every accepted candidate (compile + correctness + bench) is committed and tagged for fast revert. The Executor edits by whole file (default) or by unified diff (--edit-mode patch, opt-in) to avoid reproduction drift; either way the run outputs a reviewable unified diff and your checkout is never modified.

Measurement rigor: an op-level speedup is measured on a pinned shape with 3 repeats, then re-confirmed with interleaved fresh baseline/candidate re-benches (--final-rounds). A candidate is only called optimized if it beats the baseline by more than the measured noise, wins every round, is physically plausible (below a generous multiple of the memory roofline), and is correct on the measured shape; otherwise it is matched (correct but not faster) or fallback (rejected). Correctness-only targets (no bench) report unverified: a candidate was applied, but no speedup could be measured. The reported number is always explicit (N× = baseline/final, with the shape, ms values, noise, and bandwidth-vs-roofline).

One GPU job at a time: bench/verify/NCU/engine-E2E take an exclusive cross-process lock (.kernelopt/gpu.lock), so a campaign and a single run — or two campaigns — serialize on the device instead of corrupting each other's timing or VRAM.

Requirements

  • An NVIDIA GPU (or AMD with ROCm) + recent driver.
  • CUDA toolkit (nvcc, nsys, ncu) — or ROCm (hipcc).
  • Rust stable, Python 3.10+, cmake ≥ 3.28, ninja/make, git.
  • A target checkout (ninfer, llama.cpp, or your own repo) that builds.
  • An LLM key + a model with tool calls.

Full ordered checklist: docs/prereqs.md — driver, toolkit, profiling permission, build tools, LLM key, verify.

Build

cargo build --release        # binary: target/release/kernelopt

Setup

New machine? Start with the ordered checklist: docs/prereqs.md (GPU driver → CUDA toolkit → profiling permission → build tools → target checkout → LLM key + tool-call model → verify). Then:

cp .env.example .env     # set NINFER_REPO / LLAMACPP_REPO / KERNELOPT_MODEL / OPENCODE_API_KEY
kernelopt setup          # full prereq table
kernelopt setup --smoke  # prove the compiler works (seconds, no GPU needed)

.env is read automatically (real env wins), so --repo/--provider/--model are optional. If you pass --repo "$NINFER_REPO", export it first (set -a; source .env; set +a) — otherwise the shell expands it to empty.

Where the LLM should run

KernelOPT talks to the optimizer LLM over HTTP, so the model itself needs no local GPU — and that is the recommended setup: point it at a cloud provider (opencode-go default, openai, openrouter) or an LLM server on another machine, and keep this machine's GPU entirely for compiling, benchmarking, and profiling your model.

Serving the LLM on the same GPU (ollama, vllm, lmstudio locally) works but contends for VRAM with the target build, the benchmarks, and especially Gate 3 (engine E2E loads a whole model — tens of GB). Expect slower runs and possible OOMs; if you do it anyway, watch nvidia-smi and prefer smaller e2e models.

Documentation lookup (CUDA/HIP API docs)

When a candidate fails to compile, KernelOPT looks the offending symbol up and attaches it to the Executor's retry:

kernelopt docs cub::WarpMergeSort     # or: __reduce_max_sync, hipMalloc, rocwmma::…
kernelopt docs --login                # one-time NVIDIA login (token auto-refreshed)

Optional; falls back to local headers. Full detail: docs/docs-lookup.md.

Quickstart

First time? Use the wizard

kernelopt wizard            # detects repo/backend/targets, asks a few questions, runs
kernelopt wizard --dry-run  # just show the command it would run
kernelopt wizard --save     # also persist the choices to .env for later runs
kernelopt wizard --yes      # accept detected defaults, no prompts (scripts/CI)

It reads your .env, probes the optimizer LLM (auth + a forced tool call), lists the discovered kernels and local models, offers quick | standard | thorough presets, then prints and runs the exact command — so you can graduate to the flags once you know what you want.

One kernel

# Compiled PyTorch model (Triton kernels)
kernelopt run examples/mlp.py
kernelopt run examples/pointwise_fused.py

# ninfer Op — smallest kernel, good first run. --watch = one-terminal live view
kernelopt run-ninfer --op add_bias --iterations 3 --beam 2 --retries 3 --watch

# llama.cpp ggml-cuda kernel
kernelopt run-llamacpp --op SOFT_MAX --iterations 3 --beam 2 --watch

Every run prints a run id and writes report.md + report.diff under .kernelopt/runs/<run_id>/. --watch shows progress in the same terminal and prints those paths when the run finishes.

Go further

docs/advanced.md — harder measurements, engine verification, custom engines, campaigns, profiling, run control, outputs, testing, troubleshooting, repo layout.

Documentation

DocContents
docs/advanced.mdBeyond the basics: harder measurements, engine verification, campaigns, profiling, run control, troubleshooting
docs/cli.mdEvery command and option: what discover, profile, campaign, run* do, the full flag reference, and campaign option effects
docs/configuration.md.env, providers, thinking level, config.toml, hyperparameters
docs/prereqs.mdOrdered setup checklist: driver, CUDA toolkit, profiling permission, build tools, LLM key, verify
docs/outputs.mdArtifacts, run/campaign state, applying diffs, disk
docs/campaign.mdCampaign walkthrough: discovery, ordering, budgets, resume
docs/monitoring.mdLive progress, what the LLM is doing, pausing (Ctrl-C), watch, campaign status
docs/kernel-editing.mdHow kernels are read, edited (full-file), isolated, and surfaced as a diff
docs/docs-lookup.mdCUDA/HIP API docs lookup: NVIDIA cuda-docs MCP login, ROCm headers, kernelopt docs
docs/testing-kernel-wins.mdRST recommendations: oracles, benchmark controls, measured-shape coverage and kernel versus application wins
docs/model-e2e.mdModel-level (engine) verification: same tokens, not slower
docs/ninfer-mode.mdninfer mapping (gates, workbench, prompts)
docs/llamacpp-mode.mdllama.cpp mapping (test-backend-ops gates)
docs/custom-mode.mdcustom mode: optimize any CUDA repo via kernelopt.toml
docs/graphsignal.mdEngine-share profiling, managed install, attribution rules

Outputs

Run artifacts live in .kernelopt/runs/<run_id>/ (journal.jsonl, report.md, report.diff, candidate/bench/ncu dumps); campaigns in .kernelopt/campaigns/<id>/. Accepted candidates are committed and tagged in the worktree (kernelopt/<run_id>/<candidate>); nothing is committed to your repo.

Apply the winner with git apply, inspect with kernelopt history, or revert with kernelopt revert — details in docs/advanced.md.

License

Apache License 2.0 — see LICENSE.

Citation

This project implements the method and adapts the agent prompts from KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization (Poddar, Prasad, Samanta, Chakraborty, Goyal, Rathaur; arXiv:2609.30059).

@misc{poddar2026kernelopt,
  title        = {KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization},
  author       = {Poddar, Aheli and Prasad, Sanskar and Samanta, Arindam and
                  Chakraborty, Subha and Goyal, Vishal and Rathaur, Rohit Singh},
  year         = {2026},
  eprint       = {2609.30059},
  archivePrefix= {arXiv},
  primaryClass = {cs.DC},
  url          = {https://arxiv.org/abs/2609.30059}
}

See CITATION.cff and NOTICE. The Planner, Executor, and Summarizer prompts under prompts/ are adapted from the paper's appendix (arXiv:2609.30059). Graphsignal: https://github.com/graphsignal/graphsignal (fork adds llama.cpp/NInfer).

giveen/KernelOPT

Dispatch-aware agentic GPU kernel optimization

Rust

0

83 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

GitHub - giveen/KernelOPT: Dispatch-aware agentic GPU kernel optimization (r/LocalLLaMA)

I want to share something I've been working on, and research paper from Redhat really helped. This is my agentic GPU kernel optimizer for inference engine development Its whole goal is to look at inference engine kernels, and using cloud models, plan, test and find improvements. Nothing is changed…

2

Oct 2, 2026

README

KernelOPT

Dispatch-aware, agentic GPU-kernel optimization. Point KernelOPT at a single kernel, a directory of kernels, or a whole model — Triton kernels in a compiled PyTorch model, a CUDA op in ninfer, a ggml-cuda kernel in llama.cpp, or any CMake-based GPU repo (CUDA or HIP/ROCm) that declares its build/test/bench in kernelopt.toml. It profiles them and drives a profile → plan → edit → verify → gate loop until each is faster — or a budget runs out. Your working tree is never touched: edits happen in an isolated git worktree, every accepted candidate is committed and tagged (so it can be read or reverted instantly), and the winner is handed back as a unified diff.

An implementation of KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization (arXiv:2609.30059), extended from its Triton/Inductor setting to CUDA kernel trees (ninfer, llama.cpp) and a generic descriptor-driven backend (CUDA and HIP/ROCm), with engine-share ranking via the NVIDIA suite (nsys/ncu) or Graphsignal (CUDA or ROCm).

ModeTarget treeOptimization unitCorrectnessTiming
tritoncompiled PyTorch modelone Inductor @triton.jit kerneleager PyTorch + allclosetriton.testing.do_bench
ninferninfer checkoutone src/ops/<family>/<impl>/*.cuthe Op's own ctest suiteninfer_<op>_bench
llamacppllama.cpp checkoutone ggml/src/ggml-cuda/*.cutest-backend-ops test -o <OP>test-backend-ops perf -o <OP>
customany CMake GPU repo (CUDA or HIP/ROCm) with kernelopt.tomla file you declareyour test_cmdyour bench_cmd

How it works

discover ─▶ baseline ─▶ ┌ plan ─▶ edit ─▶ compile ─▶ correctness ─▶ bench ┐ ─▶ γ/roofline gate ─▶ shape gate ─▶ diff
                        └────────── iterate until patience/target/budget ──┘

LLM agents (Planner → Executor → Summarizer) work with an NCU/Graphsignal-guided UCB-guided beam search and shared experience memory. Five gates filter candidates: static compile, multi-seed correctness, model-level (engine) verification, a performance gate (pinned shape, interleaved fresh baseline/candidate rounds, a noise floor and a sign test), and measured-shape correctness — the op's own suite may not cover the launch path the perf gate measures. A perf win must also be physically plausible: a candidate that appears to beat the device's memory roofline is rejected as "doing less work", not accepted as a win. If nothing passes, the baseline is preserved.

The Planner can also pull context on demand (search_repo/read_file, ripgrep) instead of receiving a full file dump, and every accepted candidate (compile + correctness + bench) is committed and tagged for fast revert. The Executor edits by whole file (default) or by unified diff (--edit-mode patch, opt-in) to avoid reproduction drift; either way the run outputs a reviewable unified diff and your checkout is never modified.

Measurement rigor: an op-level speedup is measured on a pinned shape with 3 repeats, then re-confirmed with interleaved fresh baseline/candidate re-benches (--final-rounds). A candidate is only called optimized if it beats the baseline by more than the measured noise, wins every round, is physically plausible (below a generous multiple of the memory roofline), and is correct on the measured shape; otherwise it is matched (correct but not faster) or fallback (rejected). Correctness-only targets (no bench) report unverified: a candidate was applied, but no speedup could be measured. The reported number is always explicit (N× = baseline/final, with the shape, ms values, noise, and bandwidth-vs-roofline).

One GPU job at a time: bench/verify/NCU/engine-E2E take an exclusive cross-process lock (.kernelopt/gpu.lock), so a campaign and a single run — or two campaigns — serialize on the device instead of corrupting each other's timing or VRAM.

Requirements

  • An NVIDIA GPU (or AMD with ROCm) + recent driver.
  • CUDA toolkit (nvcc, nsys, ncu) — or ROCm (hipcc).
  • Rust stable, Python 3.10+, cmake ≥ 3.28, ninja/make, git.
  • A target checkout (ninfer, llama.cpp, or your own repo) that builds.
  • An LLM key + a model with tool calls.

Full ordered checklist: docs/prereqs.md — driver, toolkit, profiling permission, build tools, LLM key, verify.

Build

cargo build --release        # binary: target/release/kernelopt

Setup

New machine? Start with the ordered checklist: docs/prereqs.md (GPU driver → CUDA toolkit → profiling permission → build tools → target checkout → LLM key + tool-call model → verify). Then:

cp .env.example .env     # set NINFER_REPO / LLAMACPP_REPO / KERNELOPT_MODEL / OPENCODE_API_KEY
kernelopt setup          # full prereq table
kernelopt setup --smoke  # prove the compiler works (seconds, no GPU needed)

.env is read automatically (real env wins), so --repo/--provider/--model are optional. If you pass --repo "$NINFER_REPO", export it first (set -a; source .env; set +a) — otherwise the shell expands it to empty.

Where the LLM should run

KernelOPT talks to the optimizer LLM over HTTP, so the model itself needs no local GPU — and that is the recommended setup: point it at a cloud provider (opencode-go default, openai, openrouter) or an LLM server on another machine, and keep this machine's GPU entirely for compiling, benchmarking, and profiling your model.

Serving the LLM on the same GPU (ollama, vllm, lmstudio locally) works but contends for VRAM with the target build, the benchmarks, and especially Gate 3 (engine E2E loads a whole model — tens of GB). Expect slower runs and possible OOMs; if you do it anyway, watch nvidia-smi and prefer smaller e2e models.

Documentation lookup (CUDA/HIP API docs)

When a candidate fails to compile, KernelOPT looks the offending symbol up and attaches it to the Executor's retry:

kernelopt docs cub::WarpMergeSort     # or: __reduce_max_sync, hipMalloc, rocwmma::…
kernelopt docs --login                # one-time NVIDIA login (token auto-refreshed)

Optional; falls back to local headers. Full detail: docs/docs-lookup.md.

Quickstart

First time? Use the wizard

kernelopt wizard            # detects repo/backend/targets, asks a few questions, runs
kernelopt wizard --dry-run  # just show the command it would run
kernelopt wizard --save     # also persist the choices to .env for later runs
kernelopt wizard --yes      # accept detected defaults, no prompts (scripts/CI)

It reads your .env, probes the optimizer LLM (auth + a forced tool call), lists the discovered kernels and local models, offers quick | standard | thorough presets, then prints and runs the exact command — so you can graduate to the flags once you know what you want.

One kernel

# Compiled PyTorch model (Triton kernels)
kernelopt run examples/mlp.py
kernelopt run examples/pointwise_fused.py

# ninfer Op — smallest kernel, good first run. --watch = one-terminal live view
kernelopt run-ninfer --op add_bias --iterations 3 --beam 2 --retries 3 --watch

# llama.cpp ggml-cuda kernel
kernelopt run-llamacpp --op SOFT_MAX --iterations 3 --beam 2 --watch

Every run prints a run id and writes report.md + report.diff under .kernelopt/runs/<run_id>/. --watch shows progress in the same terminal and prints those paths when the run finishes.

Go further

docs/advanced.md — harder measurements, engine verification, custom engines, campaigns, profiling, run control, outputs, testing, troubleshooting, repo layout.

Documentation

DocContents
docs/advanced.mdBeyond the basics: harder measurements, engine verification, campaigns, profiling, run control, troubleshooting
docs/cli.mdEvery command and option: what discover, profile, campaign, run* do, the full flag reference, and campaign option effects
docs/configuration.md.env, providers, thinking level, config.toml, hyperparameters
docs/prereqs.mdOrdered setup checklist: driver, CUDA toolkit, profiling permission, build tools, LLM key, verify
docs/outputs.mdArtifacts, run/campaign state, applying diffs, disk
docs/campaign.mdCampaign walkthrough: discovery, ordering, budgets, resume
docs/monitoring.mdLive progress, what the LLM is doing, pausing (Ctrl-C), watch, campaign status
docs/kernel-editing.mdHow kernels are read, edited (full-file), isolated, and surfaced as a diff
docs/docs-lookup.mdCUDA/HIP API docs lookup: NVIDIA cuda-docs MCP login, ROCm headers, kernelopt docs
docs/testing-kernel-wins.mdRST recommendations: oracles, benchmark controls, measured-shape coverage and kernel versus application wins
docs/model-e2e.mdModel-level (engine) verification: same tokens, not slower
docs/ninfer-mode.mdninfer mapping (gates, workbench, prompts)
docs/llamacpp-mode.mdllama.cpp mapping (test-backend-ops gates)
docs/custom-mode.mdcustom mode: optimize any CUDA repo via kernelopt.toml
docs/graphsignal.mdEngine-share profiling, managed install, attribution rules

Outputs

Run artifacts live in .kernelopt/runs/<run_id>/ (journal.jsonl, report.md, report.diff, candidate/bench/ncu dumps); campaigns in .kernelopt/campaigns/<id>/. Accepted candidates are committed and tagged in the worktree (kernelopt/<run_id>/<candidate>); nothing is committed to your repo.

Apply the winner with git apply, inspect with kernelopt history, or revert with kernelopt revert — details in docs/advanced.md.

License

Apache License 2.0 — see LICENSE.

Citation

This project implements the method and adapts the agent prompts from KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization (Poddar, Prasad, Samanta, Chakraborty, Goyal, Rathaur; arXiv:2609.30059).

@misc{poddar2026kernelopt,
  title        = {KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization},
  author       = {Poddar, Aheli and Prasad, Sanskar and Samanta, Arindam and
                  Chakraborty, Subha and Goyal, Vishal and Rathaur, Rohit Singh},
  year         = {2026},
  eprint       = {2609.30059},
  archivePrefix= {arXiv},
  primaryClass = {cs.DC},
  url          = {https://arxiv.org/abs/2609.30059}
}

See CITATION.cff and NOTICE. The Planner, Executor, and Summarizer prompts under prompts/ are adapted from the paper's appendix (arXiv:2609.30059). Graphsignal: https://github.com/graphsignal/graphsignal (fork adds llama.cpp/NInfer).

Languages

Rust

89.6%

Python

10.4%