LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development
See the codeGPU kernels are central to modern compute stacks and directly determine training and inference efficiency. Kernel development is difficult because it requires hardware expertise and iterative refinement with multi step tool feedback. Since Stanford released KernelBench in February 2025, the LLM4Kernel field has grown rapidly, with increasing interest in using large language models to support or automate kernel generation, optimization, and verification.
This project provides a continuous and comprehensive survey of the field, covering both benchmarks and methods. On the methodological side, we categorize existing work into four major directions:
We include all relevant top conference papers, arXiv preprints, open source projects, technical reports, and blogs, aiming to build the most complete resource hub for LLM4Kernel research.
Online page: https://kechang.xin/Awesome-LLM4Kernel/
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
KernelBench: Can LLMs Write Efficient GPU Kernels?
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
ComputeEval: Evaluating Large Language Models for CUDA Code Generation
BackendBench: An Evaluation Suite for Testing How Well LLMs and Humans Can Write PyTorch Backends
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
robust-kbench: Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
gpuFLOPBench: Counting Without Running: Evaluating LLMs’ Reasoning About Code Complexity
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
Can Large Language Models Predict Parallel Code Performance
NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware
KernelBook: PyTorch to Triton Code Translation Dataset
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
KernelBand: Boosting LLM-based Kernel Optimization with a Hierarchical and Hardware-aware Multi-armed Bandit
Automating GPU Kernel Generation with DeepSeek-R1 and Inference Time Scaling
Tutoring LLM into a Better CUDA Optimizer
GPU Performance Portability needs Autotuning
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
EVOENGINEER: Mastering Automated CUDA Kernel Code Evolution with Large Language Models
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
AVO: Agentic Variation Operators for Autonomous Evolutionary Search
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
CUCo: An Agentic Framework for Compute and Communication Co-design
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning
AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
TritorX: Agentic Operator Generation for ML ASICs
KernelFalcon: Autonomous GPU Kernel Generation via Deep Agents
STARK: Strategic Team of Agents for Refining Kernels
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
How Many Agents Does it Take to Beat PyTorch? (surprisingly not that much)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
KForge: Program Synthesis for Diverse AI Hardware Accelerators
The AI CUDA engineer: Agentic CUDA kernel discovery, optimization and composition
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization
cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
Adaptive Self-improvement LLM Agentic System for ML Library Development
InCoder-32B: Code Foundation Model for Industrial Scenarios
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
CudaLLM: Training Language Models to Generate High-Performance CUDA Kernels
Omniwise: Predicting GPU Kernels Performance with LLMs
ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
Fine-Tuning GPT-5 for GPU Kernel Generation
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
TRITONRL: Training LLMs to Think and Code Triton Without Cheating
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement Learning
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
Kevin: Multi-Turn RL for Generating CUDA Kernels
Feel free to open an issue or submit a pull request to correct errors or add work that has not yet been included in this project. You can also email us at kcxain@gmail.com for any form of discussion and collaboration.
If you find this work useful, welcome to cite us.
@article{llm4kernel,
title={LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development},
author={Changxin Ke},
year={2025}
url={https://github.com/kcxain/Awesome-LLM4Kernel}
}
HTML
100.0%
LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development
See the codeGPU kernels are central to modern compute stacks and directly determine training and inference efficiency. Kernel development is difficult because it requires hardware expertise and iterative refinement with multi step tool feedback. Since Stanford released KernelBench in February 2025, the LLM4Kernel field has grown rapidly, with increasing interest in using large language models to support or automate kernel generation, optimization, and verification.
This project provides a continuous and comprehensive survey of the field, covering both benchmarks and methods. On the methodological side, we categorize existing work into four major directions:
We include all relevant top conference papers, arXiv preprints, open source projects, technical reports, and blogs, aiming to build the most complete resource hub for LLM4Kernel research.
Online page: https://kechang.xin/Awesome-LLM4Kernel/
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
KernelBench: Can LLMs Write Efficient GPU Kernels?
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
ComputeEval: Evaluating Large Language Models for CUDA Code Generation
BackendBench: An Evaluation Suite for Testing How Well LLMs and Humans Can Write PyTorch Backends
MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
robust-kbench: Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
gpuFLOPBench: Counting Without Running: Evaluating LLMs’ Reasoning About Code Complexity
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
Can Large Language Models Predict Parallel Code Performance
NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers
CUDABench: Benchmarking LLMs for Text-to-CUDA Generation
KernelCraft: Benchmarking for Agentic Close-to-Metal Kernel Generation on Emerging Hardware
KernelBook: PyTorch to Triton Code Translation Dataset
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
OptiML: An End-to-End Framework for Program Synthesis and CUDA Kernel Optimization
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
KernelBand: Boosting LLM-based Kernel Optimization with a Hierarchical and Hardware-aware Multi-armed Bandit
Automating GPU Kernel Generation with DeepSeek-R1 and Inference Time Scaling
Tutoring LLM into a Better CUDA Optimizer
GPU Performance Portability needs Autotuning
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
EVOENGINEER: Mastering Automated CUDA Kernel Code Evolution with Large Language Models
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
MaxCode: A Max-Reward Reinforcement Learning Framework for Automated Code Optimization
AVO: Agentic Variation Operators for Autonomous Evolutionary Search
StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning
KernelSkill: A Multi-Agent Framework for GPU Kernel Optimization
Making LLMs Optimize Multi-Scenario CUDA Kernels Like Experts
CUCo: An Agentic Framework for Compute and Communication Co-design
AscendCraft: Automatic Ascend NPU Kernel Generation via DSL-Guided Transcompilation
KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning
AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis
TritorX: Agentic Operator Generation for ML ASICs
KernelFalcon: Autonomous GPU Kernel Generation via Deep Agents
STARK: Strategic Team of Agents for Refining Kernels
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
QiMeng-TensorOp: Automatically Generating High-Performance Tensor Operators with Hardware Primitives
QiMeng-GEMM: Automatically Generating High-Performance Matrix Multiplication Code by Exploiting Large Language Models
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks
How Many Agents Does it Take to Beat PyTorch? (surprisingly not that much)
Astra: A Multi-Agent System for GPU Kernel Performance Optimization
CudaForge: An Agent Framework with Hardware Feedback for CUDA Kernel Optimization
KForge: Program Synthesis for Diverse AI Hardware Accelerators
The AI CUDA engineer: Agentic CUDA kernel discovery, optimization and composition
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
PRAGMA: A Profiling-Reasoned Multi-Agent Framework for Automatic Kernel Optimization
cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization
Adaptive Self-improvement LLM Agentic System for ML Library Development
InCoder-32B: Code Foundation Model for Industrial Scenarios
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
AutoTriton: Automatic Triton Programming with Reinforcement Learning in LLMs
QiMeng-MuPa: Mutual-Supervised Learning for Sequential-to-Parallel Code Translation
AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural Processing Units
CUDA-LLM: LLMs Can Write Efficient CUDA Kernels
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
CudaLLM: Training Language Models to Generate High-Performance CUDA Kernels
Omniwise: Predicting GPU Kernels Performance with LLMs
ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
Fine-Tuning GPT-5 for GPU Kernel Generation
QiMeng-Kernel: Macro-Thinking Micro-Coding Paradigm for LLM-Based High-Performance GPU Kernel Generation
CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
TRITONRL: Training LLMs to Think and Code Triton Without Cheating
CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning
Mastering Sparse CUDA Generation through Pretrained Models and Deep Reinforcement Learning
SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
Kevin: Multi-Turn RL for Generating CUDA Kernels
Feel free to open an issue or submit a pull request to correct errors or add work that has not yet been included in this project. You can also email us at kcxain@gmail.com for any form of discussion and collaboration.
If you find this work useful, welcome to cite us.
@article{llm4kernel,
title={LLM4Kernel: A Survey of Large Language Models for GPU Kernel Development},
author={Changxin Ke},
year={2025}
url={https://github.com/kcxain/Awesome-LLM4Kernel}
}
HTML
100.0%