LLM Inference Optimization

15 repos

Techniques and implementations for accelerating large language model inference through methods like speculative decoding, quantization, and GPU optimization. The cluster centers on Qwen model variants (particularly 27B parameters) optimized for consumer GPUs via tools like DFlash2 and NVIDIA FP4 quantization, alongside broader frameworks for deploying and experimenting with optimized inference pipelines locally. Most repositories are Python-based with CUDA kernel implementations, reflecting the practical focus on making capable models run faster and cheaper on real hardware.

Python · 4
C++ · 2
HTML · 1
Rust · 1
speculative-decoding ·7,231
qwen ·6,947
cuda ·5,694
megakernel ·5,686
local-ai ·5,686
dflash ·5,686
heterogeneous-computing ·5,686
kernel ·5,686
cuda-kernels ·5,686
llama-cpp ·5,686