LLM Inference Optimization & Serving

23 repos

Frameworks and tools for efficiently serving and running large language models in production, with emphasis on inference optimization, speculative decoding, and specialized architectures like MoE and Llama variants. The cluster centers on SGLang and related systems that enable fast, flexible LLM inference through techniques like structured generation, dynamic batching, and hardware-aware optimization. Most repositories are Python-based implementations spanning model serving, inference engines, and compiler-level optimizations.

Python · 23
inference ·36,621
moe ·35,890
qwen ·35,789
llama ·35,766
blackwell ·35,766
gpt-oss ·35,766
diffusion ·35,766
deepseek ·35,766
cuda ·35,766
attention ·35,766