LLM Inference Optimization & Serving

18 repos

Python-based frameworks and tools for optimizing large language model inference performance, memory efficiency, and deployment at scale. This cluster covers serving systems (vLLM, SGLang), inference engines with techniques like quantization and batching, and specialized optimizations for running LLMs on resource-constrained hardware. Repositories span reference implementations, production serving platforms, and experimental inference acceleration methods.

Python · 11
C++ · 2
Go · 1
JavaScript · 1
Rust · 1
vllm ·47,890
llm ·35,032
inference ·34,990
sglang ·34,710
pytorch ·21,423
cuda ·18,673
rocm ·17,556
qwen ·15,283
llm-inference ·15,159
transformers ·12,711