LLM Inference Optimization & Serving

15 repos

Python-dominant tooling for optimizing and serving large language model inference, with focus on quantization, model compression, and efficient inference engines. The cluster covers production deployment frameworks (vLLM, SGLang), quantization approaches (GPTQModel, auto-round), and distributed inference systems (gpustack, infercrane) — representing the practical engineering layer between raw model weights and deployed LLM applications.

Python · 11
C++ · 2
Go · 1
Rust · 1
vllm ·37,764
sglang ·32,677
inference ·26,729
llm ·26,729
qwen ·15,238
llm-inference ·15,057
transformers ·10,913
openai-api ·9,693
llama-cpp ·9,661
glm-5-3 ·9,564