15 repos
Python-dominant tooling for optimizing and serving large language model inference, with focus on quantization, model compression, and efficient inference engines. The cluster covers production deployment frameworks (vLLM, SGLang), quantization approaches (GPTQModel, auto-round), and distributed inference systems (gpustack, infercrane) — representing the practical engineering layer between raw model weights and deployed LLM applications.