9 repos
Python-based frameworks and tools for deploying, serving, and optimizing large language models in production. This cluster covers inference engines (vLLM, LoRAX, AICI), serving platforms (BentoML), operational automation (llm-action), and performance techniques across different hardware backends including specialized accelerators. Developers exploring this area will find systems for efficient batch processing, token scheduling, multi-LoRA serving, and practical LLM deployment patterns.