LLM Serving and Inference Optimization

9 repos

Python-based frameworks and tools for deploying, serving, and optimizing large language models in production. This cluster covers inference engines (vLLM, LoRAX, AICI), serving platforms (BentoML), operational automation (llm-action), and performance techniques across different hardware backends including specialized accelerators. Developers exploring this area will find systems for efficient batch processing, token scheduling, multi-LoRA serving, and practical LLM deployment patterns.

Python · 3
Jupyter Notebook · 2
C++ · 1
HTML · 1
Rust · 1
TeX · 1
llm ·44,621
llm-serving ·44,224
llmops ·43,066
llm-inference ·40,801
llm-training ·25,647
model-serving ·18,577
mlops ·11,690
machine-learning ·9,074
python ·9,074
ai-inference ·9,017