LLM Serving and Inference Optimization

36 repos

Python libraries and frameworks focused on deploying, serving, and optimizing large language models in production. The cluster centers on high-performance inference engines, batch processing, memory optimization, and hardware acceleration (particularly AMD support). vLLM, the dominant repository by internal connectivity, exemplifies the core work here—providing efficient LLM serving infrastructure with features like paged attention and continuous batching. Repositories in this cluster address the practical engineering challenges of running LLMs at scale, from quantization and pruning to distributed inference and dynamic scheduling.

Python · 35
JavaScript · 1
pytorch ·143,108
llm-serving ·135,299
llm ·135,299
inference ·98,285
transformer ·98,285
model-serving ·98,285
cuda ·91,518
kimi ·91,518
deepseek-v3 ·91,518
amd ·91,518

krafton-ai/vllm-omni

No description

Python

4

1,519 commits