34 repos
IBM/vllm
vLLM with support for span semantics
27
11,161 commits
tenstorrent/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
30
13,019 commits
MotifTechnologies/vllm
vLLM with Motif-3 support, based on vLLM v0.20.2
1
13,209 commits
vllm-project/vllm
92,059
16,100 commits
local-inference-lab/vllm
40
14,890 commits
lkm2835/vllm
2
15,696 commits
fort726/vllm
0
15,928 commits
CarrotShoo/vllm-old
7,439 commits
congcongchen123/vllm
6,833 commits
ttdxq/gfx906-vllm
vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
39
10,213 commits
1CatAI/1Cat-vLLM
V100 / SM70-focused vLLM engineering fork for modern LLM inference.
1,073
14,695 commits
guqiong96/Lvllmds4-x
CPU-GPU hybrid inference for DeepSeek-V4 on NVIDIA SM80+ (A100/RTX 4090 etc.), forked from…
61
14,339 commits
charlie12345/vLLM_for_AMD
vLLM optimized for AMD
20
15,902 commits
wtdcode/vllm-backport
Dogfooding vLLM backport for older GPUs.
245
16,228 commits
yhfgyyf/vllm-deepseek-v4-sm89
Run DeepSeek-V4.1-Flash / DeepSeek-V4-Flash and GLM-5.3-Flash on SM89 (Ada / RTX 4090) and SM120…
160
16,482 commits
guqiong96/Lvllmds4
A fork of jasl/vllm (codex/ds4-sm120-min-enable) with CPU-GPU hybrid inference support for…
31
15,896 commits
boson-ai/higgs-audio-vllm
Forked vLLM that supports higgs-audio model
47
5,171 commits
XiaomiMiMo/vllm
4,350 commits
finnchen11/VLLM_PromptCache
Optimize vLLM with persistent system prompt caching and block reuse for faster, memory-efficient…
53
6,565 commits
illinoisdata/lazy-attention
No description
7
5,731 commits
krafton-ai/vllm-omni
4
1,519 commits
wodex1nhaoIeng/cmu-15642
1,284 commits
Celeste-jq/vllm-omni-hunyuanimage3
1,857 commits
baicai-1145/vllm-GPT-SoVITS
1,082 commits
vllm-project/vllm-omni
A framework for efficient model inference with omni-modality models
6,872
2,863 commits
Green-Vial/openbmb-x-demo
openbmb x 昇腾挑战赛 演示demo
2,336 commits
vllm-project/recipes
Common recipes to run vLLM
1,023
670 commits
vllm-project/production-stack
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance…
2,600
697 commits
hadi-abdine/vllm_amd_sleep
1 commits
IBM/text-generation-inference
IBM development fork of https://github.com/huggingface/text-generation-inference
67
281 commits
Taichu-AI/vllm
3
ray-project/ray
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries…
43,862
29,103 commits
mitkox/vllm-turboquant
vLLM TurboQuant
619
3 commits
pytorch/benchmark
TorchBench is a collection of open source benchmarks used to evaluate PyTorch performance.
1,044
2,273 commits