Workflow Defined Engine
Python
25
188 commits
updated Nov 4, 2025
sskim-ai/Dynamo
0
ai-dynamo/dynamo
A Datacenter Scale Distributed Inference Serving Framework
8,173
ztxz16/fastllm
fastllm是后端无依赖的高性能大模型推理库。同时支持张量并行推理稠密模型和混合模式推理MOE模型,任意10G以上显卡即可推理满血DeepSeek。双路9004/9005服务器+单显卡部署Deep…
5,077
llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
4,667
CURRENTF/SparseEngine
A sparse-first inference engine (SparseEngine).
77
PaddlePaddle/FastDeploy
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
3,719
WJX3078/mini-vllm
从零实现的 mini vLLM 推理引擎:Block 级 KV Cache (PagedAttention) · Continuous Batching · Prefix Caching ·…
xLLM-AI/xllm
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI…
1,582
88.8%
Jupyter Notebook
11.1%