12 repos
bentoml/BentoML
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps,…
8,834
3,743 commits
vllm-project/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
2,800
4,658 commits
predibase/lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
3,829
876 commits
microsoft/aici
AICI: Prompts as (Wasm) Programs
2,077
1,617 commits
liguodongiot/llm-action
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
25,038
952 commits
efeslab/Nanoflow
A throughput-oriented high-performance serving framework for LLMs
977
147 commits
mosecorg/mosec
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to…
902
480 commits
OpenMachine-ai/transformer-tricks
A collection of tricks and tools to speed up transformer models
226
210 commits
rohan-paul/LLM-FineTuning-Large-Language-Models
LLM (Large Language Model) FineTuning
578
166 commits
Scottcjn/ram-coffers
LLM infrastructure cost reduction via NUMA-aware weight banking: 147 t/s (8.8x stock llama.cpp) on…
164
90 commits
skypilot-org/skypilot
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI…
10,592
5,674 commits
FedML-AI/FedML
FEDML - The unified and scalable ML library for large-scale distributed training, model serving,…
4,065
10,439 commits