LLM Inference & Deployment Frameworks

6 repos

This cluster covers infrastructure and frameworks for running large language models in production, with emphasis on conversational AI endpoints and text generation. The repositories focus on model optimization, inference serving, and compatibility with standard transformer architectures, using formats like safetensors for efficient model loading. While anchored by prominent DeepSeek model releases (R1, V3, V4 variants) and MiniMax implementations, the cluster's dominant signal is practical tooling for deploying and serving LLMs at scale.

Python · 2
conversational ·3,497
endpoints_compatible ·3,497
text-generation ·3,497
text-generation-inference ·3,497
transformers ·3,497
safetensors ·3,497
qwen2 ·1,612
qwen3 ·1,085
llama ·800