LLM inference and deployment

6 repos

A cluster focused on running, serving, and interfacing with large language models locally and at scale. The repos span multiple implementation languages (Python, Rust, Swift) and cover model inference engines (Ollama), interactive LLM interfaces (Open Interpreter), and support for popular open models like DeepSeek, Qwen, and Mistral. Someone exploring this area would find tools for deploying LLMs, building agent-like systems on top of them, and managing the infrastructure needed to run them efficiently.

Rust · 3
JavaScript · 1
Python · 1
coding-agent ·136,616
qwen ·136,616
kimi ·136,590
rust ·136,590
acp ·136,590
deepseek ·136,590
en ·26
dora ·26
moe ·26
endpoints_compatible ·26