If you tried Jan Desktop and liked it, please also check out the following awesome collection of open source and/or local AI tools and solutions.
Your contributions are always welcome!
| Repository | Description | Supported model formats | CPU/GPU Support | UI | language | Platform Type |
|---|---|---|---|---|---|---|
| llama.cpp | - Inference of LLaMA model in pure C/C++ | GGML/GGUF | Both | ❌ | C/C++ | Text-Gen |
| Cortex | - Multi-engine engine embeddable in your apps. Uses llama.cpp and more | Both | Both | ❌ | Text-Gen | |
| ollama | - CLI and local server. Uses llama.cpp | Both | Both | ❌ | Text-Gen | |
| koboldcpp | - A simple one-file way to run various GGML models with KoboldAI's UI | GGML | Both | ✅ | C/C++ | Text-Gen |
| LoLLMS | - Lord of Large Language Models Web User Interface. | Nearly ALL | Both | ✅ | Python | Text-Gen |
| ExLlama | - A more memory-efficient rewrite of the HF transformers implementation of Llama | AutoGPTQ/GPTQ | GPU | ✅ | Python/C++ | Text-Gen |
| vLLM | - vLLM is a fast and easy-to-use library for LLM inference and serving. | GGML/GGUF | Both | ❌ | Python | Text-Gen |
| SGLang | - 3-5x higher throughput than vLLM (Control flow, RadixAttention, KV cache reuse) | Safetensor / AWQ / GPTQ | GPU | ❌ | Python | Text-Gen |
| LmDeploy | - LMDeploy is a toolkit for compressing, deploying, and serving LLMs. | Pytorch / Turbomind | Both | ❌ | Python/C++ | Text-Gen |
| Tensorrt-llm | - Inference efficiently on NVIDIA GPUs | Python / C++ runtimes | Both | ❌ | Python/C++ | Text-Gen |
| CTransformers | - Python bindings for the Transformer models implemented in C/C++ using GGML library | GGML/GPTQ | Both | ❌ | C/C++ | Text-Gen |
| llama-cpp-python | - Python bindings for llama.cpp | GGUF | Both | ❌ | Python | Text-Gen |
| llama2.rs | - A fast llama2 decoder in pure Rust | GPTQ | CPU | ❌ | Rust | Text-Gen |
| ExLlamaV2 | - A fast inference library for running LLMs locally on modern consumer-class GPUs | GPTQ/EXL2 | GPU | ❌ | Python/C++ | Text-Gen |
| LoRAX | - Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs | Safetensor / AWQ / GPTQ | GPU | ❌ | Python/Rust | Text-Gen |
| text-generation-inference | - Inference serving toolbox with optimized kernels for each LLM architecture | Safetensors / AWQ / GPTQ | Both | ❌ | Python/Rust | Text-Gen |
If you tried Jan Desktop and liked it, please also check out the following awesome collection of open source and/or local AI tools and solutions.
Your contributions are always welcome!
| Repository | Description | Supported model formats | CPU/GPU Support | UI | language | Platform Type |
|---|---|---|---|---|---|---|
| llama.cpp | - Inference of LLaMA model in pure C/C++ | GGML/GGUF | Both | ❌ | C/C++ | Text-Gen |
| Cortex | - Multi-engine engine embeddable in your apps. Uses llama.cpp and more | Both | Both | ❌ | Text-Gen | |
| ollama | - CLI and local server. Uses llama.cpp | Both | Both | ❌ | Text-Gen | |
| koboldcpp | - A simple one-file way to run various GGML models with KoboldAI's UI | GGML | Both | ✅ | C/C++ | Text-Gen |
| LoLLMS | - Lord of Large Language Models Web User Interface. | Nearly ALL | Both | ✅ | Python | Text-Gen |
| ExLlama | - A more memory-efficient rewrite of the HF transformers implementation of Llama | AutoGPTQ/GPTQ | GPU | ✅ | Python/C++ | Text-Gen |
| vLLM | - vLLM is a fast and easy-to-use library for LLM inference and serving. | GGML/GGUF | Both | ❌ | Python | Text-Gen |
| SGLang | - 3-5x higher throughput than vLLM (Control flow, RadixAttention, KV cache reuse) | Safetensor / AWQ / GPTQ | GPU | ❌ | Python | Text-Gen |
| LmDeploy | - LMDeploy is a toolkit for compressing, deploying, and serving LLMs. | Pytorch / Turbomind | Both | ❌ | Python/C++ | Text-Gen |
| Tensorrt-llm | - Inference efficiently on NVIDIA GPUs | Python / C++ runtimes | Both | ❌ | Python/C++ | Text-Gen |
| CTransformers | - Python bindings for the Transformer models implemented in C/C++ using GGML library | GGML/GPTQ | Both | ❌ | C/C++ | Text-Gen |
| llama-cpp-python | - Python bindings for llama.cpp | GGUF | Both | ❌ | Python | Text-Gen |
| llama2.rs | - A fast llama2 decoder in pure Rust | GPTQ | CPU | ❌ | Rust | Text-Gen |
| ExLlamaV2 | - A fast inference library for running LLMs locally on modern consumer-class GPUs | GPTQ/EXL2 | GPU | ❌ | Python/C++ | Text-Gen |
| LoRAX | - Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs | Safetensor / AWQ / GPTQ | GPU | ❌ | Python/Rust | Text-Gen |
| text-generation-inference | - Inference serving toolbox with optimized kernels for each LLM architecture | Safetensors / AWQ / GPTQ | Both | ❌ | Python/Rust | Text-Gen |