A collection of useful AI services for AI sovereignty.
HTML
48
213 commits
updated Oct 2, 2026

A collection of useful AI services for AI sovereignty.
This repository contains a set of containerized AI services that can be run locally to provide various AI capabilities without relying on external cloud providers. Each service is designed to be easy to deploy and use.
Local LLM inference with multiple backends (vLLM, llama.cpp, SGLang, MLX) and hardware targets (RTX PRO 6000, DGX Spark, AMD Vulkan, Apple Silicon). Each family lives under models/<family>/ as a set of docker-compose.<engine>-<variant>.yml files serving an OpenAI-compatible API. See models/README.md for the full variant/benchmark matrix and a "which variant should I use?" decision tree.
| Model | Description | Location |
|---|---|---|
| Qwen3.5 | Flagship family, 0.8B–122B dense/MoE variants | models/qwen3.5 |
| Qwen3.6 | Newer hybrid arch (Gated DeltaNet + Attention): 27B dense + 35B-A3B MoE | models/qwen3.6 |
| Qwen3-Coder-Next | 80B MoE coding specialist (~3B active) | models/qwen3-coder-next |
| Qwopus | Opus-reasoning distilled 27B dense | models/qwopus |
| GLM-4.7-Flash | 30B MoE, ~3.6B active params | models/glm-4.7-flash |
| Nemotron | NVIDIA Cascade-2 / Nano family, hybrid Mamba-2 MoE (4B–120B) | models/nemotron |
| Gemma 4 | Google, Apache 2.0, multimodal (text/image/audio), E2B–31B | models/gemma4 |
| Carnice-V2-27B | Hermes-style agent SFT of Qwen3.6-27B | models/carnice-v2 |
| Mistral Medium 3.5 | Dense 128B, multimodal, 256K context | models/mistral-medium-3.5 |
| Model | Description | Location |
|---|---|---|
| Qwen3-Embedding & Reranker | RAG building blocks — embeddings + reranking/scoring APIs | models/qwen3-embedding |
| Qwen3-ASR | Speech-to-text (52 languages) + forced aligner for timestamps | models/qwen3-asr |
| Qwen3Guard | Generative safety classifier (Safe/Controversial/Unsafe, 119 languages) | models/qwen3guard |
| DeepSeek-OCR | Vision-LM, documents → markdown / HTML tables / LaTeX | models/deepseek-ocr |
Shared test and benchmark scripts live in models/shared.
| Service | Description | Location | Port |
|---|---|---|---|
| Whisper | Speech-to-text using OpenAI Whisper | speech/whisper | 8000 |
| Faster Whisper | Optimized Whisper variant | speech/faster-whisper | — |
| Orpheus TTS | High-quality voice synthesis | speech/orpheus | 5005 |
| Service | Description | Location | Port |
|---|---|---|---|
| open-genmoji | Custom emoji generation (Flux.1[dev] + LoRA, FP8 on Blackwell) | open-genmoji | 8888 |
| Service | Description | Location | Port |
|---|---|---|---|
| GPU Dashboard | Grafana + Prometheus + nvidia_gpu_exporter for GPU metrics | gpu-dashboard | 3000 (Grafana), 9090 (Prometheus), 9835 (exporter) |
| Netdata | Real-time system & GPU monitoring with auto-detected NVIDIA metrics | netdata | 19999 |
| Service | Description | Location | Port |
|---|---|---|---|
| LibreChat | Web chat UI wired to the local inference backends (OpenAI-compatible) | librechat | 3080 |
A server that runs large language models (LLMs) locally with GPU acceleration support.
A real-time voice assistant integrating WebRTC, Whisper, Gemma 3, and Orpheus for end-to-end voice chat.
Each service has its own README.md with specific setup instructions and usage examples. Generally, you can start each service using:
cd service_directory
docker compose up -d
This project would not have been possible without the great works of many people who steadily contribute to the open source community!
See the LICENSE file for details.
HTML
57.6%
Python
25.9%
Shell
13.7%
Jinja
2.1%
A collection of useful AI services for AI sovereignty.
HTML
48
213 commits
updated Oct 2, 2026

A collection of useful AI services for AI sovereignty.
This repository contains a set of containerized AI services that can be run locally to provide various AI capabilities without relying on external cloud providers. Each service is designed to be easy to deploy and use.
Local LLM inference with multiple backends (vLLM, llama.cpp, SGLang, MLX) and hardware targets (RTX PRO 6000, DGX Spark, AMD Vulkan, Apple Silicon). Each family lives under models/<family>/ as a set of docker-compose.<engine>-<variant>.yml files serving an OpenAI-compatible API. See models/README.md for the full variant/benchmark matrix and a "which variant should I use?" decision tree.
| Model | Description | Location |
|---|---|---|
| Qwen3.5 | Flagship family, 0.8B–122B dense/MoE variants | models/qwen3.5 |
| Qwen3.6 | Newer hybrid arch (Gated DeltaNet + Attention): 27B dense + 35B-A3B MoE | models/qwen3.6 |
| Qwen3-Coder-Next | 80B MoE coding specialist (~3B active) | models/qwen3-coder-next |
| Qwopus | Opus-reasoning distilled 27B dense | models/qwopus |
| GLM-4.7-Flash | 30B MoE, ~3.6B active params | models/glm-4.7-flash |
| Nemotron | NVIDIA Cascade-2 / Nano family, hybrid Mamba-2 MoE (4B–120B) | models/nemotron |
| Gemma 4 | Google, Apache 2.0, multimodal (text/image/audio), E2B–31B | models/gemma4 |
| Carnice-V2-27B | Hermes-style agent SFT of Qwen3.6-27B | models/carnice-v2 |
| Mistral Medium 3.5 | Dense 128B, multimodal, 256K context | models/mistral-medium-3.5 |
| Model | Description | Location |
|---|---|---|
| Qwen3-Embedding & Reranker | RAG building blocks — embeddings + reranking/scoring APIs | models/qwen3-embedding |
| Qwen3-ASR | Speech-to-text (52 languages) + forced aligner for timestamps | models/qwen3-asr |
| Qwen3Guard | Generative safety classifier (Safe/Controversial/Unsafe, 119 languages) | models/qwen3guard |
| DeepSeek-OCR | Vision-LM, documents → markdown / HTML tables / LaTeX | models/deepseek-ocr |
Shared test and benchmark scripts live in models/shared.
| Service | Description | Location | Port |
|---|---|---|---|
| Whisper | Speech-to-text using OpenAI Whisper | speech/whisper | 8000 |
| Faster Whisper | Optimized Whisper variant | speech/faster-whisper | — |
| Orpheus TTS | High-quality voice synthesis | speech/orpheus | 5005 |
| Service | Description | Location | Port |
|---|---|---|---|
| open-genmoji | Custom emoji generation (Flux.1[dev] + LoRA, FP8 on Blackwell) | open-genmoji | 8888 |
| Service | Description | Location | Port |
|---|---|---|---|
| GPU Dashboard | Grafana + Prometheus + nvidia_gpu_exporter for GPU metrics | gpu-dashboard | 3000 (Grafana), 9090 (Prometheus), 9835 (exporter) |
| Netdata | Real-time system & GPU monitoring with auto-detected NVIDIA metrics | netdata | 19999 |
| Service | Description | Location | Port |
|---|---|---|---|
| LibreChat | Web chat UI wired to the local inference backends (OpenAI-compatible) | librechat | 3080 |
A server that runs large language models (LLMs) locally with GPU acceleration support.
A real-time voice assistant integrating WebRTC, Whisper, Gemma 3, and Orpheus for end-to-end voice chat.
Each service has its own README.md with specific setup instructions and usage examples. Generally, you can start each service using:
cd service_directory
docker compose up -d
This project would not have been possible without the great works of many people who steadily contribute to the open source community!
See the LICENSE file for details.
HTML
57.6%
Python
25.9%
Shell
13.7%
Jinja
2.1%