AI-Guru/ai_services

A collection of useful AI services for AI sovereignty.

HTML

48

213 commits

updated Oct 2, 2026

See the code

README

AI Services

Background

A collection of useful AI services for AI sovereignty.

Video

Overview

This repository contains a set of containerized AI services that can be run locally to provide various AI capabilities without relying on external cloud providers. Each service is designed to be easy to deploy and use.

Models

Local LLM inference with multiple backends (vLLM, llama.cpp, SGLang, MLX) and hardware targets (RTX PRO 6000, DGX Spark, AMD Vulkan, Apple Silicon). Each family lives under models/<family>/ as a set of docker-compose.<engine>-<variant>.yml files serving an OpenAI-compatible API. See models/README.md for the full variant/benchmark matrix and a "which variant should I use?" decision tree.

Text generation & coding

ModelDescriptionLocation
Qwen3.5Flagship family, 0.8B–122B dense/MoE variantsmodels/qwen3.5
Qwen3.6Newer hybrid arch (Gated DeltaNet + Attention): 27B dense + 35B-A3B MoEmodels/qwen3.6
Qwen3-Coder-Next80B MoE coding specialist (~3B active)models/qwen3-coder-next
QwopusOpus-reasoning distilled 27B densemodels/qwopus
GLM-4.7-Flash30B MoE, ~3.6B active paramsmodels/glm-4.7-flash
NemotronNVIDIA Cascade-2 / Nano family, hybrid Mamba-2 MoE (4B–120B)models/nemotron
Gemma 4Google, Apache 2.0, multimodal (text/image/audio), E2B–31Bmodels/gemma4
Carnice-V2-27BHermes-style agent SFT of Qwen3.6-27Bmodels/carnice-v2
Mistral Medium 3.5Dense 128B, multimodal, 256K contextmodels/mistral-medium-3.5

Specialized

ModelDescriptionLocation
Qwen3-Embedding & RerankerRAG building blocks — embeddings + reranking/scoring APIsmodels/qwen3-embedding
Qwen3-ASRSpeech-to-text (52 languages) + forced aligner for timestampsmodels/qwen3-asr
Qwen3GuardGenerative safety classifier (Safe/Controversial/Unsafe, 119 languages)models/qwen3guard
DeepSeek-OCRVision-LM, documents → markdown / HTML tables / LaTeXmodels/deepseek-ocr

Shared test and benchmark scripts live in models/shared.

Speech Services

ServiceDescriptionLocationPort
WhisperSpeech-to-text using OpenAI Whisperspeech/whisper8000
Faster WhisperOptimized Whisper variantspeech/faster-whisper—
Orpheus TTSHigh-quality voice synthesisspeech/orpheus5005

Image Services

ServiceDescriptionLocationPort
open-genmojiCustom emoji generation (Flux.1[dev] + LoRA, FP8 on Blackwell)open-genmoji8888

Monitoring

ServiceDescriptionLocationPort
GPU DashboardGrafana + Prometheus + nvidia_gpu_exporter for GPU metricsgpu-dashboard3000 (Grafana), 9090 (Prometheus), 9835 (exporter)
NetdataReal-time system & GPU monitoring with auto-detected NVIDIA metricsnetdata19999

Chat Frontend

ServiceDescriptionLocationPort
LibreChatWeb chat UI wired to the local inference backends (OpenAI-compatible)librechat3080

Other Services

Ollama

A server that runs large language models (LLMs) locally with GPU acceleration support.

  • Features: Supports various open-source models, API access
  • Location: ollama
  • Port: 11434

Demo App (Voice Chat Assistant)

A real-time voice assistant integrating WebRTC, Whisper, Gemma 3, and Orpheus for end-to-end voice chat.

Getting Started

Each service has its own README.md with specific setup instructions and usage examples. Generally, you can start each service using:

cd service_directory
docker compose up -d

Kudos and Credits

This project would not have been possible without the great works of many people who steadily contribute to the open source community!

System Requirements

  • Docker and Docker Compose
  • NVIDIA GPU with CUDA support (recommended for optimal performance)
  • Sufficient disk space for model storage

License

See the LICENSE file for details.

AI-Guru/ai_services

A collection of useful AI services for AI sovereignty.

HTML

48

213 commits

updated Oct 2, 2026

See the code

README

AI Services

Background

A collection of useful AI services for AI sovereignty.

Video

Overview

This repository contains a set of containerized AI services that can be run locally to provide various AI capabilities without relying on external cloud providers. Each service is designed to be easy to deploy and use.

Models

Local LLM inference with multiple backends (vLLM, llama.cpp, SGLang, MLX) and hardware targets (RTX PRO 6000, DGX Spark, AMD Vulkan, Apple Silicon). Each family lives under models/<family>/ as a set of docker-compose.<engine>-<variant>.yml files serving an OpenAI-compatible API. See models/README.md for the full variant/benchmark matrix and a "which variant should I use?" decision tree.

Text generation & coding

ModelDescriptionLocation
Qwen3.5Flagship family, 0.8B–122B dense/MoE variantsmodels/qwen3.5
Qwen3.6Newer hybrid arch (Gated DeltaNet + Attention): 27B dense + 35B-A3B MoEmodels/qwen3.6
Qwen3-Coder-Next80B MoE coding specialist (~3B active)models/qwen3-coder-next
QwopusOpus-reasoning distilled 27B densemodels/qwopus
GLM-4.7-Flash30B MoE, ~3.6B active paramsmodels/glm-4.7-flash
NemotronNVIDIA Cascade-2 / Nano family, hybrid Mamba-2 MoE (4B–120B)models/nemotron
Gemma 4Google, Apache 2.0, multimodal (text/image/audio), E2B–31Bmodels/gemma4
Carnice-V2-27BHermes-style agent SFT of Qwen3.6-27Bmodels/carnice-v2
Mistral Medium 3.5Dense 128B, multimodal, 256K contextmodels/mistral-medium-3.5

Specialized

ModelDescriptionLocation
Qwen3-Embedding & RerankerRAG building blocks — embeddings + reranking/scoring APIsmodels/qwen3-embedding
Qwen3-ASRSpeech-to-text (52 languages) + forced aligner for timestampsmodels/qwen3-asr
Qwen3GuardGenerative safety classifier (Safe/Controversial/Unsafe, 119 languages)models/qwen3guard
DeepSeek-OCRVision-LM, documents → markdown / HTML tables / LaTeXmodels/deepseek-ocr

Shared test and benchmark scripts live in models/shared.

Speech Services

ServiceDescriptionLocationPort
WhisperSpeech-to-text using OpenAI Whisperspeech/whisper8000
Faster WhisperOptimized Whisper variantspeech/faster-whisper—
Orpheus TTSHigh-quality voice synthesisspeech/orpheus5005

Image Services

ServiceDescriptionLocationPort
open-genmojiCustom emoji generation (Flux.1[dev] + LoRA, FP8 on Blackwell)open-genmoji8888

Monitoring

ServiceDescriptionLocationPort
GPU DashboardGrafana + Prometheus + nvidia_gpu_exporter for GPU metricsgpu-dashboard3000 (Grafana), 9090 (Prometheus), 9835 (exporter)
NetdataReal-time system & GPU monitoring with auto-detected NVIDIA metricsnetdata19999

Chat Frontend

ServiceDescriptionLocationPort
LibreChatWeb chat UI wired to the local inference backends (OpenAI-compatible)librechat3080

Other Services

Ollama

A server that runs large language models (LLMs) locally with GPU acceleration support.

  • Features: Supports various open-source models, API access
  • Location: ollama
  • Port: 11434

Demo App (Voice Chat Assistant)

A real-time voice assistant integrating WebRTC, Whisper, Gemma 3, and Orpheus for end-to-end voice chat.

Getting Started

Each service has its own README.md with specific setup instructions and usage examples. Generally, you can start each service using:

cd service_directory
docker compose up -d

Kudos and Credits

This project would not have been possible without the great works of many people who steadily contribute to the open source community!

System Requirements

  • Docker and Docker Compose
  • NVIDIA GPU with CUDA support (recommended for optimal performance)
  • Sufficient disk space for model storage

License

See the LICENSE file for details.

Languages

HTML

57.6%

Python

25.9%

Shell

13.7%

Jinja

2.1%