A private LLM platform for a team, on your own hardware: a chat with tools, an API with a key and a budget per person, and Argus, a code index that gives the model your codebase and real documentation, within each person's GitLab permissions. Nothing leaves your network.
Ten task families, one question each, graded on facts checked against the corpus before any model ran:
| model | alone | with Argus |
|---|---|---|
qwen3.6:27b (dense, 27.8B) | 5 / 10 | 10 / 10 |
qwen3.6:35b (MoE, 36.0B) | 5 / 10 | 9 / 10 |
qwen3.8:27b (dense, 27.3B) | 5 / 9 | 9 / 10 |
All three models failed the same five tasks alone: facts too specialised to
sit in any local model's weights (a driver's IRQL, the header CreateFileW is
really declared in). Scale did not fix it; retrieval did, and answers got faster
(5.5 s to 2.4 s median). The details are in
docs/measurements/model-comparison.md.
flowchart LR
person([People and tools]) --> traefik[Traefik<br/>TLS, routing]
traefik --> web[Web<br/>src/web]
traefik --> api[API<br/>src/Llm.Api]
traefik --> gateway[LiteLLM<br/>keys, budgets]
api --> gateway --> engine[llama.cpp<br/>the models, on the GPUs]
gateway --> remote[(Other GPU servers)]
api --> argus[Argus<br/>src/Argus]
api --> sandbox[Python sandbox<br/>no network]
gateway --> media[Pictures, video,<br/>speech servers]
api --> obs[Prometheus, Loki,<br/>Alertmanager]
argus --> gitlab[(Your GitLab)]
api --> pg[(Postgres)]
Every service, network and failure mode is in docs/architecture.md.
On a Linux host with Docker (or rootless Podman) and an NVIDIA GPU:
git clone https://github.com/Binchitects/argus && cd argus/deploy
cp .env.example .env # the domain, the models folder, the first model, six secrets
docker compose up -d
Open https://llm.localhost and sign in as admin with the password from
.env. The app fetches the first model and the picture, video and speech
models by itself. Every module runs; one is left out in a line. The full
walkthrough, Podman and upgrading from v3 are in docs/deployment.md.
Argus alone, without the platform (its own small app, no sign-in service or observability): deploy/argus-standalone/.
| path | what it is |
|---|---|
src/ | Llm.Api and Llm.Core (the platform's .NET API), Argus (the code index service), web (the platform's React app), argus-web (Argus's own app) |
tests/ | Llm.Tests and Argus.Tests (xUnit), deploy (the deployment tooling) |
deploy/ | the platform's deployment: compose (and Podman's override), config, scripts; argus-standalone/ |
tools/ | development and operations tools: dn, the test GitLab, pack builds |
clients/ | MCP configurations for Claude Code, Qwen Code, Continue, DeepSeek Harness and others (the app's Connect your tools page has each tool's full setup); the arena CLI and a GitLab CI template that reviews merge requests (docs/ci.md) |
docs/ | the documentation |
evals/ | the evaluation question sets and results |
Start at docs/README.md. The most used:
See CONTRIBUTING.md for how to build, test and propose a change, and SECURITY.md for reporting a vulnerability.
GPL v3, see LICENSE. Knowledge packs carry their own upstream
licences, which are not GPL and vary per pack (CC-BY-4.0 for the Microsoft
documentation, CC-BY-SA-3.0 for cppreference, PSF-2.0 for Python, public domain
for SQLite); argus pack info <name> prints each in full.
C#
59.0%
TypeScript
30.7%
Python
7.7%
Shell
1.5%
A private LLM platform for a team, on your own hardware: a chat with tools, an API with a key and a budget per person, and Argus, a code index that gives the model your codebase and real documentation, within each person's GitLab permissions. Nothing leaves your network.
Ten task families, one question each, graded on facts checked against the corpus before any model ran:
| model | alone | with Argus |
|---|---|---|
qwen3.6:27b (dense, 27.8B) | 5 / 10 | 10 / 10 |
qwen3.6:35b (MoE, 36.0B) | 5 / 10 | 9 / 10 |
qwen3.8:27b (dense, 27.3B) | 5 / 9 | 9 / 10 |
All three models failed the same five tasks alone: facts too specialised to
sit in any local model's weights (a driver's IRQL, the header CreateFileW is
really declared in). Scale did not fix it; retrieval did, and answers got faster
(5.5 s to 2.4 s median). The details are in
docs/measurements/model-comparison.md.
flowchart LR
person([People and tools]) --> traefik[Traefik<br/>TLS, routing]
traefik --> web[Web<br/>src/web]
traefik --> api[API<br/>src/Llm.Api]
traefik --> gateway[LiteLLM<br/>keys, budgets]
api --> gateway --> engine[llama.cpp<br/>the models, on the GPUs]
gateway --> remote[(Other GPU servers)]
api --> argus[Argus<br/>src/Argus]
api --> sandbox[Python sandbox<br/>no network]
gateway --> media[Pictures, video,<br/>speech servers]
api --> obs[Prometheus, Loki,<br/>Alertmanager]
argus --> gitlab[(Your GitLab)]
api --> pg[(Postgres)]
Every service, network and failure mode is in docs/architecture.md.
On a Linux host with Docker (or rootless Podman) and an NVIDIA GPU:
git clone https://github.com/Binchitects/argus && cd argus/deploy
cp .env.example .env # the domain, the models folder, the first model, six secrets
docker compose up -d
Open https://llm.localhost and sign in as admin with the password from
.env. The app fetches the first model and the picture, video and speech
models by itself. Every module runs; one is left out in a line. The full
walkthrough, Podman and upgrading from v3 are in docs/deployment.md.
Argus alone, without the platform (its own small app, no sign-in service or observability): deploy/argus-standalone/.
| path | what it is |
|---|---|
src/ | Llm.Api and Llm.Core (the platform's .NET API), Argus (the code index service), web (the platform's React app), argus-web (Argus's own app) |
tests/ | Llm.Tests and Argus.Tests (xUnit), deploy (the deployment tooling) |
deploy/ | the platform's deployment: compose (and Podman's override), config, scripts; argus-standalone/ |
tools/ | development and operations tools: dn, the test GitLab, pack builds |
clients/ | MCP configurations for Claude Code, Qwen Code, Continue, DeepSeek Harness and others (the app's Connect your tools page has each tool's full setup); the arena CLI and a GitLab CI template that reviews merge requests (docs/ci.md) |
docs/ | the documentation |
evals/ | the evaluation question sets and results |
Start at docs/README.md. The most used:
See CONTRIBUTING.md for how to build, test and propose a change, and SECURITY.md for reporting a vulnerability.
GPL v3, see LICENSE. Knowledge packs carry their own upstream
licences, which are not GPL and vary per pack (CC-BY-4.0 for the Microsoft
documentation, CC-BY-SA-3.0 for cppreference, PSF-2.0 for Python, public domain
for SQLite); argus pack info <name> prints each in full.
C#
59.0%
TypeScript
30.7%
Python
7.7%
Shell
1.5%