Binchitects/argus

Deep Code Analysis & Review Skill

C#

0

578 commits

updated Oct 4, 2026

See the code

README

Argus Arena

A private LLM platform for a team, on your own hardware: a chat with tools, an API with a key and a budget per person, and Argus, a code index that gives the model your codebase and real documentation, within each person's GitLab permissions. Nothing leaves your network.

What you get

  • A chat for everyone: the models you run, thinking levels, branches and edits, files and Office documents, and tools that run on the server: Argus (your code and documentation), Python in a sandbox with no network (data, image, map and chart libraries), the web (sites you allow), pictures, video, speech, a calculator and dates. Sound and video in and out: voice messages and attached sound or video go to a model that hears and sees (Qwen3-Omni) or as a transcript and frames to one that does not; answers are read aloud. Tools ask before they run when you want them to. Pages, pictures, diagrams and React components the model writes run live in a sandboxed preview.
  • An API for editors, agents and scripts: OpenAI-compatible (and Anthropic's and OpenAI's Responses API), a key per person, spend and budgets per person, and fair use when the GPU is shared. Connect your tools walks each person through Claude Code, Codex, Qwen Code, OpenCode, Aider, Hermes, OpenClaw, DeepSeek Harness, Continue, Cline and more.
  • Sign-in for the whole stack: accounts, LDAP or Active Directory, two-factor, single sign-on for the services that have their own login, groups that decide who may use which model and tool.
  • Administration: people and groups; models read for what they are, several loaded at once (kept loaded, or loaded on request), each on the GPUs chosen for it, and models on other GPU servers behind the same gateway; picture, video and speech models with the same controls; every setting in one place, the audit log, the code index and knowledge packs.
  • Observability in the app: ten dashboards, every service's logs, and the alerts, what fires now and what fired before.
  • Argus, the code index: it mirrors your GitLab, extracts symbols and dependencies, serves eleven knowledge packs of API documentation, and answers over MCP. Each developer sees exactly the repositories their GitLab account can read, enforced in SQL.

Why Argus

Ten task families, one question each, graded on facts checked against the corpus before any model ran:

modelalonewith Argus
qwen3.6:27b (dense, 27.8B)5 / 1010 / 10
qwen3.6:35b (MoE, 36.0B)5 / 109 / 10
qwen3.8:27b (dense, 27.3B)5 / 99 / 10

All three models failed the same five tasks alone: facts too specialised to sit in any local model's weights (a driver's IRQL, the header CreateFileW is really declared in). Scale did not fix it; retrieval did, and answers got faster (5.5 s to 2.4 s median). The details are in docs/measurements/model-comparison.md.

Architecture

flowchart LR
  person([People and tools]) --> traefik[Traefik<br/>TLS, routing]
  traefik --> web[Web<br/>src/web]
  traefik --> api[API<br/>src/Llm.Api]
  traefik --> gateway[LiteLLM<br/>keys, budgets]
  api --> gateway --> engine[llama.cpp<br/>the models, on the GPUs]
  gateway --> remote[(Other GPU servers)]
  api --> argus[Argus<br/>src/Argus]
  api --> sandbox[Python sandbox<br/>no network]
  gateway --> media[Pictures, video,<br/>speech servers]
  api --> obs[Prometheus, Loki,<br/>Alertmanager]
  argus --> gitlab[(Your GitLab)]
  api --> pg[(Postgres)]

Every service, network and failure mode is in docs/architecture.md.

Quick start

On a Linux host with Docker (or rootless Podman) and an NVIDIA GPU:

git clone https://github.com/Binchitects/argus && cd argus/deploy
cp .env.example .env      # the domain, the models folder, the first model, six secrets
docker compose up -d

Open https://llm.localhost and sign in as admin with the password from .env. The app fetches the first model and the picture, video and speech models by itself. Every module runs; one is left out in a line. The full walkthrough, Podman and upgrading from v3 are in docs/deployment.md.

Argus alone, without the platform (its own small app, no sign-in service or observability): deploy/argus-standalone/.

Repository

pathwhat it is
src/Llm.Api and Llm.Core (the platform's .NET API), Argus (the code index service), web (the platform's React app), argus-web (Argus's own app)
tests/Llm.Tests and Argus.Tests (xUnit), deploy (the deployment tooling)
deploy/the platform's deployment: compose (and Podman's override), config, scripts; argus-standalone/
tools/development and operations tools: dn, the test GitLab, pack builds
clients/MCP configurations for Claude Code, Qwen Code, Continue, DeepSeek Harness and others (the app's Connect your tools page has each tool's full setup); the arena CLI and a GitLab CI template that reviews merge requests (docs/ci.md)
docs/the documentation
evals/the evaluation question sets and results

Documentation

Start at docs/README.md. The most used:

Contributing and security

See CONTRIBUTING.md for how to build, test and propose a change, and SECURITY.md for reporting a vulnerability.

Licence

GPL v3, see LICENSE. Knowledge packs carry their own upstream licences, which are not GPL and vary per pack (CC-BY-4.0 for the Microsoft documentation, CC-BY-SA-3.0 for cppreference, PSF-2.0 for Python, public domain for SQLite); argus pack info <name> prints each in full.

Binchitects/argus

Deep Code Analysis & Review Skill

C#

0

578 commits

updated Oct 4, 2026

See the code

README

Argus Arena

A private LLM platform for a team, on your own hardware: a chat with tools, an API with a key and a budget per person, and Argus, a code index that gives the model your codebase and real documentation, within each person's GitLab permissions. Nothing leaves your network.

What you get

  • A chat for everyone: the models you run, thinking levels, branches and edits, files and Office documents, and tools that run on the server: Argus (your code and documentation), Python in a sandbox with no network (data, image, map and chart libraries), the web (sites you allow), pictures, video, speech, a calculator and dates. Sound and video in and out: voice messages and attached sound or video go to a model that hears and sees (Qwen3-Omni) or as a transcript and frames to one that does not; answers are read aloud. Tools ask before they run when you want them to. Pages, pictures, diagrams and React components the model writes run live in a sandboxed preview.
  • An API for editors, agents and scripts: OpenAI-compatible (and Anthropic's and OpenAI's Responses API), a key per person, spend and budgets per person, and fair use when the GPU is shared. Connect your tools walks each person through Claude Code, Codex, Qwen Code, OpenCode, Aider, Hermes, OpenClaw, DeepSeek Harness, Continue, Cline and more.
  • Sign-in for the whole stack: accounts, LDAP or Active Directory, two-factor, single sign-on for the services that have their own login, groups that decide who may use which model and tool.
  • Administration: people and groups; models read for what they are, several loaded at once (kept loaded, or loaded on request), each on the GPUs chosen for it, and models on other GPU servers behind the same gateway; picture, video and speech models with the same controls; every setting in one place, the audit log, the code index and knowledge packs.
  • Observability in the app: ten dashboards, every service's logs, and the alerts, what fires now and what fired before.
  • Argus, the code index: it mirrors your GitLab, extracts symbols and dependencies, serves eleven knowledge packs of API documentation, and answers over MCP. Each developer sees exactly the repositories their GitLab account can read, enforced in SQL.

Why Argus

Ten task families, one question each, graded on facts checked against the corpus before any model ran:

modelalonewith Argus
qwen3.6:27b (dense, 27.8B)5 / 1010 / 10
qwen3.6:35b (MoE, 36.0B)5 / 109 / 10
qwen3.8:27b (dense, 27.3B)5 / 99 / 10

All three models failed the same five tasks alone: facts too specialised to sit in any local model's weights (a driver's IRQL, the header CreateFileW is really declared in). Scale did not fix it; retrieval did, and answers got faster (5.5 s to 2.4 s median). The details are in docs/measurements/model-comparison.md.

Architecture

flowchart LR
  person([People and tools]) --> traefik[Traefik<br/>TLS, routing]
  traefik --> web[Web<br/>src/web]
  traefik --> api[API<br/>src/Llm.Api]
  traefik --> gateway[LiteLLM<br/>keys, budgets]
  api --> gateway --> engine[llama.cpp<br/>the models, on the GPUs]
  gateway --> remote[(Other GPU servers)]
  api --> argus[Argus<br/>src/Argus]
  api --> sandbox[Python sandbox<br/>no network]
  gateway --> media[Pictures, video,<br/>speech servers]
  api --> obs[Prometheus, Loki,<br/>Alertmanager]
  argus --> gitlab[(Your GitLab)]
  api --> pg[(Postgres)]

Every service, network and failure mode is in docs/architecture.md.

Quick start

On a Linux host with Docker (or rootless Podman) and an NVIDIA GPU:

git clone https://github.com/Binchitects/argus && cd argus/deploy
cp .env.example .env      # the domain, the models folder, the first model, six secrets
docker compose up -d

Open https://llm.localhost and sign in as admin with the password from .env. The app fetches the first model and the picture, video and speech models by itself. Every module runs; one is left out in a line. The full walkthrough, Podman and upgrading from v3 are in docs/deployment.md.

Argus alone, without the platform (its own small app, no sign-in service or observability): deploy/argus-standalone/.

Repository

pathwhat it is
src/Llm.Api and Llm.Core (the platform's .NET API), Argus (the code index service), web (the platform's React app), argus-web (Argus's own app)
tests/Llm.Tests and Argus.Tests (xUnit), deploy (the deployment tooling)
deploy/the platform's deployment: compose (and Podman's override), config, scripts; argus-standalone/
tools/development and operations tools: dn, the test GitLab, pack builds
clients/MCP configurations for Claude Code, Qwen Code, Continue, DeepSeek Harness and others (the app's Connect your tools page has each tool's full setup); the arena CLI and a GitLab CI template that reviews merge requests (docs/ci.md)
docs/the documentation
evals/the evaluation question sets and results

Documentation

Start at docs/README.md. The most used:

Contributing and security

See CONTRIBUTING.md for how to build, test and propose a change, and SECURITY.md for reporting a vulnerability.

Licence

GPL v3, see LICENSE. Knowledge packs carry their own upstream licences, which are not GPL and vary per pack (CC-BY-4.0 for the Microsoft documentation, CC-BY-SA-3.0 for cppreference, PSF-2.0 for Python, public domain for SQLite); argus pack info <name> prints each in full.

Languages

C#

59.0%

TypeScript

30.7%

Python

7.7%

Shell

1.5%