KeyserDSoze/LlmProxy

Proxy for LLM

C#

0

766 commits

updated Oct 4, 2026

See the code

README

LlmProxy

Managed hardware: administrators can prepare Linux inference hosts with the LlmProxy Node Agent, inspect CPU/RAM/disk/GPU capacity, evaluate deployable model fit, and control vLLM model lifecycle from the Admin UI. See docs/model-hardware-management.md.

Enterprise OpenAI-compatible gateway for routing GitHub Copilot and other AI clients to on-premises LLMs running on NVIDIA DGX infrastructure.

Immutable distribution releases are generated automatically from validated main pushes, starting at v0.0.1. The legacy 0.2.0-preview.* values remain source-history metadata, not the automatic GitHub Release counter.

What this product is

LlmProxy is the control and governance boundary between AI clients and a physical inference fleet. Clients see one stable OpenAI-compatible endpoint and logical model names; the gateway resolves credentials/policies, selects an eligible DGX/model deployment, enforces distributed admission/governance and streams the response.

The supported Linux production topology is intentionally single-host for the control plane while DGX/vLLM remains on the private LAN:

GitHub Copilot / OpenAI-compatible clients
                  |
          optional Cloudflare Tunnel
                  |
+------------------------------------------------------+
| Linux production host                                |
|                                                      |
| LlmProxy + PostgreSQL + Redis                        |
| OpenTelemetry Collector                              |
| Prometheus + Tempo + Loki + Grafana                  |
+-----------------------+------------------------------+
                        |
                        | private LAN
                        v
                  DGX Spark / vLLM

PostgreSQL is durable truth, Redis provides shared runtime/coordination state, and local RAM remains the request-path configuration L1.

Core capabilities

  • OpenAI-compatible /v1/models, Chat Completions and Responses APIs.
  • Optional authenticated /v1/systemone proxy for Jev-compatible decision/classification runtimes such as Laya.
  • Incremental SSE streaming and cancellation.
  • Logical public model aliases with internal DGX/provider model identifiers.
  • Weighted least loaded, round robin and weighted round robin routing.
  • Health hysteresis and safe drain/resume maintenance.
  • Distributed physical-capacity admission with Redis leases.
  • HMAC-backed bearer credentials with administrator-recoverable encrypted secret copies for newly created/rotated keys.
  • Entra-authenticated end-user dashboard with automatic or administrator-censused user provisioning, central disable/re-enable, personal API keys, recent calls and own usage.
  • Stable Entra tid + oid user identity, Usage Groups, per-credential request-rate governance, aggregate Entra-user request quotas and output-token budgets.
  • Historical PostgreSQL usage rollups beyond raw-metric retention.
  • Transactional PostgreSQL -> Redis runtime-state outbox.
  • Metadata-only metrics/audit/OTEL plus a separate application-encrypted full-body Request Audit: administrators can inspect all retained calls, users can inspect only calls from their own personal API keys, and retention is configurable from 10 through 4015 days (11 years).
  • PostgreSQL backup/restore operators.
  • Automatic immutable SemVer releases from every green main push, with build identity, GitHub release notes, multi-arch GHCR digest evidence, SPDX SBOM and SLSA provenance.
  • Executable production environment acceptance for Linux host, direct DGX/vLLM and gateway Chat/Responses/SSE surfaces.

Repository structure

src/                    product code (.NET 10 + React/TypeScript)
tests/                  unit, frontend, integration and performance tests
docker/                 images, Compose, observability and operator scripts
docs/                   architecture/deployment/operations documentation
.github/workflows/       CI, publication, deployment and environment acceptance
AGENTS.md                mandatory engineering handover entry point
CHANGELOG.md             product-visible release history

Development quickstart

For the distributed Redis/observability development/demo bundle use:

bash docker/scripts/full-stack-init.sh
# edit docker/.env.full, especially DGX_NODE_BASE_ADDRESS and PROVIDER_MODEL_NAME
docker compose --env-file docker/.env.full -f docker/docker-compose.full.yml up -d

The generic full-stack example is intended for development/demo/acceptance setup. Production has a separate operator-owned environment file and deployment path.

Linux production deployment

For immutable GitHub Release installation/update and the llmproxyctl operator command, start with:

docs/release-installation.md

The canonical production runbook is:

docs/linux-production-deployment.md

Production uses the Redis-enabled full stack, not the legacy minimal overlay.

Preferred first installation

From a repository checkout on a new Linux host:

export GHCR_USER='<github-user>'
export GHCR_TOKEN='<token-with-package-read-access>'
export ENTRA_ENABLED=true
export ENTRA_TENANT_ID='<tenant-id>'
export ENTRA_CLIENT_ID='<client-id>'
export ENTRA_CLIENT_SECRET='<client-secret>'

sudo -E bash docker/scripts/install-linux.sh \
  --dgx-url http://10.0.0.21:8000 \
  --provider-model '<exact-vllm-model-id>' \
  --image-tag sha-df3ecf7

docker/scripts/install-linux.sh detects the distro/package manager, installs or preserves Docker Engine, ensures Docker Compose v2, prepares /opt/llmproxy, generates initial production secrets, optionally logs into GHCR, checks DGX /health and /v1/models, then invokes the canonical full-stack deployment.

Docker official repositories are used for Debian, Ubuntu, Fedora, CentOS and RHEL. Common derivative/other distributions can use apt, dnf/yum, zypper, pacman or apk; when Compose v2 is missing the installer has a CLI-plugin fallback. Existing Docker installations are preserved.

For policy-controlled hosts use --prepare-only or pre-install Docker and rerun with --skip-docker-install. --validate-only performs a no-change installer/repository compatibility check.

Generated passwords/API credential/pepper are not printed. The protected operator-owned configuration is stored at:

/opt/llmproxy/.env

Back up both LLM_PROXY_API_KEY_PEPPER and LLMPROXY_UPSTREAM_CREDENTIAL_KEY separately before treating the host as production.

Updating an existing Linux installation

If LlmProxy was installed through the release/bootstrap installer, the machine has the llmproxyctl operator command available globally.

Newer releases also install a persistent LlmProxy Update Agent on the control-plane host. Administrators can open Release Notes & Updates in the Admin UI to see the installed version, discover subsequent immutable releases, inspect each release's update plan/operator command, run Update now, or schedule the update for a later date/time. See docs/update-management.md.

Check the currently installed version and health:

llmproxyctl version
llmproxyctl status
llmproxyctl health

Update to a specific immutable release:

sudo -E llmproxyctl update 0.0.3

Replace 0.0.3 with the release version you want to install. The command is llmproxyctl update (not llmproxy update).

The update preserves the protected host configuration and Docker data:

/opt/llmproxy/.env
Docker volumes (PostgreSQL, Redis and observability data)

If GHCR requires authentication, export the credentials before running the update so sudo -E can pass them through:

export GHCR_USER='<github-user>'
export GHCR_TOKEN='<token-with-package-read-access>'
sudo -E llmproxyctl update 0.0.3

The existing super-administrator list is preserved automatically. To replace it during the update:

sudo -E llmproxyctl update 0.0.3 \
  --super-admins 'admin1@example.com;admin2@example.com'

After the update:

llmproxyctl version
llmproxyctl doctor
llmproxyctl health

If the update fails or remains in readiness checks, inspect the persistent installer log:

sudo tail -f /var/log/llmproxy/latest-install.log

If an older installed llmproxyctl fails with Unknown option: --skip-dgx-check, it is using the legacy update flag while the target release expects the hardware-neutral name. For release 0.0.6, bypass the old wrapper once and invoke the installed bootstrap helper with the current flag:

sudo -E /usr/local/lib/llmproxy/bootstrap.sh \
  --version 0.0.6 \
  --skip-docker-install \
  --skip-node-check

After that succeeds, release 0.0.6 installs the current llmproxyctl, so subsequent updates use --skip-node-check automatically. Newer installers also accept the legacy --skip-dgx-check and --dgx-url spellings as deprecated compatibility aliases.

For rollback to a release already installed locally:

sudo llmproxyctl rollback 0.0.2

See docs/release-installation.md for the complete release/update/rollback procedure.

Manual/redeployment path

After first host preparation, or when provisioning manually, deploy a published image with:

LLMPROXY_DEPLOY_DIR=/opt/llmproxy \
LLMPROXY_ENV_FILE=/opt/llmproxy/.env \
  bash docker/scripts/deploy.sh sha-df3ecf7

For controlled production changes prefer an immutable sha-<7> alias or an exact SemVer tag rather than mutable main.

deploy.sh validates production settings, refuses public Cloudflare exposure until Entra is configured, stages runtime assets under /opt/llmproxy/runtime, validates Compose, starts the full stack and requires both /healthz and /readyz.

Cloudflare is optional. Leave CLOUDFLARE_TUNNEL_TOKEN blank for private-LAN bootstrap.

Automated production deployment

.github/workflows/deploy.yml is the supported GitHub Actions deployment path after the host exists. It runs on a dedicated Linux self-hosted runner labelled:

self-hosted
linux
x64
llmproxy-prod

The runner keeps production secrets in /opt/llmproxy/.env; secrets are not committed to Git. The workflow uses the same docker/scripts/deploy.sh as manual deployment.

Production environment acceptance

0.2.0-preview.5 adds an executable target-host acceptance harness:

sudo -E bash docker/scripts/environment-acceptance.sh

It validates:

  • Linux/Docker/Compose host prerequisites;
  • direct VM -> DGX/vLLM /health, /v1/models, Chat, Responses and SSE;
  • LlmProxy /healthz, /readyz, /v1/models, Chat, Responses and SSE;
  • exact provider model and logical public model visibility.

The evidence bundle contains only metadata in summary.md and checks.tsv. Request bodies, prompts, source, generated output, response bodies and bearer/API secrets are not persisted. Canonical vLLM /health is treated as a status-only endpoint because a healthy vLLM server may return HTTP 200 with an empty body.

Read:

docs/environment-acceptance.md

After the llmproxy-prod self-hosted runner is installed, the same acceptance is manually launchable through:

.github/workflows/environment-acceptance.yml

The workflow accepts no API-key inputs. It reads the protected host configuration, uploads only summary.md and checks.tsv as a short-lived Actions artifact, and removes the runner-local evidence afterward.

This proves connectivity and functional compatibility. It does not establish production concurrency; real DGX/model Capacity Profiles still require benchmark evidence.

DGX service roots

A node stores the complete inference service root, including optional path prefix:

http://10.0.0.25:8000
http://10.0.0.25:8000/vllm
https://dgx-01.internal:8443/inference

LlmProxy derives:

<root>/health
<root>/v1/models
<root>/v1/chat/completions
<root>/v1/responses

Public API

GET  /v1/models
POST /v1/chat/completions
POST /v1/responses
POST /v1/systemone      # optional Jev-compatible classifier proxy
GET  /healthz
GET  /readyz

All /v1 surfaces use bearer credentials. /v1/systemone is a Jev-compatible classifier contract, not an OpenAI generative endpoint, and is enabled only when a private System One upstream is configured. Newly created/rotated client credentials keep an application-encrypted recovery copy so administrators can reveal/copy them later; authentication still uses the HMAC hash.

Administration

The React control plane manages nodes, models, deployments, routing, credentials, governance, usage, maintenance, runtime synchronization and audit. It also includes a model/System One playground, live administrator-only request/response payload logs, endpoint examples and a collapsed documentation accordion on each screen.

Production administration is designed for Entra ID with roles:

LlmProxy.Admin
LlmProxy.User
LlmProxy.Reader

LlmProxy.User uses /admin/me to create, rotate and revoke personal API keys, inspect own usage and see read-only aggregate user request limits. Administrators configure aggregate user request quotas from Usage & Governance. LlmProxy.Reader remains an operational read-only role.

Do not expose administrative surfaces publicly before Entra is configured and validated.

Persistence and runtime state

PostgreSQL = durable configuration/history + runtime-state outbox + usage rollups
Redis      = distributed L2 + request/token/capacity/maintenance coordination
local RAM  = per-replica request-path configuration L1

Ordinary inference configuration lookups are DB-free after startup/runtime publication.

Backup and recovery

PostgreSQL is the recovery authority; Redis is rebuildable. Bash and PowerShell backup/restore operators live under docker/scripts/.

Authentication__ApiKeyPepper and deployment secrets are external recovery dependencies and must be preserved separately from database backups.

Read docs/backup-restore.md before production restore work.

CI/CD and supply-chain evidence

Pushes/PRs execute backend, frontend, Docker/PostgreSQL and operational smokes. A container is published only after successful CI for the same main source SHA.

Published images include source/version/build identity. The publish workflow records the immutable image digest and verifies registry-native SPDX SBOM and SLSA/BuildKit provenance attestations.

Validated 0.2.0-preview.7 runtime checkpoint:

version                0.2.0-preview.7
source                 df3ecf7cb4ab6a6ff99fa6ea21b1169c44f15a38
CI                     35592623906 SUCCESS
Full Stack             35592624282 SUCCESS
Publish GHCR           35593081824 SUCCESS
image alias            sha-df3ecf7
image digest           sha256:de82c1b7fa29b6d0b7104b1e5960316b6eeea81cf85a9d23c4fcc53ac2ae4d99
attestation manifest   sha256:0da97b9a569aa974e9d77b5dd18d62082cde063fbf87221a908dc70d84fe60b8
release artifact       10635322261
artifact digest        sha256:d0884b5f48e2ecf00f55a0e52d153131b827f880f41306b89b9a31e8cd93e51b

No immutable v0.2.0-preview.7 Git tag or GitHub Release has been created.

Important production configuration

Never commit production values for:

POSTGRES_PASSWORD
REDIS_PASSWORD
LLM_PROXY_API_KEY
LLM_PROXY_API_KEY_PEPPER
GRAFANA_ADMIN_PASSWORD
ENTRA_TENANT_ID
ENTRA_CLIENT_ID
ENTRA_CLIENT_SECRET
CLOUDFLARE_TUNNEL_TOKEN
SYSTEM_ONE_API_KEY
LLMPROXY_ACCEPTANCE_DGX_API_KEY

Use docker/.env.production.example as the manual production template; the Linux installer creates the equivalent host-owned file automatically when it does not already exist.

Documentation map

Start with:

  • docs/release-installation.md — immutable GitHub Release bundle, bootstrap, llmproxyctl, update and rollback.
  • docs/update-management.md — Admin update-now/scheduling workflow, host Update Agent and per-release standard/custom update plans.
  • docs/linux-production-deployment.md — canonical zero-to-running Linux production runbook.
  • docs/environment-acceptance.md — production host/DGX/gateway acceptance and evidence rules.
  • docs/deployment.md — deployment contract and automation summary.
  • docs/full-stack.md — Redis + observability bundle details.
  • docs/operations.md — health, maintenance, release identity and audit.
  • docs/backup-restore.md — recovery procedures.
  • docs/github-copilot.md — Copilot/BYOK integration and limitations.
  • docs/capacity-control.md and docs/benchmarking.md — admission and benchmark calibration.
  • docs/project-status.md — canonical validated engineering checkpoint.
  • AGENTS.md — mandatory engineering resume protocol.

External acceptance still required

Repository automation cannot replace environment validation for:

  • actual package/repository behavior on the chosen Linux distro/version;
  • real DGX Spark/vLLM/model acceptance and benchmark sweeps;
  • representative multi-DGX coding load;
  • real Entra app/role setup;
  • real Cloudflare hostname/tunnel routing;
  • GitHub Copilot BYOK end-to-end;
  • self-hosted runner permissions/reboot behavior;
  • customer backup destination/encryption/retention;
  • customer-specific PostgreSQL/Redis/observability HA and durable storage choices.

License

Internal project. Licensing and external distribution terms will be defined before productization.

Identity and personal API-key ownership are defined in docs/identity-api-keys.md.

KeyserDSoze/LlmProxy

Proxy for LLM

C#

0

766 commits

updated Oct 4, 2026

See the code

README

LlmProxy

Managed hardware: administrators can prepare Linux inference hosts with the LlmProxy Node Agent, inspect CPU/RAM/disk/GPU capacity, evaluate deployable model fit, and control vLLM model lifecycle from the Admin UI. See docs/model-hardware-management.md.

Enterprise OpenAI-compatible gateway for routing GitHub Copilot and other AI clients to on-premises LLMs running on NVIDIA DGX infrastructure.

Immutable distribution releases are generated automatically from validated main pushes, starting at v0.0.1. The legacy 0.2.0-preview.* values remain source-history metadata, not the automatic GitHub Release counter.

What this product is

LlmProxy is the control and governance boundary between AI clients and a physical inference fleet. Clients see one stable OpenAI-compatible endpoint and logical model names; the gateway resolves credentials/policies, selects an eligible DGX/model deployment, enforces distributed admission/governance and streams the response.

The supported Linux production topology is intentionally single-host for the control plane while DGX/vLLM remains on the private LAN:

GitHub Copilot / OpenAI-compatible clients
                  |
          optional Cloudflare Tunnel
                  |
+------------------------------------------------------+
| Linux production host                                |
|                                                      |
| LlmProxy + PostgreSQL + Redis                        |
| OpenTelemetry Collector                              |
| Prometheus + Tempo + Loki + Grafana                  |
+-----------------------+------------------------------+
                        |
                        | private LAN
                        v
                  DGX Spark / vLLM

PostgreSQL is durable truth, Redis provides shared runtime/coordination state, and local RAM remains the request-path configuration L1.

Core capabilities

  • OpenAI-compatible /v1/models, Chat Completions and Responses APIs.
  • Optional authenticated /v1/systemone proxy for Jev-compatible decision/classification runtimes such as Laya.
  • Incremental SSE streaming and cancellation.
  • Logical public model aliases with internal DGX/provider model identifiers.
  • Weighted least loaded, round robin and weighted round robin routing.
  • Health hysteresis and safe drain/resume maintenance.
  • Distributed physical-capacity admission with Redis leases.
  • HMAC-backed bearer credentials with administrator-recoverable encrypted secret copies for newly created/rotated keys.
  • Entra-authenticated end-user dashboard with automatic or administrator-censused user provisioning, central disable/re-enable, personal API keys, recent calls and own usage.
  • Stable Entra tid + oid user identity, Usage Groups, per-credential request-rate governance, aggregate Entra-user request quotas and output-token budgets.
  • Historical PostgreSQL usage rollups beyond raw-metric retention.
  • Transactional PostgreSQL -> Redis runtime-state outbox.
  • Metadata-only metrics/audit/OTEL plus a separate application-encrypted full-body Request Audit: administrators can inspect all retained calls, users can inspect only calls from their own personal API keys, and retention is configurable from 10 through 4015 days (11 years).
  • PostgreSQL backup/restore operators.
  • Automatic immutable SemVer releases from every green main push, with build identity, GitHub release notes, multi-arch GHCR digest evidence, SPDX SBOM and SLSA provenance.
  • Executable production environment acceptance for Linux host, direct DGX/vLLM and gateway Chat/Responses/SSE surfaces.

Repository structure

src/                    product code (.NET 10 + React/TypeScript)
tests/                  unit, frontend, integration and performance tests
docker/                 images, Compose, observability and operator scripts
docs/                   architecture/deployment/operations documentation
.github/workflows/       CI, publication, deployment and environment acceptance
AGENTS.md                mandatory engineering handover entry point
CHANGELOG.md             product-visible release history

Development quickstart

For the distributed Redis/observability development/demo bundle use:

bash docker/scripts/full-stack-init.sh
# edit docker/.env.full, especially DGX_NODE_BASE_ADDRESS and PROVIDER_MODEL_NAME
docker compose --env-file docker/.env.full -f docker/docker-compose.full.yml up -d

The generic full-stack example is intended for development/demo/acceptance setup. Production has a separate operator-owned environment file and deployment path.

Linux production deployment

For immutable GitHub Release installation/update and the llmproxyctl operator command, start with:

docs/release-installation.md

The canonical production runbook is:

docs/linux-production-deployment.md

Production uses the Redis-enabled full stack, not the legacy minimal overlay.

Preferred first installation

From a repository checkout on a new Linux host:

export GHCR_USER='<github-user>'
export GHCR_TOKEN='<token-with-package-read-access>'
export ENTRA_ENABLED=true
export ENTRA_TENANT_ID='<tenant-id>'
export ENTRA_CLIENT_ID='<client-id>'
export ENTRA_CLIENT_SECRET='<client-secret>'

sudo -E bash docker/scripts/install-linux.sh \
  --dgx-url http://10.0.0.21:8000 \
  --provider-model '<exact-vllm-model-id>' \
  --image-tag sha-df3ecf7

docker/scripts/install-linux.sh detects the distro/package manager, installs or preserves Docker Engine, ensures Docker Compose v2, prepares /opt/llmproxy, generates initial production secrets, optionally logs into GHCR, checks DGX /health and /v1/models, then invokes the canonical full-stack deployment.

Docker official repositories are used for Debian, Ubuntu, Fedora, CentOS and RHEL. Common derivative/other distributions can use apt, dnf/yum, zypper, pacman or apk; when Compose v2 is missing the installer has a CLI-plugin fallback. Existing Docker installations are preserved.

For policy-controlled hosts use --prepare-only or pre-install Docker and rerun with --skip-docker-install. --validate-only performs a no-change installer/repository compatibility check.

Generated passwords/API credential/pepper are not printed. The protected operator-owned configuration is stored at:

/opt/llmproxy/.env

Back up both LLM_PROXY_API_KEY_PEPPER and LLMPROXY_UPSTREAM_CREDENTIAL_KEY separately before treating the host as production.

Updating an existing Linux installation

If LlmProxy was installed through the release/bootstrap installer, the machine has the llmproxyctl operator command available globally.

Newer releases also install a persistent LlmProxy Update Agent on the control-plane host. Administrators can open Release Notes & Updates in the Admin UI to see the installed version, discover subsequent immutable releases, inspect each release's update plan/operator command, run Update now, or schedule the update for a later date/time. See docs/update-management.md.

Check the currently installed version and health:

llmproxyctl version
llmproxyctl status
llmproxyctl health

Update to a specific immutable release:

sudo -E llmproxyctl update 0.0.3

Replace 0.0.3 with the release version you want to install. The command is llmproxyctl update (not llmproxy update).

The update preserves the protected host configuration and Docker data:

/opt/llmproxy/.env
Docker volumes (PostgreSQL, Redis and observability data)

If GHCR requires authentication, export the credentials before running the update so sudo -E can pass them through:

export GHCR_USER='<github-user>'
export GHCR_TOKEN='<token-with-package-read-access>'
sudo -E llmproxyctl update 0.0.3

The existing super-administrator list is preserved automatically. To replace it during the update:

sudo -E llmproxyctl update 0.0.3 \
  --super-admins 'admin1@example.com;admin2@example.com'

After the update:

llmproxyctl version
llmproxyctl doctor
llmproxyctl health

If the update fails or remains in readiness checks, inspect the persistent installer log:

sudo tail -f /var/log/llmproxy/latest-install.log

If an older installed llmproxyctl fails with Unknown option: --skip-dgx-check, it is using the legacy update flag while the target release expects the hardware-neutral name. For release 0.0.6, bypass the old wrapper once and invoke the installed bootstrap helper with the current flag:

sudo -E /usr/local/lib/llmproxy/bootstrap.sh \
  --version 0.0.6 \
  --skip-docker-install \
  --skip-node-check

After that succeeds, release 0.0.6 installs the current llmproxyctl, so subsequent updates use --skip-node-check automatically. Newer installers also accept the legacy --skip-dgx-check and --dgx-url spellings as deprecated compatibility aliases.

For rollback to a release already installed locally:

sudo llmproxyctl rollback 0.0.2

See docs/release-installation.md for the complete release/update/rollback procedure.

Manual/redeployment path

After first host preparation, or when provisioning manually, deploy a published image with:

LLMPROXY_DEPLOY_DIR=/opt/llmproxy \
LLMPROXY_ENV_FILE=/opt/llmproxy/.env \
  bash docker/scripts/deploy.sh sha-df3ecf7

For controlled production changes prefer an immutable sha-<7> alias or an exact SemVer tag rather than mutable main.

deploy.sh validates production settings, refuses public Cloudflare exposure until Entra is configured, stages runtime assets under /opt/llmproxy/runtime, validates Compose, starts the full stack and requires both /healthz and /readyz.

Cloudflare is optional. Leave CLOUDFLARE_TUNNEL_TOKEN blank for private-LAN bootstrap.

Automated production deployment

.github/workflows/deploy.yml is the supported GitHub Actions deployment path after the host exists. It runs on a dedicated Linux self-hosted runner labelled:

self-hosted
linux
x64
llmproxy-prod

The runner keeps production secrets in /opt/llmproxy/.env; secrets are not committed to Git. The workflow uses the same docker/scripts/deploy.sh as manual deployment.

Production environment acceptance

0.2.0-preview.5 adds an executable target-host acceptance harness:

sudo -E bash docker/scripts/environment-acceptance.sh

It validates:

  • Linux/Docker/Compose host prerequisites;
  • direct VM -> DGX/vLLM /health, /v1/models, Chat, Responses and SSE;
  • LlmProxy /healthz, /readyz, /v1/models, Chat, Responses and SSE;
  • exact provider model and logical public model visibility.

The evidence bundle contains only metadata in summary.md and checks.tsv. Request bodies, prompts, source, generated output, response bodies and bearer/API secrets are not persisted. Canonical vLLM /health is treated as a status-only endpoint because a healthy vLLM server may return HTTP 200 with an empty body.

Read:

docs/environment-acceptance.md

After the llmproxy-prod self-hosted runner is installed, the same acceptance is manually launchable through:

.github/workflows/environment-acceptance.yml

The workflow accepts no API-key inputs. It reads the protected host configuration, uploads only summary.md and checks.tsv as a short-lived Actions artifact, and removes the runner-local evidence afterward.

This proves connectivity and functional compatibility. It does not establish production concurrency; real DGX/model Capacity Profiles still require benchmark evidence.

DGX service roots

A node stores the complete inference service root, including optional path prefix:

http://10.0.0.25:8000
http://10.0.0.25:8000/vllm
https://dgx-01.internal:8443/inference

LlmProxy derives:

<root>/health
<root>/v1/models
<root>/v1/chat/completions
<root>/v1/responses

Public API

GET  /v1/models
POST /v1/chat/completions
POST /v1/responses
POST /v1/systemone      # optional Jev-compatible classifier proxy
GET  /healthz
GET  /readyz

All /v1 surfaces use bearer credentials. /v1/systemone is a Jev-compatible classifier contract, not an OpenAI generative endpoint, and is enabled only when a private System One upstream is configured. Newly created/rotated client credentials keep an application-encrypted recovery copy so administrators can reveal/copy them later; authentication still uses the HMAC hash.

Administration

The React control plane manages nodes, models, deployments, routing, credentials, governance, usage, maintenance, runtime synchronization and audit. It also includes a model/System One playground, live administrator-only request/response payload logs, endpoint examples and a collapsed documentation accordion on each screen.

Production administration is designed for Entra ID with roles:

LlmProxy.Admin
LlmProxy.User
LlmProxy.Reader

LlmProxy.User uses /admin/me to create, rotate and revoke personal API keys, inspect own usage and see read-only aggregate user request limits. Administrators configure aggregate user request quotas from Usage & Governance. LlmProxy.Reader remains an operational read-only role.

Do not expose administrative surfaces publicly before Entra is configured and validated.

Persistence and runtime state

PostgreSQL = durable configuration/history + runtime-state outbox + usage rollups
Redis      = distributed L2 + request/token/capacity/maintenance coordination
local RAM  = per-replica request-path configuration L1

Ordinary inference configuration lookups are DB-free after startup/runtime publication.

Backup and recovery

PostgreSQL is the recovery authority; Redis is rebuildable. Bash and PowerShell backup/restore operators live under docker/scripts/.

Authentication__ApiKeyPepper and deployment secrets are external recovery dependencies and must be preserved separately from database backups.

Read docs/backup-restore.md before production restore work.

CI/CD and supply-chain evidence

Pushes/PRs execute backend, frontend, Docker/PostgreSQL and operational smokes. A container is published only after successful CI for the same main source SHA.

Published images include source/version/build identity. The publish workflow records the immutable image digest and verifies registry-native SPDX SBOM and SLSA/BuildKit provenance attestations.

Validated 0.2.0-preview.7 runtime checkpoint:

version                0.2.0-preview.7
source                 df3ecf7cb4ab6a6ff99fa6ea21b1169c44f15a38
CI                     35592623906 SUCCESS
Full Stack             35592624282 SUCCESS
Publish GHCR           35593081824 SUCCESS
image alias            sha-df3ecf7
image digest           sha256:de82c1b7fa29b6d0b7104b1e5960316b6eeea81cf85a9d23c4fcc53ac2ae4d99
attestation manifest   sha256:0da97b9a569aa974e9d77b5dd18d62082cde063fbf87221a908dc70d84fe60b8
release artifact       10635322261
artifact digest        sha256:d0884b5f48e2ecf00f55a0e52d153131b827f880f41306b89b9a31e8cd93e51b

No immutable v0.2.0-preview.7 Git tag or GitHub Release has been created.

Important production configuration

Never commit production values for:

POSTGRES_PASSWORD
REDIS_PASSWORD
LLM_PROXY_API_KEY
LLM_PROXY_API_KEY_PEPPER
GRAFANA_ADMIN_PASSWORD
ENTRA_TENANT_ID
ENTRA_CLIENT_ID
ENTRA_CLIENT_SECRET
CLOUDFLARE_TUNNEL_TOKEN
SYSTEM_ONE_API_KEY
LLMPROXY_ACCEPTANCE_DGX_API_KEY

Use docker/.env.production.example as the manual production template; the Linux installer creates the equivalent host-owned file automatically when it does not already exist.

Documentation map

Start with:

  • docs/release-installation.md — immutable GitHub Release bundle, bootstrap, llmproxyctl, update and rollback.
  • docs/update-management.md — Admin update-now/scheduling workflow, host Update Agent and per-release standard/custom update plans.
  • docs/linux-production-deployment.md — canonical zero-to-running Linux production runbook.
  • docs/environment-acceptance.md — production host/DGX/gateway acceptance and evidence rules.
  • docs/deployment.md — deployment contract and automation summary.
  • docs/full-stack.md — Redis + observability bundle details.
  • docs/operations.md — health, maintenance, release identity and audit.
  • docs/backup-restore.md — recovery procedures.
  • docs/github-copilot.md — Copilot/BYOK integration and limitations.
  • docs/capacity-control.md and docs/benchmarking.md — admission and benchmark calibration.
  • docs/project-status.md — canonical validated engineering checkpoint.
  • AGENTS.md — mandatory engineering resume protocol.

External acceptance still required

Repository automation cannot replace environment validation for:

  • actual package/repository behavior on the chosen Linux distro/version;
  • real DGX Spark/vLLM/model acceptance and benchmark sweeps;
  • representative multi-DGX coding load;
  • real Entra app/role setup;
  • real Cloudflare hostname/tunnel routing;
  • GitHub Copilot BYOK end-to-end;
  • self-hosted runner permissions/reboot behavior;
  • customer backup destination/encryption/retention;
  • customer-specific PostgreSQL/Redis/observability HA and durable storage choices.

License

Internal project. Licensing and external distribution terms will be defined before productization.

Identity and personal API-key ownership are defined in docs/identity-api-keys.md.

Languages

C#

59.3%

TypeScript

24.9%

Shell

14.4%