A lightweight, highly secure AI API Gateway/Proxy written in Go. Acts as transparent middleware between local AI coding clients (OpenCode/Pi/Cursor) and upstream LLM providers (Gemini, DeepSeek, Zhipu z.ai).
28
stars
571
commits
Go
primary language
Sep 11, 2026
updated
AI coding clients transmit your source code, prompts, and credentials to cloud LLM providers on every request. Nenya is the gatekeeper in between: a lightweight, zero-dependency API gateway that redacts secrets before they leave your machine, keeps context payloads small, and routes across providers with fallback, caching, and transparent SSE streaming. Security-hardened: non-root execution, mlock for secrets, seccomp + no-new-privileges.
Compatible with any provider that implements the OpenAI Or Anthropic Chat Completions API. For 23 providers we ship built-in adapters with specialized handling.
Create minimal config and secrets:
mkdir -p config secrets
cat > config/config.json << 'EOF'
{
"server": { "listen_addr": ":8080" },
"agents": {
"default": {
"strategy": "fallback",
"models": ["gemini-2.5-flash"]
}
}
}
EOF
cat > secrets/provider_keys.json << 'EOF'
{
"provider_keys": {
"gemini": "AIza..."
}
}
EOF
cat > secrets/client.json << EOF
{
"client_token": "nk-$(openssl rand -hex 32)"
}
EOF
Note: the last heredoc is unquoted on purpose, so $(openssl rand -hex 32) expands once while the file is written and your token is unique.
Run the container (the same flags work with docker run):
podman run -d \
--name nenya \
-p 8080:8080 \
-v ./config:/etc/nenya:ro \
-v ./secrets:/run/secrets/nenya:ro \
-e NENYA_SECRETS_DIR=/run/secrets/nenya \
--cap-drop=ALL \
--cap-add=IPC_LOCK \
--security-opt=no-new-privileges:true \
--read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64M \
ghcr.io/gumieri/nenya:latest
Test it — the authenticated smoke test must list your configured model:
export NK=$(jq -r '.client_token' secrets/client.json)
curl -s -H "Authorization: Bearer $NK" http://localhost:8080/v1/models | jq -r '.data[].id'
Then send your first completion (streams an SSE response):
curl -N -H "Authorization: Bearer $NK" \
-d '{"model":"gemini-2.5-flash","messages":[{"role":"user","content":"Say hi in five words"}]}' \
http://localhost:8080/v1/chat/completions
No API key yet? The offline redaction demo runs Nenya against a local mock upstream — no external calls, no keys. It generates dummy secrets, sends fake AWS/GitHub credentials through the gateway, and shows what the upstream actually receives. Regenerate the GIF above with mise run demo.
Nenya provides native packages for major Linux distributions and community package managers:
| Distribution | Command |
|---|---|
| Debian/Ubuntu (.deb) | Download nenya_<version>_linux_amd64.deb from the release page and run sudo dpkg -i |
| Fedora/RHEL (.rpm) | Download nenya-<version>.x86_64.rpm from the release page and run sudo rpm -i |
| Arch Linux (.pkg.tar.zst) | Download nenya-<version>-x86_64.pkg.tar.zst from the release page and run sudo pacman -U |
| Arch Linux (AUR) | yay -S nenya-bin (or your preferred AUR helper) |
| Nix/NixOS | Add gumieri/nur-packages to your NUR registry and use nenya |
All packages install the binary to /usr/bin/nenya and include systemd service and socket units. After install, enable and start:
sudo systemctl enable --now nenya.socket
sudo systemctl enable --now nenya.service
flowchart TD
CLIENT["Client<br/>Cursor / OpenCode / Aider / etc.<br/>POST /v1/chat/completions · /v1/messages<br/>Bearer token"]
subgraph GW["Nenya Gateway"]
direction TB
AUTH["Auth + RBAC"]
RESOLVE["Parse body · resolve agent + targets<br/>strategy: fallback · round-robin · sticky"]
CACHE{"Response cache"}
MCPINJ["MCP auto-search + tool injection"]
end
CHAIN["Interceptor chain (best-effort)<br/>redact → entropy → TF-IDF → bouncer"]
TRIM["Token budget trim (hard limit)"]
subgraph LOOP["Dispatch loop — per target"]
direction TB
GUARDS["Circuit breaker · rate limits · cost guard"]
MODES["A standard forward<br/>B MCP multi-turn tool loop<br/>C context-limit retry"]
end
UPSTREAM["Upstream LLM providers<br/>23 built-in adapters"]
subgraph SSE["SSE pipeline"]
direction TB
PROBE["Stream-head probe (pre-header)<br/>empty / early-error failover"]
XFORM["Adapter transforms · format conversion"]
WATCH["Stall watchdog · stream continuation"]
ACCT["Usage accounting · cache capture · MCP auto-save"]
end
OUT["Client receives transparent SSE"]
CLIENT --> AUTH --> RESOLVE --> CACHE
CACHE -- "HIT → replay" --> OUT
CACHE -- "miss" --> MCPINJ --> CHAIN --> TRIM --> GUARDS --> MODES --> UPSTREAM
UPSTREAM --> PROBE --> XFORM --> WATCH --> ACCT --> OUT
PROBE -. "failover → next target" .-> GUARDS
classDef io fill:#e8ebf0,stroke:#57606a,color:#1f2328
classDef gw fill:#ddf4ff,stroke:#0969da,color:#1f2328
classDef pipe fill:#fff8c5,stroke:#9a6700,color:#1f2328
classDef sse fill:#dafbe1,stroke:#1a7f37,color:#1f2328
class CLIENT,OUT io
class GW,LOOP gw
class CHAIN,TRIM pipe
class SSE sse
Flow notes:
/v1/* endpoints require client bearer auth; /healthz, /statsz, /metrics do not.format attributeLimitMEMLOCK=infinity and LimitCORE=0 in systemd/tmpsystemctl reload nenya for zero-downtime config changeserror_kind field for programmatic diagnosticsAll /v1/* endpoints require Authorization: Bearer <client_token> or Bearer <api_key_token>.
API keys support RBAC enforcement — agent scoping, endpoint allowlists, role-based permissions (admin bypasses all checks).
| Endpoint | Auth | Description |
|---|---|---|
POST /v1/chat/completions | Bearer + RBAC | OpenAI-compatible chat with SSE streaming, agent fallback, MCP multi-turn |
POST /v1/messages | Bearer + RBAC | Anthropic Messages API with bidirectional format conversion |
GET /v1/models | Bearer + RBAC | Live model catalog from discovered providers + static registry (context window, max tokens) |
POST /v1/embeddings | Bearer + RBAC | Passthrough proxy |
POST /v1/responses | Bearer + RBAC | Passthrough proxy |
POST /v1/images/generations | Bearer + RBAC | Image generation (OpenAI-compatible) |
POST /v1/audio/transcriptions | Bearer + RBAC | Audio transcription (Whisper-compatible, multipart support) |
POST /v1/audio/speech | Bearer + RBAC | Text-to-speech synthesis (OpenAI-compatible) |
POST /v1/moderations | Bearer + RBAC | Content moderation (OpenAI-compatible) |
POST /v1/rerank | Bearer + RBAC | Re-ranking API (Cohere/Jina/Voyage-compatible) |
POST /v1/a2a | Bearer + RBAC | Agent-to-Agent protocol (Google A2A) |
GET/POST/DELETE /v1/files | Bearer + RBAC | File listing, upload, retrieval, deletion |
POST/GET /v1/batches | Bearer + RBAC | Batch API operations |
POST /proxy/{provider}/* | Bearer + RBAC | Arbitrary provider endpoint passthrough (all HTTP methods, SSE streaming) |
GET /healthz | None | Engine health probe |
GET /statsz | None | Token usage, circuit breaker state, MCP server status |
GET /metrics | None | Prometheus-compatible metrics |
GET /debug/pprof/* | Bearer | Go profiling endpoints (disabled by default, see debug.pprof_enabled) |
See docs/PASSTHROUGH_PROXY.md for detailed passthrough proxy usage.
Nenya supports standard environment variables for deployment portability:
| Variable | Default | Description |
|---|---|---|
PORT | 8080 | Listening port (overrides server.listen_addr) |
HOST | — | Optional bind address (e.g. 127.0.0.1). Only used when combined with PORT |
NENYA_CONFIG_DIR | /etc/nenya/ | Configuration directory path |
NENYA_CONFIG_FILE | — | Single config file path (takes precedence over NENYA_CONFIG_DIR) |
NENYA_SECRETS_DIR | — | Secrets directory (overrides CREDENTIALS_DIRECTORY) |
Example usage:
PORT=9090 HOST=127.0.0.1 ./nenya --config /path/to/config.json
Or in Docker:
docker run -e PORT=9090 -p 9090:9090 ghcr.io/gumieri/nenya:latest
| Document | Description |
|---|---|
| Providers | All 23 providers, capabilities matrix, special behaviors, adding custom providers |
| Configuration | Full config reference, directory mode, all sections and fields |
| Deploy Bare Metal | Systemd unit, config.d layout, secrets, hot reload |
| Deploy Container | Podman/Docker Compose, image verification, security notes |
| Deploy Kubernetes | Helm chart usage, ConfigMap/Secret, ingress setup |
| Passthrough Proxy | Raw provider endpoint proxying, SSE streaming, auth injection |
| Architecture | Package DAG, request lifecycle, circuit breaker, SSE pipeline |
| MCP Integration | MCP server integration, tool discovery, multi-turn execution |
| Adapters | Adapter system internals, auth styles, capability flags |
| Secrets Format | Systemd credentials, env var fallback, container/K8s deployment |
| Security | Vulnerability reporting policy |
| Disclaimer | Best-effort redaction scope and limitations |
| Changelog | Release history and notable changes |
Apache 2.0. See LICENSE.
559 commits
12 commits
Go
99.6%
A lightweight, highly secure AI API Gateway/Proxy written in Go. Acts as transparent middleware between local AI coding clients (OpenCode/Pi/Cursor) and upstream LLM providers (Gemini, DeepSeek, Zhipu z.ai).
28
stars
571
commits
Go
primary language
Sep 11, 2026
updated
AI coding clients transmit your source code, prompts, and credentials to cloud LLM providers on every request. Nenya is the gatekeeper in between: a lightweight, zero-dependency API gateway that redacts secrets before they leave your machine, keeps context payloads small, and routes across providers with fallback, caching, and transparent SSE streaming. Security-hardened: non-root execution, mlock for secrets, seccomp + no-new-privileges.
Compatible with any provider that implements the OpenAI Or Anthropic Chat Completions API. For 23 providers we ship built-in adapters with specialized handling.
Create minimal config and secrets:
mkdir -p config secrets
cat > config/config.json << 'EOF'
{
"server": { "listen_addr": ":8080" },
"agents": {
"default": {
"strategy": "fallback",
"models": ["gemini-2.5-flash"]
}
}
}
EOF
cat > secrets/provider_keys.json << 'EOF'
{
"provider_keys": {
"gemini": "AIza..."
}
}
EOF
cat > secrets/client.json << EOF
{
"client_token": "nk-$(openssl rand -hex 32)"
}
EOF
Note: the last heredoc is unquoted on purpose, so $(openssl rand -hex 32) expands once while the file is written and your token is unique.
Run the container (the same flags work with docker run):
podman run -d \
--name nenya \
-p 8080:8080 \
-v ./config:/etc/nenya:ro \
-v ./secrets:/run/secrets/nenya:ro \
-e NENYA_SECRETS_DIR=/run/secrets/nenya \
--cap-drop=ALL \
--cap-add=IPC_LOCK \
--security-opt=no-new-privileges:true \
--read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64M \
ghcr.io/gumieri/nenya:latest
Test it — the authenticated smoke test must list your configured model:
export NK=$(jq -r '.client_token' secrets/client.json)
curl -s -H "Authorization: Bearer $NK" http://localhost:8080/v1/models | jq -r '.data[].id'
Then send your first completion (streams an SSE response):
curl -N -H "Authorization: Bearer $NK" \
-d '{"model":"gemini-2.5-flash","messages":[{"role":"user","content":"Say hi in five words"}]}' \
http://localhost:8080/v1/chat/completions
No API key yet? The offline redaction demo runs Nenya against a local mock upstream — no external calls, no keys. It generates dummy secrets, sends fake AWS/GitHub credentials through the gateway, and shows what the upstream actually receives. Regenerate the GIF above with mise run demo.
Nenya provides native packages for major Linux distributions and community package managers:
| Distribution | Command |
|---|---|
| Debian/Ubuntu (.deb) | Download nenya_<version>_linux_amd64.deb from the release page and run sudo dpkg -i |
| Fedora/RHEL (.rpm) | Download nenya-<version>.x86_64.rpm from the release page and run sudo rpm -i |
| Arch Linux (.pkg.tar.zst) | Download nenya-<version>-x86_64.pkg.tar.zst from the release page and run sudo pacman -U |
| Arch Linux (AUR) | yay -S nenya-bin (or your preferred AUR helper) |
| Nix/NixOS | Add gumieri/nur-packages to your NUR registry and use nenya |
All packages install the binary to /usr/bin/nenya and include systemd service and socket units. After install, enable and start:
sudo systemctl enable --now nenya.socket
sudo systemctl enable --now nenya.service
flowchart TD
CLIENT["Client<br/>Cursor / OpenCode / Aider / etc.<br/>POST /v1/chat/completions · /v1/messages<br/>Bearer token"]
subgraph GW["Nenya Gateway"]
direction TB
AUTH["Auth + RBAC"]
RESOLVE["Parse body · resolve agent + targets<br/>strategy: fallback · round-robin · sticky"]
CACHE{"Response cache"}
MCPINJ["MCP auto-search + tool injection"]
end
CHAIN["Interceptor chain (best-effort)<br/>redact → entropy → TF-IDF → bouncer"]
TRIM["Token budget trim (hard limit)"]
subgraph LOOP["Dispatch loop — per target"]
direction TB
GUARDS["Circuit breaker · rate limits · cost guard"]
MODES["A standard forward<br/>B MCP multi-turn tool loop<br/>C context-limit retry"]
end
UPSTREAM["Upstream LLM providers<br/>23 built-in adapters"]
subgraph SSE["SSE pipeline"]
direction TB
PROBE["Stream-head probe (pre-header)<br/>empty / early-error failover"]
XFORM["Adapter transforms · format conversion"]
WATCH["Stall watchdog · stream continuation"]
ACCT["Usage accounting · cache capture · MCP auto-save"]
end
OUT["Client receives transparent SSE"]
CLIENT --> AUTH --> RESOLVE --> CACHE
CACHE -- "HIT → replay" --> OUT
CACHE -- "miss" --> MCPINJ --> CHAIN --> TRIM --> GUARDS --> MODES --> UPSTREAM
UPSTREAM --> PROBE --> XFORM --> WATCH --> ACCT --> OUT
PROBE -. "failover → next target" .-> GUARDS
classDef io fill:#e8ebf0,stroke:#57606a,color:#1f2328
classDef gw fill:#ddf4ff,stroke:#0969da,color:#1f2328
classDef pipe fill:#fff8c5,stroke:#9a6700,color:#1f2328
classDef sse fill:#dafbe1,stroke:#1a7f37,color:#1f2328
class CLIENT,OUT io
class GW,LOOP gw
class CHAIN,TRIM pipe
class SSE sse
Flow notes:
/v1/* endpoints require client bearer auth; /healthz, /statsz, /metrics do not.format attributeLimitMEMLOCK=infinity and LimitCORE=0 in systemd/tmpsystemctl reload nenya for zero-downtime config changeserror_kind field for programmatic diagnosticsAll /v1/* endpoints require Authorization: Bearer <client_token> or Bearer <api_key_token>.
API keys support RBAC enforcement — agent scoping, endpoint allowlists, role-based permissions (admin bypasses all checks).
| Endpoint | Auth | Description |
|---|---|---|
POST /v1/chat/completions | Bearer + RBAC | OpenAI-compatible chat with SSE streaming, agent fallback, MCP multi-turn |
POST /v1/messages | Bearer + RBAC | Anthropic Messages API with bidirectional format conversion |
GET /v1/models | Bearer + RBAC | Live model catalog from discovered providers + static registry (context window, max tokens) |
POST /v1/embeddings | Bearer + RBAC | Passthrough proxy |
POST /v1/responses | Bearer + RBAC | Passthrough proxy |
POST /v1/images/generations | Bearer + RBAC | Image generation (OpenAI-compatible) |
POST /v1/audio/transcriptions | Bearer + RBAC | Audio transcription (Whisper-compatible, multipart support) |
POST /v1/audio/speech | Bearer + RBAC | Text-to-speech synthesis (OpenAI-compatible) |
POST /v1/moderations | Bearer + RBAC | Content moderation (OpenAI-compatible) |
POST /v1/rerank | Bearer + RBAC | Re-ranking API (Cohere/Jina/Voyage-compatible) |
POST /v1/a2a | Bearer + RBAC | Agent-to-Agent protocol (Google A2A) |
GET/POST/DELETE /v1/files | Bearer + RBAC | File listing, upload, retrieval, deletion |
POST/GET /v1/batches | Bearer + RBAC | Batch API operations |
POST /proxy/{provider}/* | Bearer + RBAC | Arbitrary provider endpoint passthrough (all HTTP methods, SSE streaming) |
GET /healthz | None | Engine health probe |
GET /statsz | None | Token usage, circuit breaker state, MCP server status |
GET /metrics | None | Prometheus-compatible metrics |
GET /debug/pprof/* | Bearer | Go profiling endpoints (disabled by default, see debug.pprof_enabled) |
See docs/PASSTHROUGH_PROXY.md for detailed passthrough proxy usage.
Nenya supports standard environment variables for deployment portability:
| Variable | Default | Description |
|---|---|---|
PORT | 8080 | Listening port (overrides server.listen_addr) |
HOST | — | Optional bind address (e.g. 127.0.0.1). Only used when combined with PORT |
NENYA_CONFIG_DIR | /etc/nenya/ | Configuration directory path |
NENYA_CONFIG_FILE | — | Single config file path (takes precedence over NENYA_CONFIG_DIR) |
NENYA_SECRETS_DIR | — | Secrets directory (overrides CREDENTIALS_DIRECTORY) |
Example usage:
PORT=9090 HOST=127.0.0.1 ./nenya --config /path/to/config.json
Or in Docker:
docker run -e PORT=9090 -p 9090:9090 ghcr.io/gumieri/nenya:latest
| Document | Description |
|---|---|
| Providers | All 23 providers, capabilities matrix, special behaviors, adding custom providers |
| Configuration | Full config reference, directory mode, all sections and fields |
| Deploy Bare Metal | Systemd unit, config.d layout, secrets, hot reload |
| Deploy Container | Podman/Docker Compose, image verification, security notes |
| Deploy Kubernetes | Helm chart usage, ConfigMap/Secret, ingress setup |
| Passthrough Proxy | Raw provider endpoint proxying, SSE streaming, auth injection |
| Architecture | Package DAG, request lifecycle, circuit breaker, SSE pipeline |
| MCP Integration | MCP server integration, tool discovery, multi-turn execution |
| Adapters | Adapter system internals, auth styles, capability flags |
| Secrets Format | Systemd credentials, env var fallback, container/K8s deployment |
| Security | Vulnerability reporting policy |
| Disclaimer | Best-effort redaction scope and limitations |
| Changelog | Release history and notable changes |
Apache 2.0. See LICENSE.
559 commits
12 commits
Go
99.6%