askalf/dario

Use your Claude and ChatGPT subscriptions in Cursor, Cline, Aider, Claude Code and the Agent SDK — at subscription pricing, not per-token API bills. One local Anthropic + OpenAI-compatible endpoint: either plan answers either wire shape, auto-failover when one hits its limit, multi-seat pooling, live Claude Code drift tracking.

523

stars

1,059

commits

JavaScript

primary language

Sep 11, 2026

updated

www.npmjs.com/package/@askalf/dario
agent-sdk
aider
ai-gateway
anthropic
api-proxy
chatgpt
claude
claude-code
claude-code-proxy
claude-max
cline
codex
cursor
litellm
llm-proxy
llm-router
ollama
openai
openai-compat
openrouter

README

dario routes every AI tool you use to the subscriptions you already pay for. Claude Code, Cursor, Cline, Aider, Codex CLI and the Agent SDK send requests to dario at localhost:3456, which forwards each one to your Claude plan or your ChatGPT plan and fails over between them on a 429.

dario

Your Claude and ChatGPT subscriptions each work in exactly one place.
dario makes them work everywhere — at subscription pricing, not per-token API bills.

npm version Latest release CI CodeQL OpenSSF Scorecard OpenSSF Best Practices License Downloads Node version

Claude Code version the bundled template tracks (read live from master) Daily live billing canary Hourly Claude Code drift watch Live template drift watch

One local endpoint. Every AI tool you own. The subscriptions you already pay for.

npm i -g @askalf/dario · 0 runtime deps · SLSA-attested every release · nothing phones home · ~33k lines you can read in a weekend · independent, unofficial, third-party (DISCLAIMER.md)

Start · Your tools · Routing · Two plans · Pool · Drift · Trust · Risk · Commands · FAQ · Coming back after a while?


You're already paying $20, $100 or $200 a month for Claude,1 or for a ChatGPT plan. Then Cursor wants an API key. Aider wants an API key. Cline, Continue, Zed, your own scripts — every one of them bills you again, per token, while the plan you bought sits idle in the one app it shipped with.

dario is one local endpoint that routes all of them through the plans you already pay for. Point any Anthropic- or OpenAI-compatible tool at http://localhost:3456 and you're done. No per-tool config, no second bill, and when one plan hits its limit the other one takes the request.

Start in 60 seconds

# 1. Install
npm install -g @askalf/dario

# 2. Log in to your Claude subscription (Pro, Max 5x, or Max 20x)
dario login                 # or `dario login --manual` for SSH / headless

# 3. Start the local proxy
dario proxy                 # separate terminal or background

# 4. Point any Anthropic-compatible tool at it
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
Terminal: npm install -g @askalf/dario, dario login (Opening browser to sign in… Login successful!), dario proxy (dario — http://localhost:3456. Your Claude subscription is now an API. Usage: ANTHROPIC_BASE_URL=http://localhost:3456, ANTHROPIC_API_KEY=dario. OAuth healthy, Model passthrough, Pool: 1 account), then export the two variables and run aider --model sonnet.

That's the whole setup. Every tool that honors those env vars now runs on your subscription. OpenAI-shaped tools use OPENAI_BASE_URL=http://localhost:3456/v1 instead, same key.

Works with: Claude Code, Cursor, Aider, Cline, Roo Code, Kilo Code, Continue.dev, Zed, OpenHands, OpenClaw, Hermes, Codex CLI, the Claude Agent SDK, the Anthropic and OpenAI SDKs, curl, your own scripts. Per-tool snippets are one section down.

Prefer Docker? ghcr.io/askalf/dario:latest — multi-arch (amd64 + arm64), published from the same workflow as every npm release (guide). Something off? dario doctor prints one paste-ready health report.

Point your tools at it

Two base URLs, one key. Anthropic-shaped clients talk to http://localhost:3456; OpenAI-shaped clients talk to http://localhost:3456/v1. The key is dario (any value works until you set DARIO_API_KEY, which then has to match).

Claude Code — forwarded verbatim
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
claude

A genuine Claude Code request already is the Claude Code shape, so dario forwards it byte-for-byte — system prompt, tools, thinking, key order untouched — swapping in only the pool's credential, its billing tag and cache breakpoints. That covers the main loop, its Task/Agent sub-agents and the permission classifier. What you gain is the pool: several seats behind the one URL, with headroom routing and 429 failover. Background: #678.

Cursor — needs a public HTTPS tunnel, and the anthropic: prefix

Cursor's BYOK is backend-mediated: the app sends your base URL up to Cursor's servers and they make the call, behind an SSRF guard that rejects localhost by design (confirmed by Cursor staff, threads linked in the long-form guide). So:

dario proxy                                     # terminal 1
cloudflared tunnel --url http://localhost:3456   # terminal 2 → https://<random>.trycloudflare.com

In Cursor → Settings → Models: enable Override OpenAI Base URL with https://<random>.trycloudflare.com/v1, key dario, and add models as anthropic:opus / anthropic:sonnet / anthropic:haiku. The anthropic: prefix routes to the Claude backend without the claude- substring that makes Cursor switch to a tool format the OpenAI path can't parse, and it dodges Cursor's built-in-name collision. Use Agent mode (Cmd/Ctrl+I); Chat sends no tools. Treat the tunnel URL as a credential. Full walkthrough with every gotcha: agent-compat.md#cursor.

Cline · Roo Code · Kilo Code — API provider "Anthropic"

Provider Anthropic · API key dario · Anthropic Base URL http://localhost:3456 · model claude-sonnet-5 / claude-opus-5 / claude-haiku-4-5.

These clients speak an XML tool protocol. dario detects them from their system-prompt identity markers and flips into preserve-tools mode on its own, so their schemas pass through and their parsers keep working. --no-auto-detect if you'd rather choose. Details.

Aider
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
aider --model sonnet        # or opus, haiku, any claude-* id
Continue.dev
# ~/.continue/config.yaml
models:
  - name: Claude Sonnet (dario)
    provider: anthropic
    model: claude-sonnet-5
    apiBase: http://localhost:3456
    apiKey: dario
Zed
{ "language_models": { "anthropic": { "api_url": "http://localhost:3456", "version": "2023-06-01" } } }

Set ANTHROPIC_API_KEY=dario in the environment Zed launches from; the model picker then lists Claude models routed through your plan.

OpenHands
export LLM_BASE_URL=http://localhost:3456
export LLM_API_KEY=dario
export LLM_MODEL=anthropic/claude-sonnet-5

The anthropic/ prefix tells LiteLLM (OpenHands' router) to take the Anthropic path, which dario is now fronting. End-to-end walkthrough: openhands-walkthrough.md.

OpenClaw
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
openclaw "task description"

OpenClaw's exec / process / web_search / web_fetch / browser tools are translated to Claude Code's set without a flag; a tool outside the map (message, for one) rides a fallback slot instead. Newer OpenClaw reads auth-profiles.json before env vars, so a stale key there wins — the walkthrough covers it.

Codex CLI · OpenAI SDK · any OpenAI-compatible tool
export OPENAI_BASE_URL=http://localhost:3456/v1
export OPENAI_API_KEY=dario

Ask for gpt-5.5 and it is served by your ChatGPT plan once you've run dario add altman. Ask for claude-sonnet-5 on the same URL and it is served by your Claude plan, translated both ways. Ask for gpt-4o or anything else your API-key backend lists and it goes there byte-for-byte. Names that don't look like OpenAI's (llama-3.3-70b, qwen-coder) need the provider prefix below; dario refuses them rather than guess:

dario backend add openai     --key=sk-proj-...
dario backend add groq       --key=gsk_...    --base-url=https://api.groq.com/openai/v1
dario backend add openrouter --key=sk-or-...  --base-url=https://openrouter.ai/api/v1
dario backend add local      --key=anything   --base-url=http://127.0.0.1:11434/v1

Force a backend with a prefix: openai:gpt-4o, claude:opus, groq:llama-3.3-70b, local:qwen-coder.

One holdover for old configs: six legacy OpenAI names (gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.3, gpt-4, gpt-3.5-turbo) sent to /v1/chat/completions are translated to Claude models when no other provider claims them.

Claude Agent SDK · Anthropic SDK (TypeScript, Python)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "http://localhost:3456", apiKey: "dario" });
import anthropic
client = anthropic.Anthropic(base_url="http://localhost:3456", api_key="dario")

Zero code change beyond the base URL. Streaming, tool use, prompt caching and extended thinking all pass through. More in usage.md.

curl
curl http://localhost:3456/v1/messages -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","max_tokens":256,"messages":[{"role":"user","content":"Hello!"}]}'

curl http://localhost:3456/v1/chat/completions -H "content-type: application/json" \
  -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Hello!"}]}'
Docker · Kubernetes · a Pi in a closet
docker volume create dario-config
docker run --rm -it -v dario-config:/home/dario/.dario ghcr.io/askalf/dario:latest login --manual
docker run -d --name dario -p 3456:3456 -v dario-config:/home/dario/.dario \
  -e DARIO_API_KEY="$(openssl rand -hex 32)" ghcr.io/askalf/dario:latest

The image binds 0.0.0.0, so a key is mandatory; without one dario refuses to start rather than become an open relay for your subscription. No console at all? Start empty with DARIO_ADMIN=1 and provision the first account over HTTP with the admin API. Two replicas sharing accounts need the refresh lock. Docker guide.

Something not listed? If it reads ANTHROPIC_BASE_URL or OPENAI_BASE_URL, or has a "Base URL" field, it works. The compatibility matrix says which tools are exercised end-to-end, which are inferred from a shared code path, and which are untested — one honest cell per tool.

What it does with a request

You point every tool at one URL. dario reads each request, decides which plan or backend owns it, and forwards it in that backend's native protocol.

Client speaksModelRoutes toWhat happens
Anthropic Messagesclaude-* / opus / sonnet / haikuClaude poolOAuth swap + Claude Code template, then api.anthropic.com
Anthropic Messagesa slug your ChatGPT account listsCodex engineMessages→Responses translation, subscription auth
Anthropic Messagesgpt-4o, llama-*, any name no plan listsRefused400 + x-dario-upstream-rejection: model_unroutable; reach an API-key backend from this shape with a provider prefix
OpenAI Chatgpt-* / o1-* / o3-* / o4-*OpenAI-compat backendAuth swap, body forwarded byte-for-byte
OpenAI Chata slug your ChatGPT account listsCodex enginechat/completions→Responses translation, subscription auth
OpenAI Chatclaude-*Claude poolOpenAI→Anthropic translation, then the Claude path
Either<provider>:<model>Forced by prefixExplicit override

The tool doesn't know. The backend doesn't know. dario is the seam.

The full Claude lineup, autodetected. Fable 5, Opus 5, Sonnet 5 and Haiku 4.5, plus [1m] long-context variants on every family except Haiku, by full id (claude-opus-5) or shortcut (fable / opus / sonnet / haiku; append 1m for the long-context form; opus48 / opus47 / opus46 / sonnet46 pin a generation and never float). GET /v1/models reads Anthropic's live catalog (TTL-cached, baked fallback offline), so a new model resolves the day it lands with no dario release, and the model-specific request shape is applied automatically. Families pulled upstream are filtered from both the live catalog and the fallback, so /v1/models never advertises a model that 404s. A name no provider lists at all, such as a ChatGPT slug your account doesn't have or a typo that belongs to no family, is refused locally with 400 and x-dario-upstream-rejection: model_unroutable instead of spending a pool request on an upstream 404. The guard steps aside for claude-* names (Anthropic's own 404 stays authoritative there), for requests under a --model / --fast-model override, for upstream-API-key mode, and for the legacy OpenAI names the built-in map translates.

Two plans, one endpoint

Your ChatGPT plan, on both endpoints

A ChatGPT Plus or Pro plan is served on both of dario's endpoints: any client that speaks /v1/chat/completions can use it (Codex CLI, the OpenAI SDKs, your scripts), and so can any client that speaks /v1/messages (Claude Code, the Anthropic SDKs, agent runtimes). The harness never needs to know which subscription is behind it.

dario add altman            # prints an authorize URL; paste the redirect URL back
dario codex list
dario codex remove altman

dario add altman names whose plan you are attaching; dario add amodei attaches a Claude account instead. The browser lands on a localhost page that doesn't load — expected, nothing is listening there. Copy the whole address bar and paste it at the prompt; dario reads the code out of it.

curl localhost:3456/v1/models | jq -r '.data[].id'
curl localhost:3456/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"hi"}]}'

# same subscription, Anthropic wire shape — this is what Claude Code speaks
curl localhost:3456/v1/messages -H 'content-type: application/json' \
  -d '{"model":"gpt-5.5","max_tokens":64,"messages":[{"role":"user","content":"hi"}]}'

Model names are discovered, not hardcoded. The set a ChatGPT subscription may use is per-account and moves; dario asks the backend which models this account lists, caches the answer, and advertises them on GET /v1/models. Anything not on that list (gpt-4o and friends) still routes to a configured API-key backend as before. codex:<model> / chatgpt:<model> forces the route.

Streaming, tool calls and tool-result round trips work on both shapes, and chat-shape image_url parts are carried as Responses input_image parts with detail preserved: dario translates chat/completions or Messages into the Responses API the subscription backend speaks, and translates the stream back into chat.completion.chunk or Anthropic message events. There is no /v1/responses inbound yet. The Codex backend does not accept every chat field, so response_format, stop, n, logprobs, stream_options and the sampling parameters temperature, top_p, max_tokens, max_completion_tokens are intentionally lossy; with --verbose, dario reports each field that does not reach Codex once per process. Codex accounts live in ~/.dario/codex-accounts/, separate from the Claude pool.

Prompt caching: the backend caches prompt prefixes of 1,024 tokens and up on its own; what dario adds is the prompt_cache_key that routes same-prefix requests to the cache that holds them, the way the Codex CLI does with its session id. A chat/completions client that sets its own key keeps it; an Anthropic-shape request gets one per Claude Code session (a hash of metadata.user_id, never the raw ids); anything else is keyed on its model, instructions and tool names, so repeated system prompts from any caller land together. Cached tokens come back as prompt_tokens_details.cached_tokens on chat/completions and as cache_read_input_tokens on /v1/messages, and show up in /analytics and the -v usage line like a Claude request's do.

Failover between subscriptions

Two consumer plans, no API keys, and neither one able to take you down on its own.

dario proxy --pool-fallback=gpt-5.6-sol,claude-sonnet-5

That is a chain, read left to right; each provider takes the first entry it can actually serve. Prefix every entry with a tier and the same flag becomes a tier map: --pool-fallback=haiku:gpt-5.6-luna,sonnet:gpt-5.6-terra,opus:gpt-5.6-sol picks one rung per request from the tier of the model asked for (default: catches the rest, otherwise the first rung), so heartbeat work on Haiku never overflows onto a flagship. That shipped in 6.0.17, the same day a Haiku-tier fleet spent 79% of a weekly allowance doing exactly that. When the Claude pool is drained or cooling, the request is served as gpt-5.6-sol from your ChatGPT subscription. When the subscription is rate-limited or down, the request is handed back to the Claude pool as claude-sonnet-5. Every substituted response carries x-dario-pool-fallback: <model> — a silently swapped model family is exactly the surprise this project exists to avoid.

A tool sends a request to dario. The Claude plan answers 429, so dario re-serves the same request from the ChatGPT plan, which answers 200, and the response returns to the tool carrying the x-dario-pool-fallback header.

A single-entry chain is one-way and means what it always meant, so an existing config is unaffected. Failover is opt-in: without --pool-fallback, a drained pool still returns its honest 429/503. Only a 429 or 5xx fails over; a 400 surfaces, because a bad request that fails over just reproduces itself on the other provider and buries the real cause. A 429 also cools that provider for a bounded interval, its retry-after if it sent one and 60 s otherwise, never longer than 15 min, and an entry that already declined is not asked again within the same request. When every entry is cooling, the request ends on one honest 429 with a retry-after instead of a retry storm. The Claude entry has to be a model the pool can actually serve, checked positively against the live catalog, so a typo can't trade a recoverable 429 for an unrecoverable 404.

dario doctor tells you which of these you are actually in:

[ OK ]  Failover   symmetric: gpt-5.6-sol → claude-sonnet-5, across 1 Codex account
[WARN]  Failover   armed (gpt-5.6-sol) but INERT — no Codex account and no backend
                   to fall back to. Add one: `dario add altman`

That warning is the whole reason the check exists. Armed with nothing to fall back to is green on every other check and incapable of doing anything.

[!NOTE] Upgrading from v5? Nothing to do. Every v6 feature is opt-in and a single-value --pool-fallback behaves exactly as it did. CHANGELOG

Shadow compare

Once either subscription can serve either wire shape, the interesting question stops being can I reach GPT and becomes which of these is better at my work. Benchmarks answer that badly. Your own traffic answers it well.

curl localhost:3456/v1/messages \
  -H 'content-type: application/json' \
  -H 'x-dario-compare: gpt-5.6-sol' \
  -d '{"model":"claude-opus-5","max_tokens":1024,"messages":[…]}'

You get the Claude answer, exactly as you would have. Beside it, dario runs the same prompt past gpt-5.6-sol and writes both to ~/.dario/compare/<timestamp>-<model>.json, in your own wire shape, so you are comparing like with like. The comparison cannot degrade the request it observes: it only reads bytes already on their way out, your request is never held open for it, and a comparison that fails, times out or has nowhere to go is dropped with the record still written. Both sides are stored as raw payloads, because extracting text is where a bug would quietly make two answers look more alike than they are.

Many seats, one endpoint

Every dario is a pool. A plain dario login is a pool of one; there is no separate mode to switch on. Hold more than one seat — a personal Max and a work Max, a couple of Pros, team seats — and the same localhost:3456 routes every request to whichever seat has the most headroom, live, per request.

dario accounts add work
dario accounts add personal
dario proxy
The dario TUI Accounts tab: a table of pooled seats (work, personal, side) with token expiry, 5-hour and 7-day utilization, and status.

Three things it does that a round-robin doesn't:

  • Per-model headroom routing. Anthropic meters each model family separately: a 5h bucket, a 7d bucket and a per-model 7d_<family> bucket. dario reads all of them off every response and routes each request by the bucket that governs it — an Opus call to the seat with Opus room, a Sonnet call to the seat with Sonnet room, independently. Plan tiers mix freely; dario cares about headroom, not tier.
  • Session stickiness. Claude's prompt cache is scoped to {account × cache key}, so rotating a long conversation across seats on headroom alone re-pays cache-create every turn, a 5–10× token-cost multiplier on the cached portion. dario pins each conversation to one seat (hashed from its first message, deterministic) for the life of the session and rebinds only when that seat is exhausted.
  • In-flight 429 failover. A seat hits its wall mid-request and dario retries the same request against the next-best seat before your client ever sees an error. The sticky binding follows, so the next turn doesn't re-select the cold one. A seat parked on a 429 rejoins the pool on its own when its window resets.
Three pooled seats, work, personal and side, each with a headroom bar. dario routes the request to the seat with the most headroom.

--pool-strategy=fill-first concentrates new conversations on one seat until it drains, for primary/backup setups. Refresh tokens expire about 28 days after the original grant regardless of rotation, so every seat's grant age is tracked and surfaced in dario accounts list, dario doctor and GET /accounts before it becomes a silent outage. Provision over HTTP with the headless admin API; pin one request to one seat with dario accounts check <alias> (admin API required: DARIO_ADMIN=1 and a DARIO_ADMIN_TOKEN). Internals and the live /accounts + /analytics endpoints: multi-account-pool.md; covered end-to-end by test/pool-e2e.mjs.

Watch it happen

Type dario with no arguments for a full-screen control panel: live request stream, per-model burn rate, rate-limit utilization per seat, billing-bucket breakdown, and an in-place config editor that writes ~/.dario/config.json. Pure ANSI, zero new runtime deps. Tab moves between tabs, r refreshes, R resumes a halted overage guard, q quits.

The dario TUI Analytics tab: requests per minute, tokens in and out, thinking tokens, average latency, subscription percentage, a per-model bar chart, per-account rate-limit bars for the 5-hour and 7-day windows, and a billing breakdown.

Both screenshots are rendered from the real TUI against a fixture proxy by scripts/readme/tui.mjs, so a layout change shows up here instead of rotting a mock-up. The numbers are illustrative; the pixels are not.

It tracks a moving target

Claude Code's request shape changes between releases — new betas, tool renames, per-model thinking configs — usually with no subscriber-facing note. dario doesn't guess that shape: it captures it live from your own installed claude binary on every startup, diffs it against each upstream release, and replays it faithfully. That's why your subscription routes the same through dario as it does through Claude Code itself: the request that leaves your machine is the shape your plan expects. Details: wire-fidelity.md · #13 · #14.

The installed claude binary feeds its request shape into dario. A timeline of Claude Code releases ends in a node flagged as drift.

Keeping that current is the whole job, and it's automated. These watchers run unattended; each badge is the live status of that workflow's latest run, and its label is the cadence:

WatcherCatchesLive
cc-drift-watchA new Claude Code npm release that changes the wire shape. Auto-drafts the fix; cc-drift-auto-release merges and ships it within minutes.hourly
cc-drift-template-watchSame-binary remote-config drift, which no npm diff can see. Runs against a live Claude session on a self-hosted runner and opens a rebake PR with the diff inline.hourly
cc-billing-classifier-canaryClassifier drift: one real request a day must still bill to a subscription bucket.daily
wire-drift-self-hostedPer-model beta headers and billing blocks the installed claude actually sends, model by model.daily
sdk-drift-watchAgent SDK / Stainless pins drifting from what the template assumes.daily
pricing-drift-watchdario's pricing table drifting from Anthropic's published rates, so the TUI's cost figures stay honest.daily
codex-drift-watchThe ChatGPT backend's model list or wire contract moving under the translator.daily
cc-oauth-healthThe maintainer's own production proxy going unhealthy on any axis.every 30 min
dario-doctor-watchRuntime drift only a live dario doctor --obedience surfaces: identity, obedience, usage buckets.every 6 h
deployed-version-watchPublishing is not deploying: is what's running what was last released?hourly
cc-drift-watcher-livenessThe watcher itself going quiet. Lives on GitHub-hosted infrastructure on purpose, so it survives the failures it watches for.every 2 h

Guarded at PR time by live-test, a required check that runs the full suite against a live proxy on a self-hosted runner, plus compat-test-self-hosted, which replays the compat suite through a passthrough proxy on wire-shape changes. A few changes the watchers caught and shipped fixes for, same day:

Change (no subscriber-facing note)Effectdario shipped
context-1m dropped from the default beta set on the OAuth pathSubscription requests default to the 200K window on Sonnet/Opusv3.38.3–4
thinking: {type:"adaptive"} gated per-model server-sideSonnet/Opus 4-5 400 every request through any proxyv3.38.5
Per-model anthropic-beta setsProxies sending one set diverge for non-Opus modelsv4.8.53

The full ledger lives in the CHANGELOG, 500+ releases since April 2026. Setup and walkthrough: drift-monitor.md. The residual manual cases — OAuth rotation, runner re-registration — are in the recovery runbook.

Guardrails

Overage guard

During normal operation, a subscriber should never see a single response billed outside their subscription pool. If one is, something is wrong — wire-shape drift, an account misconfig, a change upstream — and forwarding more requests in the same shape either bleeds real money (accounts with extra usage enabled) or returns a wall of rejections. The first hit is the signal; the rest are damage.

Requests from five tools are stopped short of dario, which is ringed in red and labeled halted, after the Claude plan returned a response billed as overage.

So the moment any upstream response bills to something other than your subscription pool, dario halts the proxy. The check is an allow-list, not a match on one string: anything that isn't a known subscription claim (five_hour / seven_day, their _fallback and _overage_included variants, and the chatgpt_subscription claim dario stamps on Codex-served responses) and isn't the unknown no-header sentinel trips it, so a billing bucket dario has never seen still halts. Subsequent requests return 503 with an Anthropic-shaped error body until you run dario resume, press R in the TUI, or the cooldown clears (default 30 min). The halt shows across the TUI, fires a best-effort OS notification, and emits named SSE events. Tune it via ~/.dario/config.jsonoverageGuard, or --overage-behavior=warn / --no-overage-guard / --overage-cooldown=<ms>. In upstream-API-key passthrough mode (ANTHROPIC_UPSTREAM_API_KEY) the guard is off; api billing is the point there. Verified end-to-end by test/overage-guard-e2e-live.mjs. Background: #288.

The billing split, a contingency dario is built for

On 2026-05-13 Anthropic announced that, from 2026-06-15, Agent SDK and claude -p (headless) traffic would leave the subscription pool for a small separate monthly credit, then metered API rates. They paused it before that date. Those surfaces still bill subscription today, and Anthropic says it will give advance notice before any revised version. Nothing changed; no credits were issued.

The split isn't live, but it was announced once on short notice and could return, so dario is built for it either way. Every request is rebuilt into interactive Claude Code shape before it leaves your machine (and, with --stealth, the response-correlated timing an interactive session has), so your traffic sits in the subscription pool whether a split is paused or live. The daily canary above is the tripwire: it surfaces a revived split within a day instead of on a surprise invoice. Verify on your own machine right now: dario doctor --usage fires one request and prints the rate-limit headers; representative-claim should read five_hour or seven_day, both subscription buckets. Full timeline: why-now-2026-06.md.

Trust & transparency

Everything inside the box labeled your machine: the tools and dario. Only two lines leave it, to the Claude plan and the ChatGPT plan, through a padlock. Pills read 0 deps, no telemetry, MIT.
SignalStatus
Source~33k lines of TypeScript across 68 files, auditable in a weekend. One credential path since v5: the pool.
Dependencies0 runtime. Verify: npm ls --production
ProvenanceEvery release SLSA-attested via GitHub Actions + Sigstore, published with OIDC trusted publishing — no long-lived npm token exists to leak
ScanningCodeQL on every push and weekly · ClusterFuzzLite fuzzes the SSE translator and rejection parsers weekly · OpenSSF Scorecard and Best Practices badges above are live
Tests178 test files run in parallel by npm test on Node 18, 20 and 22; the live e2e / compat / stealth suites have their own entry points. Green on every release
CredentialsYour own subscription tokens, never logged, redacted from errors, 0600 on disk in 0700 dirs
NetworkBinds 127.0.0.1 by default; upstream only to configured backends over HTTPS; hardcoded SSRF allow-list; refuses a non-loopback bind without DARIO_API_KEY
TelemetryNone. No analytics, no tracking, nothing phones home
This READMECI fails if the line count above drifts from src/ or a link or anchor here stops resolving (check-readme-line-count.mjs, check-readme-links.mjs); the TUI screenshots are rendered from the real TUI and the diagrams are briefed art, not screenshots (how)
npm audit signatures
npm view @askalf/dario dist.integrity
cd $(npm root -g)/@askalf/dario && npm ls --production

Security reports go to security@askalf.org, not a public issue: SECURITY.md. API stability commitments (@stable / @experimental / @deprecated, deprecation cycles): STABILITY.md.

Will my account get suspended?

The most common question about dario, and it deserves a straight answer: I can't promise you won't be actioned, and I'd be skeptical of anyone who does. Only Anthropic decides how it enforces its terms. What I can do is lay out exactly how dario works, so you can weigh the risk yourself instead of taking anyone's word for it.

What dario does:

  • Runs entirely on your machine. Your subscription token never touches my servers or anyone else's; requests go straight from your computer to Anthropic.
  • Authenticates as you, with your own Claude login, the same OAuth credential Claude Code itself uses. It impersonates nobody and shares nothing.
  • Doesn't modify your account, billing, or subscription settings.
  • Sends requests in the shape the official client sends them, rebuilt from your own installed binary, not spoofed from a hardcoded fake.
  • Reports nothing, anywhere. No telemetry, no analytics, nothing phones home; verifiable in the source, which is the point of keeping it auditable in a weekend.

What dario does that Claude Code doesn't: it lets tools other than Claude Code use that subscription. That's the whole point of it, and it's also the part that sits outside what Anthropic's own client does. Whether that falls within your plan's terms is Anthropic's call, not mine. Read their terms, read DISCLAIMER.md, and decide deliberately. dario is a transparency tool, in that it documents request behavior Anthropic doesn't publish for subscribers, and it is also, plainly, routing subscription traffic that Anthropic's own tools bill differently. Both are true; decide with both in view.

On policy risk specifically: Anthropic's position on third-party clients has moved before and can move again. dario is built to surface that fast rather than paper over it; see the billing split for the contingency already in place and the daily canary watching for it.

Ongoing discussion, including other users' experiences: #724.

Who it's for

Best fit: developers juggling multiple LLM tools and per-tool API keys · Claude Pro/Max subscribers who want their plan usable everywhere, not just in Claude Code · ChatGPT Plus/Pro subscribers who want their plan in OpenAI-compatible harnesses · teams running local or hosted OpenAI-compat servers who want one stable local endpoint · Agent SDK users who want subscription routing with zero code change · power users wanting multi-account pooling with 429 failover.

Not a fit: you need vendor-managed production SLAs (use the provider APIs) · you want a hosted multi-tenant team platform with dashboards and SSO (dario is a single-owner local proxy) · you want a chat UI (use claude.ai).

How it compares. Only one of these routes a consumer subscription; the others route API keys, and that is the whole split.

ToolWhat it isWhen it wins
darioLocal proxy that routes your Claude and ChatGPT plans, plus any OpenAI-compatible APIYou already pay for a plan and want every tool on your machine to use it
LiteLLMPython SDK + proxy, 100+ providers via API keys, enterprise featuresYou have API keys, want central spend controls, or run a hosted multi-tenant service
OpenRouterHosted aggregator, one API key for hundreds of modelsYou want model breadth and are fine with pay-per-token
Kong AI GatewayEnterprise on-prem API gateway for LLMsYou already run Kong and need AI traffic under the same governance

Longer version, with specifics: #68.

Commands

CommandWhat it does
darioThe TUI: status, config editor, analytics, hits, accounts, backends
dario login [--manual]Log in to your Claude plan. Picks up Claude Code's credentials or runs its own OAuth flow; --manual for SSH / containers
dario proxyStart the local endpoint on :3456
dario doctor [--usage] [--probe] [--obedience] [--auth-check] [--bun-bootstrap] [--json]One aggregated health report: runtime/TLS, template and drift, OAuth, pool, refresh-grant age, failover readiness, backends
dario add altman / dario add amodeiAttach a ChatGPT plan / a Claude account, by whose it is
dario accounts list / add / remove / check <alias>Pool management; check sends one pinned request per model through the running proxy (admin API on)
dario backend list / add / removeOpenAI-compatible API-key backends
dario codex list / add / removeChatGPT accounts (the long form of dario add altman)
dario usage · dario config · dario statusBurn rate for the last hour · effective config, redacted · token health
dario resume · dario refresh · dario logout · dario upgradeClear an overage halt · force a token refresh · delete credentials · safe self-update
dario mcp · dario subagent install / remove / statusReach dario from inside any MCP client, or from inside a Claude Code session, read-only
EndpointDescription
POST /v1/messages · POST /v1/chat/completionsThe two wire shapes, any plan behind either
GET /v1/modelsLive model list: the Claude catalog plus whatever your ChatGPT plan lists
GET /health · GET /livezServiceability (503 when not) · liveness. /health?probe=1 sends one real request
GET /status · GET /accounts · GET /analyticsOAuth detail · per-seat utilization and grant age · per-account / per-model stats and burn rate
POST /v1/messages/count_tokens · POST /v1/completeToken counting and the legacy Text Completions shape
GET /analytics/stream · GET /codexLive analytics over SSE · ChatGPT-seat status, read without spending or exposing a token
/admin/*Provisioning, GET /admin/accounts, POST /admin/resume; only with DARIO_ADMIN=1 (admin API)

Flags: commands.md, plus dario --help for the ones it doesn't list yet (--effort, --max-tokens, --model-alias, --fast-model, session rotation, concurrency caps, the pacing knobs behind --stealth) · env vars grouped by task, for Docker / k8s / systemd: configuration.md · SDK examples: usage.md.

More knobs — stealth timing, system-prompt modes, client-shape overrides, VPN egress, MCP
  • Behavioral stealth (--stealth). Adds when a request arrives to what it looks like: response-length-correlated think time and session-start latency. wire-fidelity.md
  • Recover output (--system-prompt=partial). Strips Claude Code's tone and verbosity constraints for 1.2–2.8× more output on open-ended work, without changing which pool you bill to. #183 · system-prompt.md
  • Client-shape overrides. --honor-client-thinking passes a client's own thinking block through; --preserve-output-format carries a client's output_config.format JSON schema through so structured-output SDKs get schema-constrained output. Both off by default.
  • Runs any agent. A 64-entry schema-verified TOOL_MAP pre-maps Cline, Roo, Kilo, Cursor, Windsurf, Continue, Copilot, OpenHands, OpenClaw and Hermes tool names to Claude Code's native set; MCP tools (mcp__server__tool) forward verbatim. Custom schemas: --preserve-tools or --hybrid-tools. agent-compat.md
  • Model aliases and caps. --model-alias=<name=target> (repeatable) advertises a name of your choosing on /v1/models and routes it; --effort=<low|medium|high|xhigh|ultracode|max|client> and --max-tokens=<N|client> set, or pass through, per-request effort and output caps.
  • VPN / egress routing. Route dario's upstream traffic through a VPN without putting the whole host on one. vpn-routing.md
  • More than one instance, same accounts. Refresh tokens are single-use, so two replicas refreshing the same seat leave one holding a dead token; the optional refresh lock (Redis or Cloudflare) makes the loser adopt the winner's credentials. multi-instance.md
  • PII redaction in front of dario. Pair it with cordon: integrations/cordon.md
  • Reachable from inside Claude Code or any MCP client. dario subagent install registers a sub-agent for in-session diagnostics; dario mcp exposes dario as a read-only MCP server. sub-agent.md · mcp-server.md

FAQ

Does this violate Anthropic's terms?

Mechanically, dario uses your existing Claude Code OAuth tokens: it authenticates you as you, with your subscription, through Anthropic's official endpoints. Whether any particular use complies with current terms is between you and Anthropic; consult their terms and your agreement. Independent, unofficial, third-party — see DISCLAIMER.md. On the suspension question specifically: Will my account get suspended?

Do I need Claude Code installed?

Recommended, not required. With it, dario login picks up credentials automatically and the template extractor reads your binary on every startup. Without it, dario runs its own OAuth flow and falls back to the bundled (scrubbed) template snapshot, which the drift watchers keep current.

Do I need Bun?

Optional, recommended: Bun's TLS ClientHello matches Claude Code's runtime, and dario relaunches itself under Bun when it finds one on PATH. Without it dario works fine on Node; dario doctor flags the mismatch and --strict-tls hard-fails until resolved.

Can I use dario without a Claude subscription?

Yes. Skip dario login, run dario add altman for a ChatGPT plan or dario backend add openai --key=… for an API key, and you have a local router with no Claude involvement. --no-claude-auth keeps the Claude token untouched entirely.

representative-claim: seven_day in my headers — am I downgraded?

No. five_hour and seven_day are both subscription billing, different accounting buckets in the same mode. overage is the one that flips you to per-token, and the overage guard halts on it. #1

My usage through dario is higher than through Claude Code directly. Why?

Almost always the prompt-cache TTL, not proxy overhead: dario mirrors whatever cache stamp your client sends, and many harnesses send the 5-minute one, so gaps longer than five minutes between turns re-create the prefix. DARIO_CACHE_TTL_1H=1 forces the 1-hour TTL. The full breakdown, with the two-message check that tells you which case you're in: faq.md.

Will the billing split break my setup?

Not today: it was announced, then paused before it took effect, and your traffic still bills subscription. What dario already does about it, and the canary that would catch a revival: the billing split.

Why "dario"?

It's a name, not an acronym. Don't overthink it.

Full FAQ, including per-tool 401s and Team/Enterprise plans: faq.md.

Deep dives

Contributing

PRs welcome. Small TypeScript codebase, zero runtime deps. Architecture, file-by-file map and the review bar in CONTRIBUTING.md; release mechanics in RELEASING.md.

git clone https://github.com/askalf/dario && cd dario
npm install
npm run dev    # tsx, no build step
npm test       # 178 files in parallel via test/all.test.mjs
npm run e2e    # live proxy + OAuth (needs a working Claude backend)

Drift and audit runners, none of them part of npm test:

npm run drift:wire     # compare a live Claude Code capture against the baked template
npm run drift:sdk      # Agent SDK / Stainless pin drift
npm run audit:tui      # drives the real TUI through a fake TTY at 12 geometries
npm run check:overage  # overage-classifier check against live headers
npm run stress         # concurrency / queue behaviour under load
npm run cch:calibrate  # re-derive the billing-tag cch seed for a new Claude Code build
npm run readme:assets  # regenerate the diagrams and TUI screenshots above

Two easy ways to help beyond code: star the repo, the clearest signal this is useful, and file drift: open an issue when a rate-limit header flips or a tool that worked yesterday breaks today, and it gets documented in public alongside the fix. Follow @ask_alf for drift bulletins as they land.

Star history of askalf/dario

Contributors

WhoContributions
@GodsBoyProxy auth, token redaction, error sanitization (#2)
@belangertradingBilling-classification investigation (#4, #6, #7, #12, #23), multi-agent billing FAQ (#27)
@wysieESM require crash in dario login (#15), OAuth for Max-plan accounts (#18)
@earlvanzeOpenClaw tool mappings (#19), OAuth manual override (#47), HTTPS warning (#53)
@nathan-widjajaREADME positioning structure — the promise → who → first use → why-switch spine the page still runs on (#21)
@trinhnvgemOAuth login failures on first release (#22), container and headless callback binding (#28)
@adubkovThe container / headless-SSH case behind the manual OAuth code paste (#28)
@iNicholasBEmacOS keychain credential detection (#30)
@boeingchocoReverse tool-param translation (#29), SSE framing regression catch, hybrid-tool motivation (#33, #36)
@tetsucoScrubber path corruption (#35), OpenClaw reverse-mapping collisions (#37), 20x-tier report (#42)
@mikelovattSilent subscription-drain surfaced via friendly billing buckets (#34)
@ringge--no-auto-detect for text-tool auto-preserve (#40)
@rustanacexdCursor BYOK routing for Claude, and --effort=max (#190)
@daimonbotOfficial multi-arch Docker image on GHCR (#199)
@Saik0sWildcard CORS allow-headers, Opus 4.7 catalog entry (#222)
@lwsh123kcch anchored to the billing tag instead of first match (#528)
@boredlandTime-to-reset in dario doctor --usage (#550)
@pnewell--preserve-output-format for structured-output SDKs (#583)
@matteo-ramaHeadless admin bootstrap (#599), Analytics NaN and per-account rate-limit rows (#600), pool-aware /status and /health (#636), version on both (#640), Accounts TUI reads the live pool (#641)
@miklisantonMid-session /model switch 400 (#744), empty-turn guards behind the subagent 400s (#1033, #1117)
@p-i-Independent wire-fidelity audit with a re-runnable harness — version-blind bun-match, and the correction to the packet-identical claim (#813)
@jerzydziewierzTUI Config tab clipping and scrolling (#861)
@ramarro123Admin bulk re-auth (#913), shared state across instances (#993), prompt-cache behaviour under litellm (#1018), parked-seat and shared-window reporting (#1244)
@zytegalaxyThe ChatGPT/Codex engine and dario add altman (#1009)
@chaogebabaAuto-release must never fire from a fork (#1029)
@robincleUtilisation freshness — lastObservedAt / utilAgeMs on /accounts (#1032)
@anupammeRefresh-lock ownership by server-issued lock id (#1059)
@LiveNathanNever send or stamp empty text blocks (#1067), empty final user turn from CC's stream-interruption retry (#1092, as @NathanLively)

Disclaimers

dario is an independent, unofficial, third-party project. Not affiliated with, endorsed by, or sponsored by Anthropic, OpenAI, or any vendor referenced here. Provided as-is, no warranty. You are solely responsible for compliance with your subscription's terms, the security of your credentials, and the content you send through the proxy. Not for safety-critical, regulated, or production environments without your own review. Full text: DISCLAIMER.md.

License

MIT — see LICENSE and DISCLAIMER.md. The embedded README font is Space Mono under the SIL Open Font License.

Own Your Stack

dario is the routing layer of Own Your Stack, open tools for owning your AI infrastructure instead of renting it by the token. One subscription. Your box. Your terms.

Built by Thomas Sprayberry

dario is part of Own Your Stack, the open toolkit behind Sprayberry Labs, the software studio with one human on staff, run by askalf, the AI operation these tools are part of.

Built in the open, scars included. Follow the build: @ask_alf · sprayberrylabs.com/own-your-stack

Footnotes

  1. Pro at $20 a month, Max 5x at $100, Max 20x at $200, as listed on claude.com/pricing on 2026-09-06. Annual billing is cheaper; check the page for what's current.

Contributors

askalf

983 commits

dependabot[bot]

51 commits

earlvanze

3 commits

askalf/dario

Use your Claude and ChatGPT subscriptions in Cursor, Cline, Aider, Claude Code and the Agent SDK — at subscription pricing, not per-token API bills. One local Anthropic + OpenAI-compatible endpoint: either plan answers either wire shape, auto-failover when one hits its limit, multi-seat pooling, live Claude Code drift tracking.

523

stars

1,059

commits

JavaScript

primary language

Sep 11, 2026

updated

www.npmjs.com/package/@askalf/dario
agent-sdk
aider
ai-gateway
anthropic
api-proxy
chatgpt
claude
claude-code
claude-code-proxy
claude-max
cline
codex
cursor
litellm
llm-proxy
llm-router
ollama
openai
openai-compat
openrouter

README

dario routes every AI tool you use to the subscriptions you already pay for. Claude Code, Cursor, Cline, Aider, Codex CLI and the Agent SDK send requests to dario at localhost:3456, which forwards each one to your Claude plan or your ChatGPT plan and fails over between them on a 429.

dario

Your Claude and ChatGPT subscriptions each work in exactly one place.
dario makes them work everywhere — at subscription pricing, not per-token API bills.

npm version Latest release CI CodeQL OpenSSF Scorecard OpenSSF Best Practices License Downloads Node version

Claude Code version the bundled template tracks (read live from master) Daily live billing canary Hourly Claude Code drift watch Live template drift watch

One local endpoint. Every AI tool you own. The subscriptions you already pay for.

npm i -g @askalf/dario · 0 runtime deps · SLSA-attested every release · nothing phones home · ~33k lines you can read in a weekend · independent, unofficial, third-party (DISCLAIMER.md)

Start · Your tools · Routing · Two plans · Pool · Drift · Trust · Risk · Commands · FAQ · Coming back after a while?


You're already paying $20, $100 or $200 a month for Claude,1 or for a ChatGPT plan. Then Cursor wants an API key. Aider wants an API key. Cline, Continue, Zed, your own scripts — every one of them bills you again, per token, while the plan you bought sits idle in the one app it shipped with.

dario is one local endpoint that routes all of them through the plans you already pay for. Point any Anthropic- or OpenAI-compatible tool at http://localhost:3456 and you're done. No per-tool config, no second bill, and when one plan hits its limit the other one takes the request.

Start in 60 seconds

# 1. Install
npm install -g @askalf/dario

# 2. Log in to your Claude subscription (Pro, Max 5x, or Max 20x)
dario login                 # or `dario login --manual` for SSH / headless

# 3. Start the local proxy
dario proxy                 # separate terminal or background

# 4. Point any Anthropic-compatible tool at it
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
Terminal: npm install -g @askalf/dario, dario login (Opening browser to sign in… Login successful!), dario proxy (dario — http://localhost:3456. Your Claude subscription is now an API. Usage: ANTHROPIC_BASE_URL=http://localhost:3456, ANTHROPIC_API_KEY=dario. OAuth healthy, Model passthrough, Pool: 1 account), then export the two variables and run aider --model sonnet.

That's the whole setup. Every tool that honors those env vars now runs on your subscription. OpenAI-shaped tools use OPENAI_BASE_URL=http://localhost:3456/v1 instead, same key.

Works with: Claude Code, Cursor, Aider, Cline, Roo Code, Kilo Code, Continue.dev, Zed, OpenHands, OpenClaw, Hermes, Codex CLI, the Claude Agent SDK, the Anthropic and OpenAI SDKs, curl, your own scripts. Per-tool snippets are one section down.

Prefer Docker? ghcr.io/askalf/dario:latest — multi-arch (amd64 + arm64), published from the same workflow as every npm release (guide). Something off? dario doctor prints one paste-ready health report.

Point your tools at it

Two base URLs, one key. Anthropic-shaped clients talk to http://localhost:3456; OpenAI-shaped clients talk to http://localhost:3456/v1. The key is dario (any value works until you set DARIO_API_KEY, which then has to match).

Claude Code — forwarded verbatim
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
claude

A genuine Claude Code request already is the Claude Code shape, so dario forwards it byte-for-byte — system prompt, tools, thinking, key order untouched — swapping in only the pool's credential, its billing tag and cache breakpoints. That covers the main loop, its Task/Agent sub-agents and the permission classifier. What you gain is the pool: several seats behind the one URL, with headroom routing and 429 failover. Background: #678.

Cursor — needs a public HTTPS tunnel, and the anthropic: prefix

Cursor's BYOK is backend-mediated: the app sends your base URL up to Cursor's servers and they make the call, behind an SSRF guard that rejects localhost by design (confirmed by Cursor staff, threads linked in the long-form guide). So:

dario proxy                                     # terminal 1
cloudflared tunnel --url http://localhost:3456   # terminal 2 → https://<random>.trycloudflare.com

In Cursor → Settings → Models: enable Override OpenAI Base URL with https://<random>.trycloudflare.com/v1, key dario, and add models as anthropic:opus / anthropic:sonnet / anthropic:haiku. The anthropic: prefix routes to the Claude backend without the claude- substring that makes Cursor switch to a tool format the OpenAI path can't parse, and it dodges Cursor's built-in-name collision. Use Agent mode (Cmd/Ctrl+I); Chat sends no tools. Treat the tunnel URL as a credential. Full walkthrough with every gotcha: agent-compat.md#cursor.

Cline · Roo Code · Kilo Code — API provider "Anthropic"

Provider Anthropic · API key dario · Anthropic Base URL http://localhost:3456 · model claude-sonnet-5 / claude-opus-5 / claude-haiku-4-5.

These clients speak an XML tool protocol. dario detects them from their system-prompt identity markers and flips into preserve-tools mode on its own, so their schemas pass through and their parsers keep working. --no-auto-detect if you'd rather choose. Details.

Aider
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
aider --model sonnet        # or opus, haiku, any claude-* id
Continue.dev
# ~/.continue/config.yaml
models:
  - name: Claude Sonnet (dario)
    provider: anthropic
    model: claude-sonnet-5
    apiBase: http://localhost:3456
    apiKey: dario
Zed
{ "language_models": { "anthropic": { "api_url": "http://localhost:3456", "version": "2023-06-01" } } }

Set ANTHROPIC_API_KEY=dario in the environment Zed launches from; the model picker then lists Claude models routed through your plan.

OpenHands
export LLM_BASE_URL=http://localhost:3456
export LLM_API_KEY=dario
export LLM_MODEL=anthropic/claude-sonnet-5

The anthropic/ prefix tells LiteLLM (OpenHands' router) to take the Anthropic path, which dario is now fronting. End-to-end walkthrough: openhands-walkthrough.md.

OpenClaw
export ANTHROPIC_BASE_URL=http://localhost:3456
export ANTHROPIC_API_KEY=dario
openclaw "task description"

OpenClaw's exec / process / web_search / web_fetch / browser tools are translated to Claude Code's set without a flag; a tool outside the map (message, for one) rides a fallback slot instead. Newer OpenClaw reads auth-profiles.json before env vars, so a stale key there wins — the walkthrough covers it.

Codex CLI · OpenAI SDK · any OpenAI-compatible tool
export OPENAI_BASE_URL=http://localhost:3456/v1
export OPENAI_API_KEY=dario

Ask for gpt-5.5 and it is served by your ChatGPT plan once you've run dario add altman. Ask for claude-sonnet-5 on the same URL and it is served by your Claude plan, translated both ways. Ask for gpt-4o or anything else your API-key backend lists and it goes there byte-for-byte. Names that don't look like OpenAI's (llama-3.3-70b, qwen-coder) need the provider prefix below; dario refuses them rather than guess:

dario backend add openai     --key=sk-proj-...
dario backend add groq       --key=gsk_...    --base-url=https://api.groq.com/openai/v1
dario backend add openrouter --key=sk-or-...  --base-url=https://openrouter.ai/api/v1
dario backend add local      --key=anything   --base-url=http://127.0.0.1:11434/v1

Force a backend with a prefix: openai:gpt-4o, claude:opus, groq:llama-3.3-70b, local:qwen-coder.

One holdover for old configs: six legacy OpenAI names (gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, gpt-5.3, gpt-4, gpt-3.5-turbo) sent to /v1/chat/completions are translated to Claude models when no other provider claims them.

Claude Agent SDK · Anthropic SDK (TypeScript, Python)
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({ baseURL: "http://localhost:3456", apiKey: "dario" });
import anthropic
client = anthropic.Anthropic(base_url="http://localhost:3456", api_key="dario")

Zero code change beyond the base URL. Streaming, tool use, prompt caching and extended thinking all pass through. More in usage.md.

curl
curl http://localhost:3456/v1/messages -H "content-type: application/json" \
  -d '{"model":"claude-sonnet-5","max_tokens":256,"messages":[{"role":"user","content":"Hello!"}]}'

curl http://localhost:3456/v1/chat/completions -H "content-type: application/json" \
  -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"Hello!"}]}'
Docker · Kubernetes · a Pi in a closet
docker volume create dario-config
docker run --rm -it -v dario-config:/home/dario/.dario ghcr.io/askalf/dario:latest login --manual
docker run -d --name dario -p 3456:3456 -v dario-config:/home/dario/.dario \
  -e DARIO_API_KEY="$(openssl rand -hex 32)" ghcr.io/askalf/dario:latest

The image binds 0.0.0.0, so a key is mandatory; without one dario refuses to start rather than become an open relay for your subscription. No console at all? Start empty with DARIO_ADMIN=1 and provision the first account over HTTP with the admin API. Two replicas sharing accounts need the refresh lock. Docker guide.

Something not listed? If it reads ANTHROPIC_BASE_URL or OPENAI_BASE_URL, or has a "Base URL" field, it works. The compatibility matrix says which tools are exercised end-to-end, which are inferred from a shared code path, and which are untested — one honest cell per tool.

What it does with a request

You point every tool at one URL. dario reads each request, decides which plan or backend owns it, and forwards it in that backend's native protocol.

Client speaksModelRoutes toWhat happens
Anthropic Messagesclaude-* / opus / sonnet / haikuClaude poolOAuth swap + Claude Code template, then api.anthropic.com
Anthropic Messagesa slug your ChatGPT account listsCodex engineMessages→Responses translation, subscription auth
Anthropic Messagesgpt-4o, llama-*, any name no plan listsRefused400 + x-dario-upstream-rejection: model_unroutable; reach an API-key backend from this shape with a provider prefix
OpenAI Chatgpt-* / o1-* / o3-* / o4-*OpenAI-compat backendAuth swap, body forwarded byte-for-byte
OpenAI Chata slug your ChatGPT account listsCodex enginechat/completions→Responses translation, subscription auth
OpenAI Chatclaude-*Claude poolOpenAI→Anthropic translation, then the Claude path
Either<provider>:<model>Forced by prefixExplicit override

The tool doesn't know. The backend doesn't know. dario is the seam.

The full Claude lineup, autodetected. Fable 5, Opus 5, Sonnet 5 and Haiku 4.5, plus [1m] long-context variants on every family except Haiku, by full id (claude-opus-5) or shortcut (fable / opus / sonnet / haiku; append 1m for the long-context form; opus48 / opus47 / opus46 / sonnet46 pin a generation and never float). GET /v1/models reads Anthropic's live catalog (TTL-cached, baked fallback offline), so a new model resolves the day it lands with no dario release, and the model-specific request shape is applied automatically. Families pulled upstream are filtered from both the live catalog and the fallback, so /v1/models never advertises a model that 404s. A name no provider lists at all, such as a ChatGPT slug your account doesn't have or a typo that belongs to no family, is refused locally with 400 and x-dario-upstream-rejection: model_unroutable instead of spending a pool request on an upstream 404. The guard steps aside for claude-* names (Anthropic's own 404 stays authoritative there), for requests under a --model / --fast-model override, for upstream-API-key mode, and for the legacy OpenAI names the built-in map translates.

Two plans, one endpoint

Your ChatGPT plan, on both endpoints

A ChatGPT Plus or Pro plan is served on both of dario's endpoints: any client that speaks /v1/chat/completions can use it (Codex CLI, the OpenAI SDKs, your scripts), and so can any client that speaks /v1/messages (Claude Code, the Anthropic SDKs, agent runtimes). The harness never needs to know which subscription is behind it.

dario add altman            # prints an authorize URL; paste the redirect URL back
dario codex list
dario codex remove altman

dario add altman names whose plan you are attaching; dario add amodei attaches a Claude account instead. The browser lands on a localhost page that doesn't load — expected, nothing is listening there. Copy the whole address bar and paste it at the prompt; dario reads the code out of it.

curl localhost:3456/v1/models | jq -r '.data[].id'
curl localhost:3456/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"gpt-5.5","messages":[{"role":"user","content":"hi"}]}'

# same subscription, Anthropic wire shape — this is what Claude Code speaks
curl localhost:3456/v1/messages -H 'content-type: application/json' \
  -d '{"model":"gpt-5.5","max_tokens":64,"messages":[{"role":"user","content":"hi"}]}'

Model names are discovered, not hardcoded. The set a ChatGPT subscription may use is per-account and moves; dario asks the backend which models this account lists, caches the answer, and advertises them on GET /v1/models. Anything not on that list (gpt-4o and friends) still routes to a configured API-key backend as before. codex:<model> / chatgpt:<model> forces the route.

Streaming, tool calls and tool-result round trips work on both shapes, and chat-shape image_url parts are carried as Responses input_image parts with detail preserved: dario translates chat/completions or Messages into the Responses API the subscription backend speaks, and translates the stream back into chat.completion.chunk or Anthropic message events. There is no /v1/responses inbound yet. The Codex backend does not accept every chat field, so response_format, stop, n, logprobs, stream_options and the sampling parameters temperature, top_p, max_tokens, max_completion_tokens are intentionally lossy; with --verbose, dario reports each field that does not reach Codex once per process. Codex accounts live in ~/.dario/codex-accounts/, separate from the Claude pool.

Prompt caching: the backend caches prompt prefixes of 1,024 tokens and up on its own; what dario adds is the prompt_cache_key that routes same-prefix requests to the cache that holds them, the way the Codex CLI does with its session id. A chat/completions client that sets its own key keeps it; an Anthropic-shape request gets one per Claude Code session (a hash of metadata.user_id, never the raw ids); anything else is keyed on its model, instructions and tool names, so repeated system prompts from any caller land together. Cached tokens come back as prompt_tokens_details.cached_tokens on chat/completions and as cache_read_input_tokens on /v1/messages, and show up in /analytics and the -v usage line like a Claude request's do.

Failover between subscriptions

Two consumer plans, no API keys, and neither one able to take you down on its own.

dario proxy --pool-fallback=gpt-5.6-sol,claude-sonnet-5

That is a chain, read left to right; each provider takes the first entry it can actually serve. Prefix every entry with a tier and the same flag becomes a tier map: --pool-fallback=haiku:gpt-5.6-luna,sonnet:gpt-5.6-terra,opus:gpt-5.6-sol picks one rung per request from the tier of the model asked for (default: catches the rest, otherwise the first rung), so heartbeat work on Haiku never overflows onto a flagship. That shipped in 6.0.17, the same day a Haiku-tier fleet spent 79% of a weekly allowance doing exactly that. When the Claude pool is drained or cooling, the request is served as gpt-5.6-sol from your ChatGPT subscription. When the subscription is rate-limited or down, the request is handed back to the Claude pool as claude-sonnet-5. Every substituted response carries x-dario-pool-fallback: <model> — a silently swapped model family is exactly the surprise this project exists to avoid.

A tool sends a request to dario. The Claude plan answers 429, so dario re-serves the same request from the ChatGPT plan, which answers 200, and the response returns to the tool carrying the x-dario-pool-fallback header.

A single-entry chain is one-way and means what it always meant, so an existing config is unaffected. Failover is opt-in: without --pool-fallback, a drained pool still returns its honest 429/503. Only a 429 or 5xx fails over; a 400 surfaces, because a bad request that fails over just reproduces itself on the other provider and buries the real cause. A 429 also cools that provider for a bounded interval, its retry-after if it sent one and 60 s otherwise, never longer than 15 min, and an entry that already declined is not asked again within the same request. When every entry is cooling, the request ends on one honest 429 with a retry-after instead of a retry storm. The Claude entry has to be a model the pool can actually serve, checked positively against the live catalog, so a typo can't trade a recoverable 429 for an unrecoverable 404.

dario doctor tells you which of these you are actually in:

[ OK ]  Failover   symmetric: gpt-5.6-sol → claude-sonnet-5, across 1 Codex account
[WARN]  Failover   armed (gpt-5.6-sol) but INERT — no Codex account and no backend
                   to fall back to. Add one: `dario add altman`

That warning is the whole reason the check exists. Armed with nothing to fall back to is green on every other check and incapable of doing anything.

[!NOTE] Upgrading from v5? Nothing to do. Every v6 feature is opt-in and a single-value --pool-fallback behaves exactly as it did. CHANGELOG

Shadow compare

Once either subscription can serve either wire shape, the interesting question stops being can I reach GPT and becomes which of these is better at my work. Benchmarks answer that badly. Your own traffic answers it well.

curl localhost:3456/v1/messages \
  -H 'content-type: application/json' \
  -H 'x-dario-compare: gpt-5.6-sol' \
  -d '{"model":"claude-opus-5","max_tokens":1024,"messages":[…]}'

You get the Claude answer, exactly as you would have. Beside it, dario runs the same prompt past gpt-5.6-sol and writes both to ~/.dario/compare/<timestamp>-<model>.json, in your own wire shape, so you are comparing like with like. The comparison cannot degrade the request it observes: it only reads bytes already on their way out, your request is never held open for it, and a comparison that fails, times out or has nowhere to go is dropped with the record still written. Both sides are stored as raw payloads, because extracting text is where a bug would quietly make two answers look more alike than they are.

Many seats, one endpoint

Every dario is a pool. A plain dario login is a pool of one; there is no separate mode to switch on. Hold more than one seat — a personal Max and a work Max, a couple of Pros, team seats — and the same localhost:3456 routes every request to whichever seat has the most headroom, live, per request.

dario accounts add work
dario accounts add personal
dario proxy
The dario TUI Accounts tab: a table of pooled seats (work, personal, side) with token expiry, 5-hour and 7-day utilization, and status.

Three things it does that a round-robin doesn't:

  • Per-model headroom routing. Anthropic meters each model family separately: a 5h bucket, a 7d bucket and a per-model 7d_<family> bucket. dario reads all of them off every response and routes each request by the bucket that governs it — an Opus call to the seat with Opus room, a Sonnet call to the seat with Sonnet room, independently. Plan tiers mix freely; dario cares about headroom, not tier.
  • Session stickiness. Claude's prompt cache is scoped to {account × cache key}, so rotating a long conversation across seats on headroom alone re-pays cache-create every turn, a 5–10× token-cost multiplier on the cached portion. dario pins each conversation to one seat (hashed from its first message, deterministic) for the life of the session and rebinds only when that seat is exhausted.
  • In-flight 429 failover. A seat hits its wall mid-request and dario retries the same request against the next-best seat before your client ever sees an error. The sticky binding follows, so the next turn doesn't re-select the cold one. A seat parked on a 429 rejoins the pool on its own when its window resets.
Three pooled seats, work, personal and side, each with a headroom bar. dario routes the request to the seat with the most headroom.

--pool-strategy=fill-first concentrates new conversations on one seat until it drains, for primary/backup setups. Refresh tokens expire about 28 days after the original grant regardless of rotation, so every seat's grant age is tracked and surfaced in dario accounts list, dario doctor and GET /accounts before it becomes a silent outage. Provision over HTTP with the headless admin API; pin one request to one seat with dario accounts check <alias> (admin API required: DARIO_ADMIN=1 and a DARIO_ADMIN_TOKEN). Internals and the live /accounts + /analytics endpoints: multi-account-pool.md; covered end-to-end by test/pool-e2e.mjs.

Watch it happen

Type dario with no arguments for a full-screen control panel: live request stream, per-model burn rate, rate-limit utilization per seat, billing-bucket breakdown, and an in-place config editor that writes ~/.dario/config.json. Pure ANSI, zero new runtime deps. Tab moves between tabs, r refreshes, R resumes a halted overage guard, q quits.

The dario TUI Analytics tab: requests per minute, tokens in and out, thinking tokens, average latency, subscription percentage, a per-model bar chart, per-account rate-limit bars for the 5-hour and 7-day windows, and a billing breakdown.

Both screenshots are rendered from the real TUI against a fixture proxy by scripts/readme/tui.mjs, so a layout change shows up here instead of rotting a mock-up. The numbers are illustrative; the pixels are not.

It tracks a moving target

Claude Code's request shape changes between releases — new betas, tool renames, per-model thinking configs — usually with no subscriber-facing note. dario doesn't guess that shape: it captures it live from your own installed claude binary on every startup, diffs it against each upstream release, and replays it faithfully. That's why your subscription routes the same through dario as it does through Claude Code itself: the request that leaves your machine is the shape your plan expects. Details: wire-fidelity.md · #13 · #14.

The installed claude binary feeds its request shape into dario. A timeline of Claude Code releases ends in a node flagged as drift.

Keeping that current is the whole job, and it's automated. These watchers run unattended; each badge is the live status of that workflow's latest run, and its label is the cadence:

WatcherCatchesLive
cc-drift-watchA new Claude Code npm release that changes the wire shape. Auto-drafts the fix; cc-drift-auto-release merges and ships it within minutes.hourly
cc-drift-template-watchSame-binary remote-config drift, which no npm diff can see. Runs against a live Claude session on a self-hosted runner and opens a rebake PR with the diff inline.hourly
cc-billing-classifier-canaryClassifier drift: one real request a day must still bill to a subscription bucket.daily
wire-drift-self-hostedPer-model beta headers and billing blocks the installed claude actually sends, model by model.daily
sdk-drift-watchAgent SDK / Stainless pins drifting from what the template assumes.daily
pricing-drift-watchdario's pricing table drifting from Anthropic's published rates, so the TUI's cost figures stay honest.daily
codex-drift-watchThe ChatGPT backend's model list or wire contract moving under the translator.daily
cc-oauth-healthThe maintainer's own production proxy going unhealthy on any axis.every 30 min
dario-doctor-watchRuntime drift only a live dario doctor --obedience surfaces: identity, obedience, usage buckets.every 6 h
deployed-version-watchPublishing is not deploying: is what's running what was last released?hourly
cc-drift-watcher-livenessThe watcher itself going quiet. Lives on GitHub-hosted infrastructure on purpose, so it survives the failures it watches for.every 2 h

Guarded at PR time by live-test, a required check that runs the full suite against a live proxy on a self-hosted runner, plus compat-test-self-hosted, which replays the compat suite through a passthrough proxy on wire-shape changes. A few changes the watchers caught and shipped fixes for, same day:

Change (no subscriber-facing note)Effectdario shipped
context-1m dropped from the default beta set on the OAuth pathSubscription requests default to the 200K window on Sonnet/Opusv3.38.3–4
thinking: {type:"adaptive"} gated per-model server-sideSonnet/Opus 4-5 400 every request through any proxyv3.38.5
Per-model anthropic-beta setsProxies sending one set diverge for non-Opus modelsv4.8.53

The full ledger lives in the CHANGELOG, 500+ releases since April 2026. Setup and walkthrough: drift-monitor.md. The residual manual cases — OAuth rotation, runner re-registration — are in the recovery runbook.

Guardrails

Overage guard

During normal operation, a subscriber should never see a single response billed outside their subscription pool. If one is, something is wrong — wire-shape drift, an account misconfig, a change upstream — and forwarding more requests in the same shape either bleeds real money (accounts with extra usage enabled) or returns a wall of rejections. The first hit is the signal; the rest are damage.

Requests from five tools are stopped short of dario, which is ringed in red and labeled halted, after the Claude plan returned a response billed as overage.

So the moment any upstream response bills to something other than your subscription pool, dario halts the proxy. The check is an allow-list, not a match on one string: anything that isn't a known subscription claim (five_hour / seven_day, their _fallback and _overage_included variants, and the chatgpt_subscription claim dario stamps on Codex-served responses) and isn't the unknown no-header sentinel trips it, so a billing bucket dario has never seen still halts. Subsequent requests return 503 with an Anthropic-shaped error body until you run dario resume, press R in the TUI, or the cooldown clears (default 30 min). The halt shows across the TUI, fires a best-effort OS notification, and emits named SSE events. Tune it via ~/.dario/config.jsonoverageGuard, or --overage-behavior=warn / --no-overage-guard / --overage-cooldown=<ms>. In upstream-API-key passthrough mode (ANTHROPIC_UPSTREAM_API_KEY) the guard is off; api billing is the point there. Verified end-to-end by test/overage-guard-e2e-live.mjs. Background: #288.

The billing split, a contingency dario is built for

On 2026-05-13 Anthropic announced that, from 2026-06-15, Agent SDK and claude -p (headless) traffic would leave the subscription pool for a small separate monthly credit, then metered API rates. They paused it before that date. Those surfaces still bill subscription today, and Anthropic says it will give advance notice before any revised version. Nothing changed; no credits were issued.

The split isn't live, but it was announced once on short notice and could return, so dario is built for it either way. Every request is rebuilt into interactive Claude Code shape before it leaves your machine (and, with --stealth, the response-correlated timing an interactive session has), so your traffic sits in the subscription pool whether a split is paused or live. The daily canary above is the tripwire: it surfaces a revived split within a day instead of on a surprise invoice. Verify on your own machine right now: dario doctor --usage fires one request and prints the rate-limit headers; representative-claim should read five_hour or seven_day, both subscription buckets. Full timeline: why-now-2026-06.md.

Trust & transparency

Everything inside the box labeled your machine: the tools and dario. Only two lines leave it, to the Claude plan and the ChatGPT plan, through a padlock. Pills read 0 deps, no telemetry, MIT.
SignalStatus
Source~33k lines of TypeScript across 68 files, auditable in a weekend. One credential path since v5: the pool.
Dependencies0 runtime. Verify: npm ls --production
ProvenanceEvery release SLSA-attested via GitHub Actions + Sigstore, published with OIDC trusted publishing — no long-lived npm token exists to leak
ScanningCodeQL on every push and weekly · ClusterFuzzLite fuzzes the SSE translator and rejection parsers weekly · OpenSSF Scorecard and Best Practices badges above are live
Tests178 test files run in parallel by npm test on Node 18, 20 and 22; the live e2e / compat / stealth suites have their own entry points. Green on every release
CredentialsYour own subscription tokens, never logged, redacted from errors, 0600 on disk in 0700 dirs
NetworkBinds 127.0.0.1 by default; upstream only to configured backends over HTTPS; hardcoded SSRF allow-list; refuses a non-loopback bind without DARIO_API_KEY
TelemetryNone. No analytics, no tracking, nothing phones home
This READMECI fails if the line count above drifts from src/ or a link or anchor here stops resolving (check-readme-line-count.mjs, check-readme-links.mjs); the TUI screenshots are rendered from the real TUI and the diagrams are briefed art, not screenshots (how)
npm audit signatures
npm view @askalf/dario dist.integrity
cd $(npm root -g)/@askalf/dario && npm ls --production

Security reports go to security@askalf.org, not a public issue: SECURITY.md. API stability commitments (@stable / @experimental / @deprecated, deprecation cycles): STABILITY.md.

Will my account get suspended?

The most common question about dario, and it deserves a straight answer: I can't promise you won't be actioned, and I'd be skeptical of anyone who does. Only Anthropic decides how it enforces its terms. What I can do is lay out exactly how dario works, so you can weigh the risk yourself instead of taking anyone's word for it.

What dario does:

  • Runs entirely on your machine. Your subscription token never touches my servers or anyone else's; requests go straight from your computer to Anthropic.
  • Authenticates as you, with your own Claude login, the same OAuth credential Claude Code itself uses. It impersonates nobody and shares nothing.
  • Doesn't modify your account, billing, or subscription settings.
  • Sends requests in the shape the official client sends them, rebuilt from your own installed binary, not spoofed from a hardcoded fake.
  • Reports nothing, anywhere. No telemetry, no analytics, nothing phones home; verifiable in the source, which is the point of keeping it auditable in a weekend.

What dario does that Claude Code doesn't: it lets tools other than Claude Code use that subscription. That's the whole point of it, and it's also the part that sits outside what Anthropic's own client does. Whether that falls within your plan's terms is Anthropic's call, not mine. Read their terms, read DISCLAIMER.md, and decide deliberately. dario is a transparency tool, in that it documents request behavior Anthropic doesn't publish for subscribers, and it is also, plainly, routing subscription traffic that Anthropic's own tools bill differently. Both are true; decide with both in view.

On policy risk specifically: Anthropic's position on third-party clients has moved before and can move again. dario is built to surface that fast rather than paper over it; see the billing split for the contingency already in place and the daily canary watching for it.

Ongoing discussion, including other users' experiences: #724.

Who it's for

Best fit: developers juggling multiple LLM tools and per-tool API keys · Claude Pro/Max subscribers who want their plan usable everywhere, not just in Claude Code · ChatGPT Plus/Pro subscribers who want their plan in OpenAI-compatible harnesses · teams running local or hosted OpenAI-compat servers who want one stable local endpoint · Agent SDK users who want subscription routing with zero code change · power users wanting multi-account pooling with 429 failover.

Not a fit: you need vendor-managed production SLAs (use the provider APIs) · you want a hosted multi-tenant team platform with dashboards and SSO (dario is a single-owner local proxy) · you want a chat UI (use claude.ai).

How it compares. Only one of these routes a consumer subscription; the others route API keys, and that is the whole split.

ToolWhat it isWhen it wins
darioLocal proxy that routes your Claude and ChatGPT plans, plus any OpenAI-compatible APIYou already pay for a plan and want every tool on your machine to use it
LiteLLMPython SDK + proxy, 100+ providers via API keys, enterprise featuresYou have API keys, want central spend controls, or run a hosted multi-tenant service
OpenRouterHosted aggregator, one API key for hundreds of modelsYou want model breadth and are fine with pay-per-token
Kong AI GatewayEnterprise on-prem API gateway for LLMsYou already run Kong and need AI traffic under the same governance

Longer version, with specifics: #68.

Commands

CommandWhat it does
darioThe TUI: status, config editor, analytics, hits, accounts, backends
dario login [--manual]Log in to your Claude plan. Picks up Claude Code's credentials or runs its own OAuth flow; --manual for SSH / containers
dario proxyStart the local endpoint on :3456
dario doctor [--usage] [--probe] [--obedience] [--auth-check] [--bun-bootstrap] [--json]One aggregated health report: runtime/TLS, template and drift, OAuth, pool, refresh-grant age, failover readiness, backends
dario add altman / dario add amodeiAttach a ChatGPT plan / a Claude account, by whose it is
dario accounts list / add / remove / check <alias>Pool management; check sends one pinned request per model through the running proxy (admin API on)
dario backend list / add / removeOpenAI-compatible API-key backends
dario codex list / add / removeChatGPT accounts (the long form of dario add altman)
dario usage · dario config · dario statusBurn rate for the last hour · effective config, redacted · token health
dario resume · dario refresh · dario logout · dario upgradeClear an overage halt · force a token refresh · delete credentials · safe self-update
dario mcp · dario subagent install / remove / statusReach dario from inside any MCP client, or from inside a Claude Code session, read-only
EndpointDescription
POST /v1/messages · POST /v1/chat/completionsThe two wire shapes, any plan behind either
GET /v1/modelsLive model list: the Claude catalog plus whatever your ChatGPT plan lists
GET /health · GET /livezServiceability (503 when not) · liveness. /health?probe=1 sends one real request
GET /status · GET /accounts · GET /analyticsOAuth detail · per-seat utilization and grant age · per-account / per-model stats and burn rate
POST /v1/messages/count_tokens · POST /v1/completeToken counting and the legacy Text Completions shape
GET /analytics/stream · GET /codexLive analytics over SSE · ChatGPT-seat status, read without spending or exposing a token
/admin/*Provisioning, GET /admin/accounts, POST /admin/resume; only with DARIO_ADMIN=1 (admin API)

Flags: commands.md, plus dario --help for the ones it doesn't list yet (--effort, --max-tokens, --model-alias, --fast-model, session rotation, concurrency caps, the pacing knobs behind --stealth) · env vars grouped by task, for Docker / k8s / systemd: configuration.md · SDK examples: usage.md.

More knobs — stealth timing, system-prompt modes, client-shape overrides, VPN egress, MCP
  • Behavioral stealth (--stealth). Adds when a request arrives to what it looks like: response-length-correlated think time and session-start latency. wire-fidelity.md
  • Recover output (--system-prompt=partial). Strips Claude Code's tone and verbosity constraints for 1.2–2.8× more output on open-ended work, without changing which pool you bill to. #183 · system-prompt.md
  • Client-shape overrides. --honor-client-thinking passes a client's own thinking block through; --preserve-output-format carries a client's output_config.format JSON schema through so structured-output SDKs get schema-constrained output. Both off by default.
  • Runs any agent. A 64-entry schema-verified TOOL_MAP pre-maps Cline, Roo, Kilo, Cursor, Windsurf, Continue, Copilot, OpenHands, OpenClaw and Hermes tool names to Claude Code's native set; MCP tools (mcp__server__tool) forward verbatim. Custom schemas: --preserve-tools or --hybrid-tools. agent-compat.md
  • Model aliases and caps. --model-alias=<name=target> (repeatable) advertises a name of your choosing on /v1/models and routes it; --effort=<low|medium|high|xhigh|ultracode|max|client> and --max-tokens=<N|client> set, or pass through, per-request effort and output caps.
  • VPN / egress routing. Route dario's upstream traffic through a VPN without putting the whole host on one. vpn-routing.md
  • More than one instance, same accounts. Refresh tokens are single-use, so two replicas refreshing the same seat leave one holding a dead token; the optional refresh lock (Redis or Cloudflare) makes the loser adopt the winner's credentials. multi-instance.md
  • PII redaction in front of dario. Pair it with cordon: integrations/cordon.md
  • Reachable from inside Claude Code or any MCP client. dario subagent install registers a sub-agent for in-session diagnostics; dario mcp exposes dario as a read-only MCP server. sub-agent.md · mcp-server.md

FAQ

Does this violate Anthropic's terms?

Mechanically, dario uses your existing Claude Code OAuth tokens: it authenticates you as you, with your subscription, through Anthropic's official endpoints. Whether any particular use complies with current terms is between you and Anthropic; consult their terms and your agreement. Independent, unofficial, third-party — see DISCLAIMER.md. On the suspension question specifically: Will my account get suspended?

Do I need Claude Code installed?

Recommended, not required. With it, dario login picks up credentials automatically and the template extractor reads your binary on every startup. Without it, dario runs its own OAuth flow and falls back to the bundled (scrubbed) template snapshot, which the drift watchers keep current.

Do I need Bun?

Optional, recommended: Bun's TLS ClientHello matches Claude Code's runtime, and dario relaunches itself under Bun when it finds one on PATH. Without it dario works fine on Node; dario doctor flags the mismatch and --strict-tls hard-fails until resolved.

Can I use dario without a Claude subscription?

Yes. Skip dario login, run dario add altman for a ChatGPT plan or dario backend add openai --key=… for an API key, and you have a local router with no Claude involvement. --no-claude-auth keeps the Claude token untouched entirely.

representative-claim: seven_day in my headers — am I downgraded?

No. five_hour and seven_day are both subscription billing, different accounting buckets in the same mode. overage is the one that flips you to per-token, and the overage guard halts on it. #1

My usage through dario is higher than through Claude Code directly. Why?

Almost always the prompt-cache TTL, not proxy overhead: dario mirrors whatever cache stamp your client sends, and many harnesses send the 5-minute one, so gaps longer than five minutes between turns re-create the prefix. DARIO_CACHE_TTL_1H=1 forces the 1-hour TTL. The full breakdown, with the two-message check that tells you which case you're in: faq.md.

Will the billing split break my setup?

Not today: it was announced, then paused before it took effect, and your traffic still bills subscription. What dario already does about it, and the canary that would catch a revival: the billing split.

Why "dario"?

It's a name, not an acronym. Don't overthink it.

Full FAQ, including per-tool 401s and Team/Enterprise plans: faq.md.

Deep dives

Contributing

PRs welcome. Small TypeScript codebase, zero runtime deps. Architecture, file-by-file map and the review bar in CONTRIBUTING.md; release mechanics in RELEASING.md.

git clone https://github.com/askalf/dario && cd dario
npm install
npm run dev    # tsx, no build step
npm test       # 178 files in parallel via test/all.test.mjs
npm run e2e    # live proxy + OAuth (needs a working Claude backend)

Drift and audit runners, none of them part of npm test:

npm run drift:wire     # compare a live Claude Code capture against the baked template
npm run drift:sdk      # Agent SDK / Stainless pin drift
npm run audit:tui      # drives the real TUI through a fake TTY at 12 geometries
npm run check:overage  # overage-classifier check against live headers
npm run stress         # concurrency / queue behaviour under load
npm run cch:calibrate  # re-derive the billing-tag cch seed for a new Claude Code build
npm run readme:assets  # regenerate the diagrams and TUI screenshots above

Two easy ways to help beyond code: star the repo, the clearest signal this is useful, and file drift: open an issue when a rate-limit header flips or a tool that worked yesterday breaks today, and it gets documented in public alongside the fix. Follow @ask_alf for drift bulletins as they land.

Star history of askalf/dario

Contributors

WhoContributions
@GodsBoyProxy auth, token redaction, error sanitization (#2)
@belangertradingBilling-classification investigation (#4, #6, #7, #12, #23), multi-agent billing FAQ (#27)
@wysieESM require crash in dario login (#15), OAuth for Max-plan accounts (#18)
@earlvanzeOpenClaw tool mappings (#19), OAuth manual override (#47), HTTPS warning (#53)
@nathan-widjajaREADME positioning structure — the promise → who → first use → why-switch spine the page still runs on (#21)
@trinhnvgemOAuth login failures on first release (#22), container and headless callback binding (#28)
@adubkovThe container / headless-SSH case behind the manual OAuth code paste (#28)
@iNicholasBEmacOS keychain credential detection (#30)
@boeingchocoReverse tool-param translation (#29), SSE framing regression catch, hybrid-tool motivation (#33, #36)
@tetsucoScrubber path corruption (#35), OpenClaw reverse-mapping collisions (#37), 20x-tier report (#42)
@mikelovattSilent subscription-drain surfaced via friendly billing buckets (#34)
@ringge--no-auto-detect for text-tool auto-preserve (#40)
@rustanacexdCursor BYOK routing for Claude, and --effort=max (#190)
@daimonbotOfficial multi-arch Docker image on GHCR (#199)
@Saik0sWildcard CORS allow-headers, Opus 4.7 catalog entry (#222)
@lwsh123kcch anchored to the billing tag instead of first match (#528)
@boredlandTime-to-reset in dario doctor --usage (#550)
@pnewell--preserve-output-format for structured-output SDKs (#583)
@matteo-ramaHeadless admin bootstrap (#599), Analytics NaN and per-account rate-limit rows (#600), pool-aware /status and /health (#636), version on both (#640), Accounts TUI reads the live pool (#641)
@miklisantonMid-session /model switch 400 (#744), empty-turn guards behind the subagent 400s (#1033, #1117)
@p-i-Independent wire-fidelity audit with a re-runnable harness — version-blind bun-match, and the correction to the packet-identical claim (#813)
@jerzydziewierzTUI Config tab clipping and scrolling (#861)
@ramarro123Admin bulk re-auth (#913), shared state across instances (#993), prompt-cache behaviour under litellm (#1018), parked-seat and shared-window reporting (#1244)
@zytegalaxyThe ChatGPT/Codex engine and dario add altman (#1009)
@chaogebabaAuto-release must never fire from a fork (#1029)
@robincleUtilisation freshness — lastObservedAt / utilAgeMs on /accounts (#1032)
@anupammeRefresh-lock ownership by server-issued lock id (#1059)
@LiveNathanNever send or stamp empty text blocks (#1067), empty final user turn from CC's stream-interruption retry (#1092, as @NathanLively)

Disclaimers

dario is an independent, unofficial, third-party project. Not affiliated with, endorsed by, or sponsored by Anthropic, OpenAI, or any vendor referenced here. Provided as-is, no warranty. You are solely responsible for compliance with your subscription's terms, the security of your credentials, and the content you send through the proxy. Not for safety-critical, regulated, or production environments without your own review. Full text: DISCLAIMER.md.

License

MIT — see LICENSE and DISCLAIMER.md. The embedded README font is Space Mono under the SIL Open Font License.

Own Your Stack

dario is the routing layer of Own Your Stack, open tools for owning your AI infrastructure instead of renting it by the token. One subscription. Your box. Your terms.

Built by Thomas Sprayberry

dario is part of Own Your Stack, the open toolkit behind Sprayberry Labs, the software studio with one human on staff, run by askalf, the AI operation these tools are part of.

Built in the open, scars included. Follow the build: @ask_alf · sprayberrylabs.com/own-your-stack

Footnotes

  1. Pro at $20 a month, Max 5x at $100, Max 20x at $200, as listed on claude.com/pricing on 2026-09-06. Annual billing is cheaper; check the page for what's current.

Contributors

askalf

983 commits

dependabot[bot]

51 commits

earlvanze

3 commits

Languages

JavaScript

60.4%

TypeScript

39.0%