aayushpokhrel1/delegation-pipeline

85.6% fewer orchestrator tokens, measured: offload diff-verifiable grunt work from your paid Claude Code session to free or cheap models. Zero-dependency Python, OpenAI-compatible, ships as a Claude Code plugin.

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Delegation Pipeline – offload Claude Code grunt work to cheap models (86% fewer tokens) (r/SideProject)

Claude Code usage goes fast, and a lot of mine was going into work that didn't need a frontier model: docstrings, mechanical refactors, filling in a module from a contract I'd already specified. I thought of this project as leveraging what free resources in AI we have and having opus be the mind…

1

Oct 2, 2026

README

Delegation Pipeline

GitHub stars License: MIT Python 3.8+ Zero dependencies Claude Code plugin PRs welcome Orchestrator tokens saved

Offload token-heavy grunt work from your paid Claude Code session to free or cheap models, so your Claude usage goes to the thinking, not the typing.

delegate is a zero-dependency Python worker (stdlib only) that runs a small agentic loop over any OpenAI-compatible endpoint. It can read, search, and edit files in the current repo. It cannot run shell commands or use git, by design: you (or your Claude orchestrator) review the resulting git diff and commit.

~/.claude/bin/delegate [backend] "<task>"

The backend is optional and defaults to free ($0). When free cannot be reached, the run escalates once to deepseek automatically, so you do not have to guess a tier per task. See Backends.

demo

Measured impact

Generated 2026-10-02 from 51 delegated runs across 4 repos.

MetricValue
Runs delegated51
Tokens offloaded to workers27,147,998 (measured)
Worker API calls998
Paid on the worker tier$7.60
Claude tokens avoided~23,238,686 (estimated)
Opus-equivalent value~$348.58

By backend:

deepseek  ████████████████████  100%   27,147,998 tokens

By month:

2026-09  ████████████████████   20,873,682 tokens
2026-10  ██████░░░░░░░░░░░░░░   6,274,316 tokens

Offloaded tokens and worker calls are measured from each backend's own usage fields. "Claude tokens avoided" is an estimate: 85.6% of the offloaded total, the median of 9 benchmark tasks in bench/RESULTS.md whose individual savings ranged from 60.7% to 89.7%. Reviewing a worker's diff costs about 14.4% of doing the task inline. Dollar figures are rough blended per-tier prices, for scale, not billing.

That ratio was measured with deepseek-chat on both sides of the benchmark: the same model did each task inline and reviewed the worker's diff. It is applied above to worker tokens as a stand-in for what the orchestrator model would have spent on the same work, which is an assumption the benchmark does not test. Read the direction as sound and the absolute figures as indicative.

Every figure above can be recomputed from the redacted ledger committed at bench/ledger-snapshot.jsonl, using this exact command:

python delegate.py --stats --ledger bench/ledger-snapshot.jsonl

Proof: it measurably cuts orchestrator tokens

Not just a claim. bench/ runs a reproducible A/B where every task is done two ways and measured from real OpenAI-compatible usage fields, and each result is gated by a pytest that fails until the work is actually correct, so a row only counts if the change really worked.

savings% = 1 - T_review / T_do: when it delegates, the orchestrator pays only T_review (write a brief + skim the diff) instead of T_do (doing the whole task itself). Both figures below were measured with deepseek-chat on both sides, so they are deepseek reviewing versus deepseek doing, not Opus versus Opus. Prices scale across models; token counts do not, so read the direction as sound and the absolute number as indicative. See bench/README.md, under "What it measures".

Median: 85.6% fewer orchestrator tokens per task across 9 tasks, spread 60.7-89.7%, with 71,116 tokens of real coding work pushed onto the cheap tier. Every task passed both arms (inline ok / worker ok).

The per-task table lives in bench/RESULTS.md and is not copied here on purpose. It used to be, and the copy silently drifted: it still showed a 4-task run with a "74-93% band" long after the benchmark widened to 9 tasks spanning 60.7-89.7%. RESULTS.md is generated by bench/run_bench.py, and a test asserts it agrees with REVIEW_RATIO in delegate.py, so that one is checked and a hand-maintained second copy never could be.

A real datapoint from building this benchmark: bench/run_bench.py itself (~250 lines, stdlib only) was written by a deepseek worker that consumed 109,058 tokens; the Opus orchestrator paid only the brief plus one diff review, the same asymmetry, on the real orchestrator, on a non-toy file.

Reproduce it in under a minute:

pip install pytest
python bench/run_bench.py --worker deepseek --baseline deepseek   # or --worker free

Every run also prints a per-task TOKENS: total=… trailer (work that ran off your subscription) and rewrites bench/RESULTS.md. Full methodology and the honest caveats live in bench/README.md.

Backends

BackendProviderCostNotes
freeOmniRoute (local)$0Local gateway, auto/coding model
nvidiaNVIDIA API Catalogfree creditsbuild.nvidia.com; strong models, free tier
openrouterOpenRouter APIfree / cheapPin a specific :free or vision model; :free ids are rate-limited
deepseekDeepSeek APIcheapReliable; good default for real work
kimiKimi / Moonshot APIpremiumStrongest; use when quality matters

All of them are just an OpenAI-compatible base URL + model + key, configured in ~/.claude/delegate.config.json.

Default tier and escalation

Naming no backend gets you free. If free fails before the worker has edited anything, the run escalates once to deepseek and says so on stderr:

delegate: free failed with no edits made, escalating to deepseek: ...

Why it works this way, measured on bench/tasks.json on 2026-10-01: reviewing a free worker's diff costs the same as reviewing a paid one (85.4% vs 85.6% median savings), so free is not the lower-quality tier, only the less reliable one. It completed 5 of 9 tasks run back-to-back, while the 4 that failed all passed when retried alone, because the OmniRoute pool drains under sustained load. Defaulting to free and escalating on failure gets the $0 attempt without the caller having to predict which tasks it will finish.

Two deliberate limits:

  • Escalation only happens on a zero-edit failure. Once the worker has written to the tree, the next backend would start from a half-applied change nobody reviewed, so the run stops and tells you to look at the tree instead.
  • --model disables escalation. A model id belongs to one backend, so carrying it to another would 404. Pinning a model means you chose the tier too.

Both attempts are written to the ledger: the failed one with "ok": false and "escalated_to", the successful one with "escalated_from". A tier showing 0 runs and 0 failed attempts was never routed to at all, which the ledger previously could not distinguish from a tier that was tried and died.

Install

The /delegate + /orchestrate commands and the delegate / orchestrate skills ship as a Claude Code plugin, which makes them available in every project (and the skills model-invokable, so "orchestrate it" works anywhere):

claude plugin marketplace add aayushpokhrel1/delegation-pipeline
claude plugin install delegation-pipeline

(Or the interactive /plugin marketplace add ... and /plugin install ....)

That is the whole install. The plugin carries the commands, the skills, and the worker itself: a SessionStart hook runs scripts/ensure_launcher.py, which writes the ~/.claude/bin/delegate launcher pointing at the plugin's own copy of delegate.py. It appears at the start of your next Claude Code session, with no clone and no installer.

The backends still need their own setup, which is the one manual step: the free tier wants a local OmniRoute gateway running, and the remote tiers want an API key. See Keys and Autostart.

Clone-based install

The clone-based install is optional and aimed at people who want to work on this repo:

bash install.sh
./install.ps1

This writes a delegate launcher into ~/.claude/bin/ pointing at your checkout rather than the plugin's copy, so your edits to delegate.py take effect immediately. It also seeds ~/.claude/delegate.config.json from the example and installs the post-commit plugin-refresh hook described below. Re-run after moving the repo. It does not copy the slash commands, those ship with the plugin above.

The session-start hook will not undo this: it leaves any launcher alone whose target still exists, and only writes one when none exists or when the path it points at has gone away.

If you develop this repo, point the marketplace at your checkout

A GitHub marketplace pins a commit, so the installed skills can silently lag your working copy by weeks, and stale skill prose is live instruction in every session. Use a directory marketplace instead:

claude plugin marketplace remove delegation-pipeline
claude plugin marketplace add /path/to/your/checkout
claude plugin install delegation-pipeline

Removing a marketplace uninstalls its plugins, so the install is required, not optional.

claude plugin update does not refresh a directory marketplace. It compares the version in .claude-plugin/plugin.json against the installed version and skips the copy when they match, reporting updateOutcome: "up_to_date" however much the files changed. The version string is the cache key: installs land in ~/.claude/plugins/cache/<marketplace>/<plugin>/<version>/. So either bump the version, or sidestep version bookkeeping:

claude plugin uninstall delegation-pipeline && claude plugin install delegation-pipeline

install.sh / install.ps1 wire scripts/post-commit-refresh-plugin.sh into .git/hooks/post-commit so that second form runs for you whenever a commit touches skills/, commands/ or .claude-plugin/. A written reminder would rot the same way the pinned clone did, and these files are instructions: stale ones do not break, they steer. The hook no-ops unless this checkout is the marketplace source, so it is harmless on a machine installing from GitHub. DELEGATE_SKIP_PLUGIN_REFRESH=1 disables it.

Install copies your working tree, not HEAD: uncommitted edits and even gitignored paths (__pycache__/, graphify-out/) land in the installed copy. Convenient while iterating, but it means a half-finished SKILL.md becomes live guidance, so finish an edit before reinstalling.

One piece, not two. The plugin ships the commands, the skills, and the worker, and sets up its own launcher. A clone only changes which copy of delegate.py that launcher points at, which is what you want while developing this repo and nothing you need otherwise.

Keys

  • free: needs a local OmniRoute gateway and, on current OmniRoute (3.8.50+), an API key. Set up autostart (below) so the gateway is always there, then create a key at http://localhost:20128/dashboard/api-manager and put it in delegate.config.json:

    "free": { "base_url": "http://localhost:20128/v1", "api_key": "...", "model": "auto/coding" }
    

    The key is shown once at creation (ALLOW_API_KEY_REVEAL=false), so copy it there and then. Older docs (and OmniRoute's own startup banner, plus REQUIRE_API_KEY=false in its .env) claim the local gateway is keyless. It is not: /v1/* always requires a Bearer key, and Settings > Security states that authorization model with no toggle to turn it off. Without a key delegate sends no Authorization header at all and every call fails 401 invalid_api_key, which looks exactly like "OmniRoute is down".

  • deepseek / kimi: set env vars, or paste the key into delegate.config.json (that file is git-ignored):

    export DEEPSEEK_API_KEY=sk-...
    export MOONSHOT_API_KEY=sk-...
    
  • nvidia: get a free key at https://build.nvidia.com (any model page → "Get API Key", it starts with nvapi-), then:

    export NVIDIA_API_KEY=nvapi-...
    

    The nvidia backend is OpenAI-compatible via https://integrate.api.nvidia.com/v1. It defaults to nvidia/nemotron-3-super-120b-a12b (verified tool-calling); override per run with --model. The catalog changes fast (models EOL or are account-gated, and the previous deepseek-ai/deepseek-v4-flash-0731 default now hangs), so use --list-models and see MODELS.md. This is a great free path when OmniRoute's routes are dry.

    You can also plug NVIDIA into OmniRoute itself (so the free/auto router can use it): open http://localhost:20128 → provider keys → add the NVIDIA key. Either way works; the direct nvidia backend is the more predictable of the two.

  • openrouter: get a key at https://openrouter.ai/keys, then:

    export OPENROUTER_API_KEY=sk-or-...
    

    OpenAI-compatible via https://openrouter.ai/api/v1. Defaults to nvidia/nemotron-3.5-lightning:free; free models carry a :free suffix and are rate-limited (~50 requests/day at $0, ~1000/day with ~$10 of credits). The worker needs tool calling and the catalog churns, so --list-models first and see MODELS.md. OmniRoute's free backend already routes through OpenRouter, so reach for this direct backend when you want to pin a specific fast/free or vision model instead of auto-routing.

Autostart

So you never have to start the gateway by hand. Logs land in ~/.claude/omniroute.log, and every installer prefers a global omniroute (npm i -g omniroute) and falls back to npx --yes omniroute.

Install OmniRoute globally, don't rely on the npx fallback. npx re-extracts the package and discards Next's .next build cache on every launch, which maximises the cold-start cost below. On npm 11+ the global install silently skips lifecycle scripts, leaving native deps (koffi, keytar, onnxruntime-node, esbuild) unbuilt; if you see an install-scripts warning, re-run with the --allow-scripts=<list> that npm prints.

The starter warms the gateway's routes before exiting. OmniRoute ships as a Next.js dev server, so every route compiles on its first request after a restart: cold on Windows, /api/monitoring/health takes over 90s and /v1/models about 10s, both under 0.1s once warm. Whichever client calls first otherwise pays that bill, times out, and reports the gateway "down" while it is really just compiling. OmniRoute's own CLI hits this too: its readiness probe hardcodes 60s, so npx omniroute often prints "Server did not respond within 60s" about a server that comes up fine seconds later. Treat that warning as noise, and check curl http://127.0.0.1:20128/api/monitoring/health instead.

Because the warm-up waits out that compile, a cold start-omniroute.ps1 can take a few minutes before it exits. That is fine for a hidden logon task where nobody is waiting, and it is the whole point: the task eats the latency so your sessions never do.

The Windows task runs the starter from this repo's path, so moving or renaming the checkout breaks autostart. Re-run install-autostart.ps1 afterwards.

Windows (Scheduled Task, launches hidden at logon):

powershell -ExecutionPolicy Bypass -File scripts\install-autostart.ps1
Start-ScheduledTask -TaskName "OmniRoute Gateway"                 # start now, no reboot
powershell -ExecutionPolicy Bypass -File scripts\uninstall-autostart.ps1   # remove

macOS (launchd agent, RunAtLoad + KeepAlive so it restarts if it dies):

bash scripts/install-autostart.sh          # starts immediately and at every login
bash scripts/uninstall-autostart.sh        # remove

Linux (systemd --user service, restarts on failure):

bash scripts/install-autostart.sh          # enable + start now, and at every login
sudo loginctl enable-linger "$USER"        # optional: keep running without an active login
bash scripts/uninstall-autostart.sh        # remove

All installers are idempotent, they won't start a second copy if one is already running. For a one-off manual start on macOS/Linux without installing autostart, use bash scripts/start-omniroute.sh.

OmniRoute ships with only a couple of keyless free providers (OpenCode Zen, Felo). Their shared quota is frequently exhausted (403 insufficient_quota / 429), which makes the bare free backend unreliable. Add your own free-tier API keys so the router has healthy routes to fall back to. All of these give a free key at signup:

ProviderFree tierGet a key
GroqFast, generous free tierhttps://console.groq.com/keys
Cerebras~1M tokens/day freehttps://cloud.cerebras.ai
Google GeminiFree tier (Flash models)https://aistudio.google.com/apikey
MistralFree experiment tierhttps://console.mistral.ai
OpenRouter:free modelshttps://openrouter.ai/keys
GitHub ModelsFree (rate-limited)GitHub settings → developer token

Add them in the OmniRoute dashboard (open http://localhost:20128 → provider/keys page, keys are encrypted at rest), or via CLI (omniroute keys). Groq + Cerebras + Gemini alone make free solidly usable for grunt work. When free is dry, deepseek remains the reliable paid-but-cheap fallback.

The worker also sends a normal User-Agent (some Cloudflare-fronted free providers return 1010 browser_signature_banned to the default Python-urllib signature) and retries transient provider errors (401/403/429/5xx). Because the free pool is load-balanced, a retry re-rolls to a different provider, so one bad route no longer fails the whole run.

Use it

# smoke test (needs OmniRoute running)
~/.claude/bin/delegate free "Summarize what this project does in 3 bullets"

# real grunt work
~/.claude/bin/delegate deepseek "Add type hints to every function in src/parser.py"
~/.claude/bin/delegate free --model auto/cheap "Write a docstring for each function in utils/"
~/.claude/bin/delegate nvidia "Convert callbacks to async/await in api/client.py"

# vision: attach an image (needs a vision-capable model)
~/.claude/bin/delegate openrouter --model inclusionai/ling-3.0-flash-vl:free \
  --image mockup.png "Build the login form in src/Login.jsx to match this mockup"

Step logs stream to stderr; the final summary prints to stdout.

Flags: --dir <path> (repo root, default cwd), --model <id> (override), --image <path|url> (attach an image, repeatable, needs a vision model; the worker can also fetch images mid-task via its view_image tool), --max-steps N, --list-models (print the backend's catalog and exit), --stats (print the running savings tally and exit), --ledger <path> (read the ledger from somewhere else, e.g. the published snapshot), plus --verify / --commit (below).

The usage ledger

Every delegate run appends one line to ~/.claude/delegate-usage.jsonl (override the path with the DELEGATE_LEDGER env var). The ledger accumulates across all sessions and all repos, so it is a running record of everything you have offloaded, not just the current run.

~/.claude/bin/delegate --stats

--stats reads the ledger back and prints the tally: total offloaded tokens and worker calls, what you paid on the worker tier, the Claude tokens avoided, an Opus-equivalent value, verify failures and commits, plus breakdowns by backend, model, repo, and month.

The offloaded-token count is measured from the backend's own usage fields. The "claude tokens avoided" figure is an estimate: it applies the median review ratio from bench/RESULTS.md (reviewing a worker's diff costs about 13.9% of doing the task inline, so roughly 86.1% of the offloaded tokens never reach your subscription). The dollar figures are rough blended per-tier prices, useful for a sense of scale, not billing.

The file is plain JSONL, one run per line, so it can be grepped, piped into jq, or deleted freely. Deleting it just resets the tally; nothing else depends on it.

bench/ledger-snapshot.jsonl is a committed, redacted copy of that ledger: repo names are stripped, every token count is intact. It is refreshed by the same delegate --stats --readme run that regenerates the block below, so the two cannot drift. That means anyone can recompute the published table from the repo alone:

python delegate.py --stats --ledger bench/ledger-snapshot.jsonl

Redaction costs exactly one thing: with repo names stripped, a recompute from the snapshot reports a single repo instead of the real count. Every token figure matches to the digit.

delegate --stats --readme regenerates the "Measured impact" block at the top of this README from the ledger. It never includes repository names, only a count of how many repos are involved, so the block is safe to publish. A weekly scheduled task can keep it current: scripts/install-stats-task.ps1 registers it on Windows.

Letting Claude pick the model

The tool never auto-selects a model beyond each backend's default, that's the orchestrator's job. When you say "delegate" (or run /delegate, or Opus drives the raw CLI on its own), Claude sizes up the task and chooses both the backend and the model/route, using --list-models to see the live catalog and MODELS.md as the shortlist. Two flavors of choice:

  • free (OmniRoute) → Claude picks an auto/* route (auto/coding, auto/cheap, auto/fast, ...) and lets OmniRoute's router pick the concrete model, so it fails over across your provider keys. Prefer routes over pinning one model here.
  • nvidia / paid backends → Claude picks a concrete model id (there's no router in front), restricted to tool-calling-capable models since the worker edits via function calls.

You can always force a specific backend/model/route by naming it: delegate free --model auto/cheap "..." or delegate nvidia --model nvidia/nemotron-3-super-120b-a12b "...".

From inside Claude Code: /delegate

The installer also drops a /delegate slash command into ~/.claude/commands/, so you can trigger a delegation without leaving your Claude session:

/delegate deepseek Add type hints to every function in src/parser.py
/delegate free Write a docstring for each exported function in utils/

The command has Claude expand your request into a tight, self-contained spec, run the worker, then review the resulting git diff and report back. First token is the backend (free if omitted); the rest is the task.

A free worker can take minutes, since the free model is the slow part. Run it with a long foreground timeout so it isn't detached to the background mid-run. If your shell or tool does background it at a default timeout, that only cuts off the wait: the worker keeps running and its file writes still land. Just confirm the run finished before treating the diff as complete, reading files while a backgrounded worker is still going can show a half-applied tree. deepseek/nvidia are faster if you want to avoid the wait entirely.

How Claude should drive it

The orchestrator protocol lives in your global ~/.claude/CLAUDE.md. In short:

  1. Write a tight, self-contained instruction (name the files, describe the change, point at a pattern to mirror). The worker has no conversation context.
  2. Delegate on a clean tree so the resulting git diff is attributable.
  3. Review the diff. Fix small issues yourself; re-delegate with a sharper spec if it drifted. Workers never commit.
  4. You run tests and commit, after review.

Route by complexity, the gate is a tight verifiable spec, not low difficulty: free for trivial/mechanical work (boilerplate, repetitive edits, docstrings, stubs); deepseek (roughly Sonnet-class) for substantial well-specified work (a whole module, a real refactor, a non-trivial first draft), do not cap delegation at boilerplate. Keep design, tricky debugging, security-sensitive code, deep-context work, and one-liners in the Claude session.

Proactive delegation (no /delegate needed)

The orchestrator (Opus) is set up to delegate on its own whenever a request contains work it can spec and verify, mechanical or substantial, without you typing /delegate. It announces the backend/model in one line, runs the worker, reviews the diff, and folds the result in. /delegate and saying "delegate this" still work as manual triggers; they're just not required. Say "don't delegate this" to keep a task (or a sensitive repo) in-session.

Note: delegating sends repo content to external model providers (deepseek/nvidia are remote; free/OmniRoute routes out too). That's the intended trade; disable it per-repo when the code is sensitive.

This behavior lives in your global ~/.claude/CLAUDE.md, not in this repo, so it travels with that file, not with a clone. To enable it on another device, add a block like this to that device's ~/.claude/CLAUDE.md:

# Delegation to cheap-model workers
Use `~/.claude/bin/delegate <backend> [--model <id>] "<task>"` to offload work.
Delegate proactively (no /delegate needed): announce the backend/model in one line, run
the worker, then review the git diff. Route by COMPLEXITY (the gate is a tight verifiable
spec, not low difficulty): `free` for trivial/mechanical work; `deepseek` (roughly
Sonnet-class) for substantial well-specified work like a whole module or a real refactor,
do not cap it at boilerplate; `nvidia` when free is dry (tool-calling models only, see
MODELS.md). Keep design, tricky debugging, security-sensitive code, deep-context work, and
one-liners in-session. Stop if told "don't delegate this".

(The full version is in this repo's git history / the author's own CLAUDE.md.)

Combined Orchestration

The delegation pipeline and Claude subagents are two ways to move work off the paid orchestrator. Combined orchestration routes each task to the cheapest tier that can do it, so Claude tokens go to planning and judgment, not typing.

The model

One orchestrator, two worker tiers.

  • Orchestrator = Claude Opus in a Claude Code session. Holds the plan, writes a tight per-task brief, routes each task, and adjudicates reviews. It is the only component that spends Claude tokens, and it never writes implementation code itself.
  • Tier B = this delegation pipeline (free / deepseek / nvidia). Zero Claude tokens. This is the default tier for both implementation and review.
  • Tier A = Claude subagents (the Agent tool: haiku / sonnet / opus). Costs Claude tokens. An escape hatch, not a default.

Routing table

Route each task on three axes. Any single axis landing in the right column sends the task to Tier A; otherwise it stays on Tier B.

AxisTier B (default, no Claude tokens)Tier A (escalate, costs tokens)
Complexityanything specifiable -> delegate (defaults to free, escalates to deepseek on failure); force the paid tier with delegate deepseek when a task must not be retriedneeds broad codebase judgment -> subagent sonnet / opus
Iteration depthverifiable in ~one shot (orchestrator runs verify once, commits)long autonomous run-fail-edit loop, or needs a live service / Docker / the app running
Sensitivityordinary codesecurity-sensitive, or full-context debugging (Tier A, or the orchestrator itself)

Reviews default to Tier B too: a deepseek worker reviews the diff against the brief for zero Claude tokens.

Cost gate. Delegation is not free of the orchestrator's tokens: the brief, the review, and any re-run all cost Claude tokens. Before routing a task to Tier B, the orchestrator weighs that overhead against doing it inline, and keeps the task in-session when the brief plus review would cost as much as the edit itself. The pipeline only saves tokens when the work is bulkier than its description, which is why one-liners stay in-session.

The loop (per task)

  1. Brief. The orchestrator writes a tight, self-contained brief: exact files, the change, a pattern to mirror, exact values. The worker has no conversation context.
  2. Route and run. Pick a backend by the table and hand off. The worker edits; the CLI verifies and commits on green:
    ~/.claude/bin/delegate deepseek \
      --verify "npm test" \
      --commit "feat: <what changed>, per brief" \
      "<the full brief>"
    
  3. On red. A failed verify exits 2 and commits nothing, printing the command output. The orchestrator reads that output and either re-delegates with the failure folded in as a sharper spec, or escalates the task to a Tier A subagent.
  4. Review. A second worker reviews the committed diff against the brief, for zero Claude tokens:
    git show HEAD > review.diff
    ~/.claude/bin/delegate deepseek \
      "Review the diff in review.diff against the brief in brief.md. Report spec compliance and code quality, most severe first. Do not edit anything."
    
    The orchestrator adjudicates: accept, fix inline, or re-delegate.
  5. Escalate. Go to Tier A only when the loop cannot converge, or the task hits one of the axes above.

The verify/commit flow

--verify "<cmd>" runs the command after the worker finishes editing. A non-zero exit prints the output and exits 2 without committing. --commit "<msg>" stages and commits only the files the worker edited, but only when verify passed (or no --verify was given). Together they make a delegated task self-contained: the orchestrator hands off a brief and gets back a tested, committed result, without babysitting the test-and-commit cycle and without spending Claude tokens on it.

# mechanical, free tier, verified and committed in one shot
~/.claude/bin/delegate free \
  --verify "python -m pytest tests/test_utils.py -q" \
  --commit "test: cover utils edge cases" \
  "Add the three missing edge-case tests to tests/test_utils.py, mirroring the table-driven style already there."

# check the outcome
echo $?   # 0 = verified and committed; 2 = verify failed, nothing committed

On Windows, --verify runs through cmd.exe, so pass one command (npm test, npx tsc --noEmit), not a bash && / ; chain.

The trust boundary holds. The worker model still never runs shell or git. Only the caller's --verify and --commit flags run commands, and only the orchestrator sets them. See Tools the worker has.

The /orchestrate command and the orchestrate skill encode this whole table and loop, so an orchestrator invokes one thing instead of re-deriving it each session.

Tools the worker has

list_dir, find_files, read_file, search_text, write_file, edit_file. No shell, no network beyond the model endpoint, no git. File access is sandboxed to the working directory. The --verify / --commit flags are the one exception, and they are run by the CLI harness (the caller), never by the worker model.

Portability

Everything the worker needs is one delegate.py file and the stdlib, so it runs on any device with Python 3.8+. Sync = clone this repo + run the installer. Real keys live in ~/.claude/delegate.config.json (git-ignored), never in the repo.

Same spirit, different mechanism. The graphify skill builds knowledge graphs; its semantic extraction (docs, papers, images) otherwise dispatches Claude subagents, which spends your Claude tokens on the initial build. Set a Gemini key once and graphify routes that work to Gemini instead:

[Environment]::SetEnvironmentVariable("GEMINI_API_KEY", "<your-key>", "User")
[Environment]::SetEnvironmentVariable("GOOGLE_API_KEY", "<your-key>", "User")
export GEMINI_API_KEY="<your-key>"    # or GOOGLE_API_KEY; graphify accepts either

Code-only corpuses skip semantic extraction entirely and never cost subagent tokens. Already-running Claude Code sessions must be fully restarted to pick up a newly set key. Details live in the graphify skill's SKILL.md ("Save Claude tokens: set a Gemini key"). Get a key at https://aistudio.google.com/apikey.

This repo is graphify-integrated

This repo ships a graphify knowledge-graph integration so Claude consults the graph before falling back to raw search, and keeps it current automatically:

  • CLAUDE.md: guidance telling Claude to run graphify query/path/explain before answering codebase questions, and graphify update . after code changes.
  • .claude/settings.json: a PreToolUse hook-guard that nudges toward the graph on search/read tools.
  • .gitattributes: a merge driver for the generated graph.json.
  • The generated graph itself lives in graphify-out/ (git-ignored, rebuilt locally).

After cloning, run this once to install the local git hooks (post-commit and post-checkout rebuild the graph; hooks live in .git/ and are never version-controlled):

graphify hook install

The .claude/settings.json hook-guard invokes graphify from your PATH, so it's portable across machines as long as graphify is installed and on PATH.

License

This project is licensed under the MIT License. See LICENSE for details.

ai-agents
claude
claude-code
cli
deepseek
developer-tools
llm
openai-compatible
productivity
token-optimization

aayushpokhrel1/delegation-pipeline

85.6% fewer orchestrator tokens, measured: offload diff-verifiable grunt work from your paid Claude Code session to free or cheap models. Zero-dependency Python, OpenAI-compatible, ships as a Claude Code plugin.

See the code

See what people are saying

SourceMessageScoreDate

Show HN: Delegation Pipeline – offload Claude Code grunt work to cheap models (86% fewer tokens) (r/SideProject)

Claude Code usage goes fast, and a lot of mine was going into work that didn't need a frontier model: docstrings, mechanical refactors, filling in a module from a contract I'd already specified. I thought of this project as leveraging what free resources in AI we have and having opus be the mind…

1

Oct 2, 2026

README

Delegation Pipeline

GitHub stars License: MIT Python 3.8+ Zero dependencies Claude Code plugin PRs welcome Orchestrator tokens saved

Offload token-heavy grunt work from your paid Claude Code session to free or cheap models, so your Claude usage goes to the thinking, not the typing.

delegate is a zero-dependency Python worker (stdlib only) that runs a small agentic loop over any OpenAI-compatible endpoint. It can read, search, and edit files in the current repo. It cannot run shell commands or use git, by design: you (or your Claude orchestrator) review the resulting git diff and commit.

~/.claude/bin/delegate [backend] "<task>"

The backend is optional and defaults to free ($0). When free cannot be reached, the run escalates once to deepseek automatically, so you do not have to guess a tier per task. See Backends.

demo

Measured impact

Generated 2026-10-02 from 51 delegated runs across 4 repos.

MetricValue
Runs delegated51
Tokens offloaded to workers27,147,998 (measured)
Worker API calls998
Paid on the worker tier$7.60
Claude tokens avoided~23,238,686 (estimated)
Opus-equivalent value~$348.58

By backend:

deepseek  ████████████████████  100%   27,147,998 tokens

By month:

2026-09  ████████████████████   20,873,682 tokens
2026-10  ██████░░░░░░░░░░░░░░   6,274,316 tokens

Offloaded tokens and worker calls are measured from each backend's own usage fields. "Claude tokens avoided" is an estimate: 85.6% of the offloaded total, the median of 9 benchmark tasks in bench/RESULTS.md whose individual savings ranged from 60.7% to 89.7%. Reviewing a worker's diff costs about 14.4% of doing the task inline. Dollar figures are rough blended per-tier prices, for scale, not billing.

That ratio was measured with deepseek-chat on both sides of the benchmark: the same model did each task inline and reviewed the worker's diff. It is applied above to worker tokens as a stand-in for what the orchestrator model would have spent on the same work, which is an assumption the benchmark does not test. Read the direction as sound and the absolute figures as indicative.

Every figure above can be recomputed from the redacted ledger committed at bench/ledger-snapshot.jsonl, using this exact command:

python delegate.py --stats --ledger bench/ledger-snapshot.jsonl

Proof: it measurably cuts orchestrator tokens

Not just a claim. bench/ runs a reproducible A/B where every task is done two ways and measured from real OpenAI-compatible usage fields, and each result is gated by a pytest that fails until the work is actually correct, so a row only counts if the change really worked.

savings% = 1 - T_review / T_do: when it delegates, the orchestrator pays only T_review (write a brief + skim the diff) instead of T_do (doing the whole task itself). Both figures below were measured with deepseek-chat on both sides, so they are deepseek reviewing versus deepseek doing, not Opus versus Opus. Prices scale across models; token counts do not, so read the direction as sound and the absolute number as indicative. See bench/README.md, under "What it measures".

Median: 85.6% fewer orchestrator tokens per task across 9 tasks, spread 60.7-89.7%, with 71,116 tokens of real coding work pushed onto the cheap tier. Every task passed both arms (inline ok / worker ok).

The per-task table lives in bench/RESULTS.md and is not copied here on purpose. It used to be, and the copy silently drifted: it still showed a 4-task run with a "74-93% band" long after the benchmark widened to 9 tasks spanning 60.7-89.7%. RESULTS.md is generated by bench/run_bench.py, and a test asserts it agrees with REVIEW_RATIO in delegate.py, so that one is checked and a hand-maintained second copy never could be.

A real datapoint from building this benchmark: bench/run_bench.py itself (~250 lines, stdlib only) was written by a deepseek worker that consumed 109,058 tokens; the Opus orchestrator paid only the brief plus one diff review, the same asymmetry, on the real orchestrator, on a non-toy file.

Reproduce it in under a minute:

pip install pytest
python bench/run_bench.py --worker deepseek --baseline deepseek   # or --worker free

Every run also prints a per-task TOKENS: total=… trailer (work that ran off your subscription) and rewrites bench/RESULTS.md. Full methodology and the honest caveats live in bench/README.md.

Backends

BackendProviderCostNotes
freeOmniRoute (local)$0Local gateway, auto/coding model
nvidiaNVIDIA API Catalogfree creditsbuild.nvidia.com; strong models, free tier
openrouterOpenRouter APIfree / cheapPin a specific :free or vision model; :free ids are rate-limited
deepseekDeepSeek APIcheapReliable; good default for real work
kimiKimi / Moonshot APIpremiumStrongest; use when quality matters

All of them are just an OpenAI-compatible base URL + model + key, configured in ~/.claude/delegate.config.json.

Default tier and escalation

Naming no backend gets you free. If free fails before the worker has edited anything, the run escalates once to deepseek and says so on stderr:

delegate: free failed with no edits made, escalating to deepseek: ...

Why it works this way, measured on bench/tasks.json on 2026-10-01: reviewing a free worker's diff costs the same as reviewing a paid one (85.4% vs 85.6% median savings), so free is not the lower-quality tier, only the less reliable one. It completed 5 of 9 tasks run back-to-back, while the 4 that failed all passed when retried alone, because the OmniRoute pool drains under sustained load. Defaulting to free and escalating on failure gets the $0 attempt without the caller having to predict which tasks it will finish.

Two deliberate limits:

  • Escalation only happens on a zero-edit failure. Once the worker has written to the tree, the next backend would start from a half-applied change nobody reviewed, so the run stops and tells you to look at the tree instead.
  • --model disables escalation. A model id belongs to one backend, so carrying it to another would 404. Pinning a model means you chose the tier too.

Both attempts are written to the ledger: the failed one with "ok": false and "escalated_to", the successful one with "escalated_from". A tier showing 0 runs and 0 failed attempts was never routed to at all, which the ledger previously could not distinguish from a tier that was tried and died.

Install

The /delegate + /orchestrate commands and the delegate / orchestrate skills ship as a Claude Code plugin, which makes them available in every project (and the skills model-invokable, so "orchestrate it" works anywhere):

claude plugin marketplace add aayushpokhrel1/delegation-pipeline
claude plugin install delegation-pipeline

(Or the interactive /plugin marketplace add ... and /plugin install ....)

That is the whole install. The plugin carries the commands, the skills, and the worker itself: a SessionStart hook runs scripts/ensure_launcher.py, which writes the ~/.claude/bin/delegate launcher pointing at the plugin's own copy of delegate.py. It appears at the start of your next Claude Code session, with no clone and no installer.

The backends still need their own setup, which is the one manual step: the free tier wants a local OmniRoute gateway running, and the remote tiers want an API key. See Keys and Autostart.

Clone-based install

The clone-based install is optional and aimed at people who want to work on this repo:

bash install.sh
./install.ps1

This writes a delegate launcher into ~/.claude/bin/ pointing at your checkout rather than the plugin's copy, so your edits to delegate.py take effect immediately. It also seeds ~/.claude/delegate.config.json from the example and installs the post-commit plugin-refresh hook described below. Re-run after moving the repo. It does not copy the slash commands, those ship with the plugin above.

The session-start hook will not undo this: it leaves any launcher alone whose target still exists, and only writes one when none exists or when the path it points at has gone away.

If you develop this repo, point the marketplace at your checkout

A GitHub marketplace pins a commit, so the installed skills can silently lag your working copy by weeks, and stale skill prose is live instruction in every session. Use a directory marketplace instead:

claude plugin marketplace remove delegation-pipeline
claude plugin marketplace add /path/to/your/checkout
claude plugin install delegation-pipeline

Removing a marketplace uninstalls its plugins, so the install is required, not optional.

claude plugin update does not refresh a directory marketplace. It compares the version in .claude-plugin/plugin.json against the installed version and skips the copy when they match, reporting updateOutcome: "up_to_date" however much the files changed. The version string is the cache key: installs land in ~/.claude/plugins/cache/<marketplace>/<plugin>/<version>/. So either bump the version, or sidestep version bookkeeping:

claude plugin uninstall delegation-pipeline && claude plugin install delegation-pipeline

install.sh / install.ps1 wire scripts/post-commit-refresh-plugin.sh into .git/hooks/post-commit so that second form runs for you whenever a commit touches skills/, commands/ or .claude-plugin/. A written reminder would rot the same way the pinned clone did, and these files are instructions: stale ones do not break, they steer. The hook no-ops unless this checkout is the marketplace source, so it is harmless on a machine installing from GitHub. DELEGATE_SKIP_PLUGIN_REFRESH=1 disables it.

Install copies your working tree, not HEAD: uncommitted edits and even gitignored paths (__pycache__/, graphify-out/) land in the installed copy. Convenient while iterating, but it means a half-finished SKILL.md becomes live guidance, so finish an edit before reinstalling.

One piece, not two. The plugin ships the commands, the skills, and the worker, and sets up its own launcher. A clone only changes which copy of delegate.py that launcher points at, which is what you want while developing this repo and nothing you need otherwise.

Keys

  • free: needs a local OmniRoute gateway and, on current OmniRoute (3.8.50+), an API key. Set up autostart (below) so the gateway is always there, then create a key at http://localhost:20128/dashboard/api-manager and put it in delegate.config.json:

    "free": { "base_url": "http://localhost:20128/v1", "api_key": "...", "model": "auto/coding" }
    

    The key is shown once at creation (ALLOW_API_KEY_REVEAL=false), so copy it there and then. Older docs (and OmniRoute's own startup banner, plus REQUIRE_API_KEY=false in its .env) claim the local gateway is keyless. It is not: /v1/* always requires a Bearer key, and Settings > Security states that authorization model with no toggle to turn it off. Without a key delegate sends no Authorization header at all and every call fails 401 invalid_api_key, which looks exactly like "OmniRoute is down".

  • deepseek / kimi: set env vars, or paste the key into delegate.config.json (that file is git-ignored):

    export DEEPSEEK_API_KEY=sk-...
    export MOONSHOT_API_KEY=sk-...
    
  • nvidia: get a free key at https://build.nvidia.com (any model page → "Get API Key", it starts with nvapi-), then:

    export NVIDIA_API_KEY=nvapi-...
    

    The nvidia backend is OpenAI-compatible via https://integrate.api.nvidia.com/v1. It defaults to nvidia/nemotron-3-super-120b-a12b (verified tool-calling); override per run with --model. The catalog changes fast (models EOL or are account-gated, and the previous deepseek-ai/deepseek-v4-flash-0731 default now hangs), so use --list-models and see MODELS.md. This is a great free path when OmniRoute's routes are dry.

    You can also plug NVIDIA into OmniRoute itself (so the free/auto router can use it): open http://localhost:20128 → provider keys → add the NVIDIA key. Either way works; the direct nvidia backend is the more predictable of the two.

  • openrouter: get a key at https://openrouter.ai/keys, then:

    export OPENROUTER_API_KEY=sk-or-...
    

    OpenAI-compatible via https://openrouter.ai/api/v1. Defaults to nvidia/nemotron-3.5-lightning:free; free models carry a :free suffix and are rate-limited (~50 requests/day at $0, ~1000/day with ~$10 of credits). The worker needs tool calling and the catalog churns, so --list-models first and see MODELS.md. OmniRoute's free backend already routes through OpenRouter, so reach for this direct backend when you want to pin a specific fast/free or vision model instead of auto-routing.

Autostart

So you never have to start the gateway by hand. Logs land in ~/.claude/omniroute.log, and every installer prefers a global omniroute (npm i -g omniroute) and falls back to npx --yes omniroute.

Install OmniRoute globally, don't rely on the npx fallback. npx re-extracts the package and discards Next's .next build cache on every launch, which maximises the cold-start cost below. On npm 11+ the global install silently skips lifecycle scripts, leaving native deps (koffi, keytar, onnxruntime-node, esbuild) unbuilt; if you see an install-scripts warning, re-run with the --allow-scripts=<list> that npm prints.

The starter warms the gateway's routes before exiting. OmniRoute ships as a Next.js dev server, so every route compiles on its first request after a restart: cold on Windows, /api/monitoring/health takes over 90s and /v1/models about 10s, both under 0.1s once warm. Whichever client calls first otherwise pays that bill, times out, and reports the gateway "down" while it is really just compiling. OmniRoute's own CLI hits this too: its readiness probe hardcodes 60s, so npx omniroute often prints "Server did not respond within 60s" about a server that comes up fine seconds later. Treat that warning as noise, and check curl http://127.0.0.1:20128/api/monitoring/health instead.

Because the warm-up waits out that compile, a cold start-omniroute.ps1 can take a few minutes before it exits. That is fine for a hidden logon task where nobody is waiting, and it is the whole point: the task eats the latency so your sessions never do.

The Windows task runs the starter from this repo's path, so moving or renaming the checkout breaks autostart. Re-run install-autostart.ps1 afterwards.

Windows (Scheduled Task, launches hidden at logon):

powershell -ExecutionPolicy Bypass -File scripts\install-autostart.ps1
Start-ScheduledTask -TaskName "OmniRoute Gateway"                 # start now, no reboot
powershell -ExecutionPolicy Bypass -File scripts\uninstall-autostart.ps1   # remove

macOS (launchd agent, RunAtLoad + KeepAlive so it restarts if it dies):

bash scripts/install-autostart.sh          # starts immediately and at every login
bash scripts/uninstall-autostart.sh        # remove

Linux (systemd --user service, restarts on failure):

bash scripts/install-autostart.sh          # enable + start now, and at every login
sudo loginctl enable-linger "$USER"        # optional: keep running without an active login
bash scripts/uninstall-autostart.sh        # remove

All installers are idempotent, they won't start a second copy if one is already running. For a one-off manual start on macOS/Linux without installing autostart, use bash scripts/start-omniroute.sh.

OmniRoute ships with only a couple of keyless free providers (OpenCode Zen, Felo). Their shared quota is frequently exhausted (403 insufficient_quota / 429), which makes the bare free backend unreliable. Add your own free-tier API keys so the router has healthy routes to fall back to. All of these give a free key at signup:

ProviderFree tierGet a key
GroqFast, generous free tierhttps://console.groq.com/keys
Cerebras~1M tokens/day freehttps://cloud.cerebras.ai
Google GeminiFree tier (Flash models)https://aistudio.google.com/apikey
MistralFree experiment tierhttps://console.mistral.ai
OpenRouter:free modelshttps://openrouter.ai/keys
GitHub ModelsFree (rate-limited)GitHub settings → developer token

Add them in the OmniRoute dashboard (open http://localhost:20128 → provider/keys page, keys are encrypted at rest), or via CLI (omniroute keys). Groq + Cerebras + Gemini alone make free solidly usable for grunt work. When free is dry, deepseek remains the reliable paid-but-cheap fallback.

The worker also sends a normal User-Agent (some Cloudflare-fronted free providers return 1010 browser_signature_banned to the default Python-urllib signature) and retries transient provider errors (401/403/429/5xx). Because the free pool is load-balanced, a retry re-rolls to a different provider, so one bad route no longer fails the whole run.

Use it

# smoke test (needs OmniRoute running)
~/.claude/bin/delegate free "Summarize what this project does in 3 bullets"

# real grunt work
~/.claude/bin/delegate deepseek "Add type hints to every function in src/parser.py"
~/.claude/bin/delegate free --model auto/cheap "Write a docstring for each function in utils/"
~/.claude/bin/delegate nvidia "Convert callbacks to async/await in api/client.py"

# vision: attach an image (needs a vision-capable model)
~/.claude/bin/delegate openrouter --model inclusionai/ling-3.0-flash-vl:free \
  --image mockup.png "Build the login form in src/Login.jsx to match this mockup"

Step logs stream to stderr; the final summary prints to stdout.

Flags: --dir <path> (repo root, default cwd), --model <id> (override), --image <path|url> (attach an image, repeatable, needs a vision model; the worker can also fetch images mid-task via its view_image tool), --max-steps N, --list-models (print the backend's catalog and exit), --stats (print the running savings tally and exit), --ledger <path> (read the ledger from somewhere else, e.g. the published snapshot), plus --verify / --commit (below).

The usage ledger

Every delegate run appends one line to ~/.claude/delegate-usage.jsonl (override the path with the DELEGATE_LEDGER env var). The ledger accumulates across all sessions and all repos, so it is a running record of everything you have offloaded, not just the current run.

~/.claude/bin/delegate --stats

--stats reads the ledger back and prints the tally: total offloaded tokens and worker calls, what you paid on the worker tier, the Claude tokens avoided, an Opus-equivalent value, verify failures and commits, plus breakdowns by backend, model, repo, and month.

The offloaded-token count is measured from the backend's own usage fields. The "claude tokens avoided" figure is an estimate: it applies the median review ratio from bench/RESULTS.md (reviewing a worker's diff costs about 13.9% of doing the task inline, so roughly 86.1% of the offloaded tokens never reach your subscription). The dollar figures are rough blended per-tier prices, useful for a sense of scale, not billing.

The file is plain JSONL, one run per line, so it can be grepped, piped into jq, or deleted freely. Deleting it just resets the tally; nothing else depends on it.

bench/ledger-snapshot.jsonl is a committed, redacted copy of that ledger: repo names are stripped, every token count is intact. It is refreshed by the same delegate --stats --readme run that regenerates the block below, so the two cannot drift. That means anyone can recompute the published table from the repo alone:

python delegate.py --stats --ledger bench/ledger-snapshot.jsonl

Redaction costs exactly one thing: with repo names stripped, a recompute from the snapshot reports a single repo instead of the real count. Every token figure matches to the digit.

delegate --stats --readme regenerates the "Measured impact" block at the top of this README from the ledger. It never includes repository names, only a count of how many repos are involved, so the block is safe to publish. A weekly scheduled task can keep it current: scripts/install-stats-task.ps1 registers it on Windows.

Letting Claude pick the model

The tool never auto-selects a model beyond each backend's default, that's the orchestrator's job. When you say "delegate" (or run /delegate, or Opus drives the raw CLI on its own), Claude sizes up the task and chooses both the backend and the model/route, using --list-models to see the live catalog and MODELS.md as the shortlist. Two flavors of choice:

  • free (OmniRoute) → Claude picks an auto/* route (auto/coding, auto/cheap, auto/fast, ...) and lets OmniRoute's router pick the concrete model, so it fails over across your provider keys. Prefer routes over pinning one model here.
  • nvidia / paid backends → Claude picks a concrete model id (there's no router in front), restricted to tool-calling-capable models since the worker edits via function calls.

You can always force a specific backend/model/route by naming it: delegate free --model auto/cheap "..." or delegate nvidia --model nvidia/nemotron-3-super-120b-a12b "...".

From inside Claude Code: /delegate

The installer also drops a /delegate slash command into ~/.claude/commands/, so you can trigger a delegation without leaving your Claude session:

/delegate deepseek Add type hints to every function in src/parser.py
/delegate free Write a docstring for each exported function in utils/

The command has Claude expand your request into a tight, self-contained spec, run the worker, then review the resulting git diff and report back. First token is the backend (free if omitted); the rest is the task.

A free worker can take minutes, since the free model is the slow part. Run it with a long foreground timeout so it isn't detached to the background mid-run. If your shell or tool does background it at a default timeout, that only cuts off the wait: the worker keeps running and its file writes still land. Just confirm the run finished before treating the diff as complete, reading files while a backgrounded worker is still going can show a half-applied tree. deepseek/nvidia are faster if you want to avoid the wait entirely.

How Claude should drive it

The orchestrator protocol lives in your global ~/.claude/CLAUDE.md. In short:

  1. Write a tight, self-contained instruction (name the files, describe the change, point at a pattern to mirror). The worker has no conversation context.
  2. Delegate on a clean tree so the resulting git diff is attributable.
  3. Review the diff. Fix small issues yourself; re-delegate with a sharper spec if it drifted. Workers never commit.
  4. You run tests and commit, after review.

Route by complexity, the gate is a tight verifiable spec, not low difficulty: free for trivial/mechanical work (boilerplate, repetitive edits, docstrings, stubs); deepseek (roughly Sonnet-class) for substantial well-specified work (a whole module, a real refactor, a non-trivial first draft), do not cap delegation at boilerplate. Keep design, tricky debugging, security-sensitive code, deep-context work, and one-liners in the Claude session.

Proactive delegation (no /delegate needed)

The orchestrator (Opus) is set up to delegate on its own whenever a request contains work it can spec and verify, mechanical or substantial, without you typing /delegate. It announces the backend/model in one line, runs the worker, reviews the diff, and folds the result in. /delegate and saying "delegate this" still work as manual triggers; they're just not required. Say "don't delegate this" to keep a task (or a sensitive repo) in-session.

Note: delegating sends repo content to external model providers (deepseek/nvidia are remote; free/OmniRoute routes out too). That's the intended trade; disable it per-repo when the code is sensitive.

This behavior lives in your global ~/.claude/CLAUDE.md, not in this repo, so it travels with that file, not with a clone. To enable it on another device, add a block like this to that device's ~/.claude/CLAUDE.md:

# Delegation to cheap-model workers
Use `~/.claude/bin/delegate <backend> [--model <id>] "<task>"` to offload work.
Delegate proactively (no /delegate needed): announce the backend/model in one line, run
the worker, then review the git diff. Route by COMPLEXITY (the gate is a tight verifiable
spec, not low difficulty): `free` for trivial/mechanical work; `deepseek` (roughly
Sonnet-class) for substantial well-specified work like a whole module or a real refactor,
do not cap it at boilerplate; `nvidia` when free is dry (tool-calling models only, see
MODELS.md). Keep design, tricky debugging, security-sensitive code, deep-context work, and
one-liners in-session. Stop if told "don't delegate this".

(The full version is in this repo's git history / the author's own CLAUDE.md.)

Combined Orchestration

The delegation pipeline and Claude subagents are two ways to move work off the paid orchestrator. Combined orchestration routes each task to the cheapest tier that can do it, so Claude tokens go to planning and judgment, not typing.

The model

One orchestrator, two worker tiers.

  • Orchestrator = Claude Opus in a Claude Code session. Holds the plan, writes a tight per-task brief, routes each task, and adjudicates reviews. It is the only component that spends Claude tokens, and it never writes implementation code itself.
  • Tier B = this delegation pipeline (free / deepseek / nvidia). Zero Claude tokens. This is the default tier for both implementation and review.
  • Tier A = Claude subagents (the Agent tool: haiku / sonnet / opus). Costs Claude tokens. An escape hatch, not a default.

Routing table

Route each task on three axes. Any single axis landing in the right column sends the task to Tier A; otherwise it stays on Tier B.

AxisTier B (default, no Claude tokens)Tier A (escalate, costs tokens)
Complexityanything specifiable -> delegate (defaults to free, escalates to deepseek on failure); force the paid tier with delegate deepseek when a task must not be retriedneeds broad codebase judgment -> subagent sonnet / opus
Iteration depthverifiable in ~one shot (orchestrator runs verify once, commits)long autonomous run-fail-edit loop, or needs a live service / Docker / the app running
Sensitivityordinary codesecurity-sensitive, or full-context debugging (Tier A, or the orchestrator itself)

Reviews default to Tier B too: a deepseek worker reviews the diff against the brief for zero Claude tokens.

Cost gate. Delegation is not free of the orchestrator's tokens: the brief, the review, and any re-run all cost Claude tokens. Before routing a task to Tier B, the orchestrator weighs that overhead against doing it inline, and keeps the task in-session when the brief plus review would cost as much as the edit itself. The pipeline only saves tokens when the work is bulkier than its description, which is why one-liners stay in-session.

The loop (per task)

  1. Brief. The orchestrator writes a tight, self-contained brief: exact files, the change, a pattern to mirror, exact values. The worker has no conversation context.
  2. Route and run. Pick a backend by the table and hand off. The worker edits; the CLI verifies and commits on green:
    ~/.claude/bin/delegate deepseek \
      --verify "npm test" \
      --commit "feat: <what changed>, per brief" \
      "<the full brief>"
    
  3. On red. A failed verify exits 2 and commits nothing, printing the command output. The orchestrator reads that output and either re-delegates with the failure folded in as a sharper spec, or escalates the task to a Tier A subagent.
  4. Review. A second worker reviews the committed diff against the brief, for zero Claude tokens:
    git show HEAD > review.diff
    ~/.claude/bin/delegate deepseek \
      "Review the diff in review.diff against the brief in brief.md. Report spec compliance and code quality, most severe first. Do not edit anything."
    
    The orchestrator adjudicates: accept, fix inline, or re-delegate.
  5. Escalate. Go to Tier A only when the loop cannot converge, or the task hits one of the axes above.

The verify/commit flow

--verify "<cmd>" runs the command after the worker finishes editing. A non-zero exit prints the output and exits 2 without committing. --commit "<msg>" stages and commits only the files the worker edited, but only when verify passed (or no --verify was given). Together they make a delegated task self-contained: the orchestrator hands off a brief and gets back a tested, committed result, without babysitting the test-and-commit cycle and without spending Claude tokens on it.

# mechanical, free tier, verified and committed in one shot
~/.claude/bin/delegate free \
  --verify "python -m pytest tests/test_utils.py -q" \
  --commit "test: cover utils edge cases" \
  "Add the three missing edge-case tests to tests/test_utils.py, mirroring the table-driven style already there."

# check the outcome
echo $?   # 0 = verified and committed; 2 = verify failed, nothing committed

On Windows, --verify runs through cmd.exe, so pass one command (npm test, npx tsc --noEmit), not a bash && / ; chain.

The trust boundary holds. The worker model still never runs shell or git. Only the caller's --verify and --commit flags run commands, and only the orchestrator sets them. See Tools the worker has.

The /orchestrate command and the orchestrate skill encode this whole table and loop, so an orchestrator invokes one thing instead of re-deriving it each session.

Tools the worker has

list_dir, find_files, read_file, search_text, write_file, edit_file. No shell, no network beyond the model endpoint, no git. File access is sandboxed to the working directory. The --verify / --commit flags are the one exception, and they are run by the CLI harness (the caller), never by the worker model.

Portability

Everything the worker needs is one delegate.py file and the stdlib, so it runs on any device with Python 3.8+. Sync = clone this repo + run the installer. Real keys live in ~/.claude/delegate.config.json (git-ignored), never in the repo.

Same spirit, different mechanism. The graphify skill builds knowledge graphs; its semantic extraction (docs, papers, images) otherwise dispatches Claude subagents, which spends your Claude tokens on the initial build. Set a Gemini key once and graphify routes that work to Gemini instead:

[Environment]::SetEnvironmentVariable("GEMINI_API_KEY", "<your-key>", "User")
[Environment]::SetEnvironmentVariable("GOOGLE_API_KEY", "<your-key>", "User")
export GEMINI_API_KEY="<your-key>"    # or GOOGLE_API_KEY; graphify accepts either

Code-only corpuses skip semantic extraction entirely and never cost subagent tokens. Already-running Claude Code sessions must be fully restarted to pick up a newly set key. Details live in the graphify skill's SKILL.md ("Save Claude tokens: set a Gemini key"). Get a key at https://aistudio.google.com/apikey.

This repo is graphify-integrated

This repo ships a graphify knowledge-graph integration so Claude consults the graph before falling back to raw search, and keeps it current automatically:

  • CLAUDE.md: guidance telling Claude to run graphify query/path/explain before answering codebase questions, and graphify update . after code changes.
  • .claude/settings.json: a PreToolUse hook-guard that nudges toward the graph on search/read tools.
  • .gitattributes: a merge driver for the generated graph.json.
  • The generated graph itself lives in graphify-out/ (git-ignored, rebuilt locally).

After cloning, run this once to install the local git hooks (post-commit and post-checkout rebuild the graph; hooks live in .git/ and are never version-controlled):

graphify hook install

The .claude/settings.json hook-guard invokes graphify from your PATH, so it's portable across machines as long as graphify is installed and on PATH.

License

This project is licensed under the MIT License. See LICENSE for details.

ai-agents
claude
claude-code
cli
deepseek
developer-tools
llm
openai-compatible
productivity
token-optimization

Languages

Python

82.7%

PowerShell

9.9%

Shell

7.4%