Runs CLI coding agents (Claude Code, Codex, Copilot, Cursor, Gemini, opencode) against a task queue. Each works in an isolated VM; output is checked by configurable auditors, then merged via git. Pools multiple provider subscriptions with quota-aware routing and fallback. C#/.NET 10, MIT-licensed.
C#
7
1,238 commits
updated Oct 5, 2026
An autonomous coding orchestrator you can hand work to and walk away from. Give it a task — a title and a prompt against one of your repos — and CodeyBox picks a coding agent, runs it inside a throwaway VM, and then does the part that makes walking away possible: it puts the change in front of a panel of auditors, sends every failing finding back to the agent as rework, and only lands the change on your branch (and on GitHub, if you point it there) once every auditor on the panel passes. You stay in the loop for product decisions; it handles the delivery grind.
The audit panel is why you don't have to watch it. An agent that says it is done is not trusted to be done. Every item goes through a default panel of sixteen auditors, each an independent hard gate — no averaging into "good enough", no single reviewer to talk round:
On top of that panel sit 64 auditor plugins you can switch on for the stack you actually have — linters for a dozen languages, SAST, dependency vulnerabilities, secrets, infrastructure-as-code, licensing, schema and API compatibility, documentation — and language presets for C#, Python, Node, Go and Rust. Plugins are off until you enable them, so audit time tracks what you chose to check. See Quality gates you control.
It drives a fleet of twenty-five agent CLIs — Claude Code, OpenAI Codex, GitHub Copilot, Cursor, Devin, Gemini, opencode, Aider, Goose and more — and routes each task to whichever one is best and available, falling back automatically when a provider hits a rate limit. No coding agent ever runs on your host: every model call that touches a repository happens through an agent CLI inside a sandbox, boxed in a real VM with its own kernel (and, on Linux, behind a host-enforced firewall) — see Security: defense in depth.
Built in C#/.NET 10. Managed repos can be any stack — Python, Node, Go, Rust, C#, or your own — through config-driven auditors.

CodeyBox ships its own web admin on the host. It is two views of one fleet.
Map view is the default, and the picture above is what it is for. You file
a feature as a chain of small items with explicit dependencies; the map draws
the graph, and the orchestrator derives the execution order from it. Every
edge names the item it waits on by title rather than by id, and the (+)
beside a card files a new item that queues behind it.
Everything is placed against a time axis: landed work to the left, what is running now in the middle, the queue forecast to the right. Zoom is semantic — chains at a distance, cards close up, an item's full stage pipeline when you zoom into it. Idle time is compressed rather than scrolled through, so months of history stay on one screen, and positions are anchored to the work rather than to the clock, so nothing drifts under the cursor while you read it.
![]() The whole fleet — chains with their counts, the fan-out of everything waiting on one item, and the predicted dispatch batches to the right of now. Four hundred chains and five hundred landed items on one canvas. | ![]() Queue view — the same fleet as a list when you want one: filter by state, reorder dispatch, and act on a row without leaving it. |
Every item goes plan → work → audit → merge → landed, and audit is a gate,
not a step: its fail path returns the item to work. That loop is drawn rather
than described, and an item's record keeps the whole history — how many times
it worked, how many times audit sent it back, every finding and which auditor
raised it.
![]() What it took to land — seven work attempts, five audits, one rejection that sent it back, and twenty-three findings across the auditor panel, each named, timed and quoted against the file it came from. | ![]() The loop, live — a running item on its third attempt with a passed audit, and the returns that got it there labelled on the arcs: operator retries and interruptions, stated in plain language underneath. |
What needs a decision is pinned to the side of the canvas rather than waiting to be found, with the actions that actually apply — answer the agent's question, retry from work, retry from audit, delegate a repair turn. Suggestions raised by agents while they work appear as ghost cards beside the item that produced them, and promote to real work items in one click.
![]() Landed, failed, running — three states of one chain side by side, with the failure tethered to its entry in the rail and offering the three things you can do about it. | ![]() Suggestions — an agent noticed the work-item prompts point at a directory that does not exist, and proposed the fix. Promote it and it becomes a work item; dismiss it and it goes away. |
The remaining page screenshots in screenshots/ predate this
rework and still show the old sidebar shell; they are generated against a
deterministic seeded instance by tools/screenshots/ and
are being regenerated.
CodeyBox is an orchestrator as well as an application. It exposes a REST API, a SignalR stream and a typed CLI, and it is designed to be left running without anyone watching it.
When you want to steer it from another machine or from your phone, use Agnes — a separate product, a remote interface to coding CLIs, which ships a first-class CodeyBox client. Point it at your orchestrator and you get the screens below. Neither product requires the other.
![]() Overview — a plain-language verdict, quota-to-reset per agent, thirty days of cumulative flow, and a "needs a look" list ranked by how stuck something is rather than by age. | ![]() Work queue — now, next in dispatch order, waiting on you, and landed. A failed item explains itself and offers the three things you can actually do about it. |
sudo on your machine. Every agent runs
in a real VM with its own kernel, so a compromised agent can't reach your
host — and on Linux hosts a host-enforced firewall stops it exfiltrating past
its allowlist.
Phases 1–3 are atomic: the change lands cleanly or not at all. A clean merge is
pure git plumbing on the host — git merge-tree then git commit-tree, no VM,
no agent — and only a genuine content conflict is handed to an in-VM agent, then
checked by a deterministic host-side scope fence. Push is a separate retryable
tier, so a flaky remote never corrupts your local result.
The optional Plan phase runs first when a work item sets the plan knob:
the agent drafts a plan artifact that reviewers evaluate before any code is
written, which is worth the extra cycle on larger or higher-risk changes. The
full state machine is in docs/concepts/architecture.md.
Most agent orchestrators run the model in a container or straight on the host. CodeyBox stacks several independent layers between an agent and your machine, so a prompt-injected or actively malicious agent has to defeat all of them:
sudo in its
sandbox still can't reach your LAN, cloud-metadata endpoints, or anything off
its allowlist — it can't flush a firewall it can't see. Providers that can't
enforce this are labelled egress not enforced, and work that requires an
enforced network profile is never placed on them. One concession exists: a
provider-host filter outside the guest (Tart Softnet on a Mac) may serve
profiled work per sandbox, only after the host's own canary passes — and it
is never ranked above host enforcement.Honest caveat: this is defense in depth, not a guarantee. A determined
adversary — especially one targeting a weaker coding agent you've installed —
may still find a path, and a misconfigured egress profile or an over-broad
project setup weakens the model. Sandbox-escape and egress-bypass testing on a
live KVM host is still outstanding. Read
docs/concepts/security.md before you trust it with
anything that matters.
On a Linux host, the fastest path — it checks prerequisites, installs what is missing, offers to set up host network isolation, builds, and writes a starter config:
curl -fsSL https://raw.githubusercontent.com/AdamFrisby/CodeyBox/main/install.sh | bash
It is idempotent, prompts before anything with side effects, and refuses to continue silently if host network isolation could not be set up. It does steps 1 to 3 for you and prints where it put the config, so when it finishes go straight to step 4.
Because the script arrives on stdin, flags need bash -s --:
curl -fsSL https://raw.githubusercontent.com/AdamFrisby/CodeyBox/main/install.sh | bash -s -- --yes
curl -fsSL https://raw.githubusercontent.com/AdamFrisby/CodeyBox/main/install.sh | bash -s -- --help
Follow all four steps to set up by hand. Use them on macOS and Windows too, where the installer does not run and only the remote-executor topology is supported.
1. Install prerequisites — the .NET 10 SDK, Git, a sandbox provider, and at least one authenticated agent CLI.
2. Clone and build. Use ./build.sh on Linux and macOS — it heals an
unwritable NuGet home first (see below). On Windows use ./build.ps1, which
forwards to dotnet with the same telemetry settings.
git clone https://github.com/AdamFrisby/CodeyBox.git
cd CodeyBox
./build.sh # Windows: ./build.ps1
If restore fails with
Failed to read NuGet.Config due to unauthorized access: This applies toinstall.shtoo, since it builds the same way. NuGet probes user-level configuration under$HOME/.nuget/NuGet/regardless of what the repository pins, so it needs that directory to be writable. A home baked read-only, or owned by another user, aborts restore for every project — and a checked-in config or--configfiledoes not help, because NuGet probes the user settings directory anyway.
./build.shhandles this for you: it sourcesscripts/nuget-home-heal.sh, which is the single source of truth for the repair and is shared with the audit path. It relocates an unwritable tree aside (no root needed), preserves the populated package cache by symlink so restore stays offline-safe, and seeds a readable user config. If$HOMEitself cannot be written to — an inherited read-only mount, say — then even moving the tree aside is impossible, so it instead redirectsDOTNET_CLI_HOMEto a writable scratch directory for that process tree../build.sh # builds, healing the NuGet home first if needed . scripts/nuget-home-heal.sh # or just heal the current shell
3. Configure a project. Drop a JSON file somewhere and point
CODEYBOX_EXTRA_CONFIG at it (it hot-reloads on change):
{
"CodeyBox": {
"SandboxProvider": "multipass",
"Projects": [
{
"Id": "my-app",
"RepositoryUrl": "https://github.com/you/my-app.git",
"BaseBranch": "main",
"Agent": "claude"
}
]
}
}
4. Run:
dotnet run --project tools/CodeyBox.Cli -- queue add \
--project my-app \
--title "Add a hello file" \
--prompt "Add hello.txt containing the word hello."
dotnet run --project tools/CodeyBox.Cli -- queue watch WORK_ITEM_ID
The step-by-step version — host networking, a minimal config, the first work
item, and what to check when it fails — is in
docs/getting-started.md.
CodeyBox trades wall-clock time and tokens for review depth. Throughput is bounded by host CPU and agent quota, because each concurrent phase runs a VM. Small, dependent tasks generally converge faster than monolithic prompts.
Tune concurrency, agent classes, auditors, iteration limits, and budgets for
your workload. Watch state transitions and updated timestamps — not only
completed-item count — to tell a quota-limited queue apart from a stuck one.
Recovery procedures are in
docs/operating/running.md and
docs/operating/recovery.md.
docs/concepts/agent-classes.mddocs/operating/host-firewall.mddocs/operating/quota.mddocs/operating/recovery.mddocs/reference/api.md,
docs/reference/webhooks.mddocs/operating/remote-executors.mddocs/extending/plugins.mdAuditors stack. You choose exactly which checks gate a merge — built-in tool auditors (formatting, build, the full test suite, coverage, mutation rigor, gitleaks secret scanning, semgrep SAST) and LLM reviewers over six audit types (security, architecture, quality, completeness, cheating, tests) — plus any of the plugin catalogue, or your own. Each runs in its own capability-scoped sandbox, and the tool-only ones hold no agent credentials.
The plugin catalogue (each one disabled until you enable it):
| Category | Auditors |
|---|---|
| Linting (22) | Biome, clang-tidy, Clippy, Cppcheck, Credo, detekt, ESLint, golangci-lint, ReSharper InspectCode, Knip, mypy, Oxlint, PHPStan, PMD, Pyright, Roslynator, RuboCop, Ruff, SpotBugs, Staticcheck, SwiftLint, and a SARIF example to build your own |
| SAST (5) | Bandit, Brakeman, CodeQL, DevSkim, Semgrep |
| Dependency vulnerabilities (8) | cargo-audit, cargo-deny, OWASP Dependency-Check, govulncheck, Grype, OSV-Scanner, Socket, Trivy |
| Secrets (4) | Betterleaks, detect-secrets, Gitleaks, TruffleHog (with live credential verification) |
| Infrastructure (10) | actionlint, cfn-lint, Checkov, Conftest, Hadolint, KICS, KubeLinter, kubeconform, TFLint, zizmor |
| Schema (3) | Spectral, SQLFluff, Squawk |
| API compatibility (3) | Buf breaking, cargo-semver-checks, GraphQL Inspector |
| Architecture (3) | dependency-cruiser, Import Linter, file-size limits |
| Documentation (3) | lychee, markdownlint, Vale |
| Scripting (2) | PSScriptAnalyzer, ShellCheck |
| Licensing (2) | REUSE, ScanCode Toolkit |
Enabling one adds its tool to the sandbox baseline; disabling it takes it back
out. → docs/extending/auditor-plugins.md
Test-heavy suites can opt into regression test selection: after every
merge CodeyBox records which lines each test covers, and audits run only the
tests a change can reach. It ships shadow-first — the full suite still runs and
the would-be selection is scored — and only switches to enforcing once a
calibration window shows it never skips a test that would have failed.
→ docs/quality/test-selection.md
The gate is hard: when any auditor fails, its findings go straight back to the
agent, which reworks and resubmits — the loop repeats until every gate passes
or it hits the iteration cap, at which point the item is flagged AuditFailed
and is not merged. The auditor set, the failing-severity threshold, and the
iteration cap are all per-project config.
→ docs/quality/audit.md
CodeyBox tracks token usage and estimated spend for every work item, broken down by phase (work, each rework, each audit iteration, merge) and by agent/model. So you can answer "what did this bugfix actually cost to run?" — and build a real feel for the economics of automated work before you scale it up.
Costs are normalised to pay-per-API list prices — even on subscription plans, and accounting for cached tokens — so they're comparable across agents and over time. Query per item or per project:
curl -H "authorization: Bearer $CODEYBOX_API_KEY" \
http://localhost:5036/workitems/<id>/costs # one item, broken out by phase
curl -H "authorization: Bearer $CODEYBOX_API_KEY" \
http://localhost:5036/projects/my-app/costs # the whole project
The admin dashboard's Costs tab charts the same data.
→ docs/operating/costs.md
codeybox is a typed client for the whole API — no more curl + jq. Run it from
source (dotnet run --project tools/CodeyBox.Cli -- <command>) or publish a
self-contained binary:
dotnet publish tools/CodeyBox.Cli -c Release -r linux-x64 -o ./bin/codeybox
codeybox configure # save API URL + token to ~/.config/codeybox
Everyday use:
# Queue a task (inline, --prompt-file, or piped in) and follow it live
ID=$(codeybox queue add --project my-app --title "Add /healthz" \
--prompt "Add a /healthz endpoint returning 200." --quiet)
codeybox queue watch "$ID" # streams state transitions over SSE
codeybox queue ls --state Working,Auditing # what's in flight
codeybox queue show <id> # full detail for one item
codeybox queue retry <id> --from audit # re-drive a failed item
codeybox queue cancel <id>
queue add also takes --agent, --work-branch, --base-branch,
--auditor-profile, --push-upstream, and --depends-on (to chain dependent
items); --json / --quiet make every command pipe-friendly.
→ docs/reference/cli.md
Twenty-five agent CLIs are supported today:
claude · codex · copilot · cursor · devin · gemini · opencode ·
antigravity · crock · aider · goose · pi · prime · autohand ·
vibe · cline · kilo · omp · continue · qwen · cmd · crush ·
caveman · dotnet-opencode · unreal
Each lives in src/CodeyBox.Agents.<Name> and implements IAgentRunner — a
subclass of CliAgentRunnerBase that builds one non-interactive invocation.
Adding another is that class, a credential mapping, an install line in the
sandbox baseline, and two smoke probes.
Agents are interchangeable. A class lists members with quality scores; the router
prefers the highest-scoring one that's within quota and under its concurrency
cap. See
docs/concepts/agents.md for each agent's auth, its
sandbox install command, and its known quirks.
| Orchestrator host | incus | multipass (local) | tart (plugin) | multipass-remote | sprites | bubblewrap | process (dev-only) |
|---|---|---|---|---|---|---|---|
| Linux | ✅ VM, egress enforced on host | ✅ VM, egress enforced on host | ❌ macOS only | ✅ VM, egress enforced on executor | ✅ VM, egress enforced on executor | ⚠️ shared kernel, no egress | ⚠️ no isolation, dev only |
| macOS | ❌ | ❌ | ✅ VM (macOS or Linux guests), egress verified per sandbox via Softnet (see below) | ✅ VM, egress enforced on executor | ✅ VM, egress enforced on executor | ❌ | ❌ |
| Windows | ❌ | ❌ | ❌ | ✅ VM, egress enforced on executor | ✅ VM, egress enforced on executor | ❌ | ❌ |
On a Mac, run the orchestrator locally (./build.sh) and give agents
local VMs with the Tart plugin — a
fresh VM with its own kernel per work item, with macOS guests as well as Linux
ones, so Apple-platform work can run too. A Tart VM is VM isolation (its own
kernel — the primary boundary), but its egress is NotEnforced by default:
guest network follows the Mac's. Opt into Softnet mode plus host-owned
per-sandbox canary verification (CodeyBox:EgressVerification with tart
opted in) and each sandbox is handed over only after its own canary passes —
the verified grant (EnforcedOnProviderHostVerified) is deliberately never
stronger than host enforcement. The fail-closed and IPv6 properties are
established only by the Mac-only operator procedure
(scripts/verify-tart-softnet.sh); until it has run on real hardware the path
is documented as unverified. Use Tart for the
work you would trust with that, and a Linux host or a remote Linux executor for
the rest — work that requires an enforced network profile is placed there
automatically.
On Windows, run the orchestrator locally (./build.ps1) with VMs on a
remote Linux executor host, where the allowlist holds.
An unenforced allowlist is never described as isolation, and unsupported
provider + host combinations fail fast at startup with a message pointing at
the matrix (docs/concepts/host-platforms.md).
Pick with CodeyBox.SandboxProvider:
| Provider | Setup | Isolation |
|---|---|---|
incus | Incus 6.3+ and existing ZFS/Btrfs pool | KVM; fast, space-efficient copy-on-write baseline clones |
multipass | snap install multipass | KVM; simplest setup |
multipass-remote | Multipass on a remote host + SSH | KVM, VMs offloaded to another machine over SSH — orchestrator stays local |
sprites | a Fly.io Sprites account | Firecracker microVMs over an HTTP/WebSocket API; writable host mounts sync back at teardown, not per exec |
bubblewrap | apt install bubblewrap | namespaces, shared kernel; integration-tested |
process | none | none — testing only, never with untrusted prompts |
More backends ship as plugins (disabled by default): cloud VMs on any
OpenStack cloud (openstack, with a sample config for Infomaniak Public Cloud),
hosted sandboxes on Daytona, E2B, Modal, Runloop and Blaxel,
local microVMs with BoxLite and microsandbox, and macOS guests with
Tart. Their egress is classified not enforced — the host can't put its
firewall in front of a machine it doesn't own — so placement keeps any work that
requires an enforced network profile on a host-enforced provider, and each
plugin's doc says exactly what isolation it does and doesn't give. The one
exception is Tart in Softnet mode with host-owned canary verification (above):
a per-sandbox verified grant, never above host enforcement.
→ docs/extending/sandbox-plugins.md
Choose explicitly: prefer incus for persistent or high-throughput headless
installations, and multipass for the simplest setup. Multipass baseline clones
copy full VM images; Incus ZFS/Btrfs clones are copy-on-write, reducing launch
time, disk use, and repeated SSD writes. multipass-remote runs the same VMs on
a separate host over SSH while the orchestrator — state, git, merge, auditors —
stays local, so you can offload VM CPU without splitting the brain.
A graphical flavour (a desktop plus VNC/X display, and a computer-use bridge
exposing screenshots and input synthesis through the sandbox API) is available on
both Incus and Multipass. Turn it on per project with
"GraphicalSandbox": true, not by selecting a provider.
→ docs/concepts/sandboxes.md
docs/concepts/sandboxes.md, including Incus
storage-pool and service-identity prerequisites.scripts/setup-host-networks.sh
creates a Linux bridge per network profile and writes nftables rules that drop
anything not on the profile's allowlist. A compromised agent with sudo can't
disable this, because it lives on the host, not in the guest.
→ docs/operating/host-firewall.mddocs/concepts/security.md — the threat
model, the trust boundaries, the sharp edges, and the known gaps. This is not
optional.Credentials are tiered: tool-only audit sandboxes hold no agent secrets, and upstream remote credentials (e.g. a GitHub PAT) live only in the orchestrator process and never cross into a sandbox.
docs/ is the full reference, indexed by task. Good entry
points:
getting-started.md — clean host to merged changeconcepts/architecture.md — the system, its boundaries, the state machineconcepts/security.md — threat model (read before deploying)concepts/projects.md — project, auditor, and upstream configconcepts/agent-classes.md — routing, quotas, and fallbackextending/plugins.md — the plugin SDKreference/api.md — the full REST referenceCodeyBox is under active development and builds clean against .NET 10. Incus is
recommended for persistent, high-throughput headless deployments; Multipass is
the simpler option. The process provider is for constrained testing only and
gives no isolation. Issues and contributions are welcome.
Because CodeyBox builds itself, its roadmap is its own work queue — and most of what's described above was built that way, by agents working through this same audit panel. Recently landed: the plugin catalogue (64 auditors plus forges, work sources, notifications, credential and sandbox backends), the majordomo, remote executors, and the coverage-baseline producer for test selection. The threads currently moving: calibrating test selection toward enforcement, verifying a merge's combined result builds before it lands, counting audit sessions against per-agent concurrency caps, autonomous exploratory testing that emits replayable regression artifacts, and smarter quota drain scheduling.
100 followers · starred Sep 2026
Runs CLI coding agents (Claude Code, Codex, Copilot, Cursor, Gemini, opencode) against a task queue. Each works in an isolated VM; output is checked by configurable auditors, then merged via git. Pools multiple provider subscriptions with quota-aware routing and fallback. C#/.NET 10, MIT-licensed.
C#
7
1,238 commits
updated Oct 5, 2026
An autonomous coding orchestrator you can hand work to and walk away from. Give it a task — a title and a prompt against one of your repos — and CodeyBox picks a coding agent, runs it inside a throwaway VM, and then does the part that makes walking away possible: it puts the change in front of a panel of auditors, sends every failing finding back to the agent as rework, and only lands the change on your branch (and on GitHub, if you point it there) once every auditor on the panel passes. You stay in the loop for product decisions; it handles the delivery grind.
The audit panel is why you don't have to watch it. An agent that says it is done is not trusted to be done. Every item goes through a default panel of sixteen auditors, each an independent hard gate — no averaging into "good enough", no single reviewer to talk round:
On top of that panel sit 64 auditor plugins you can switch on for the stack you actually have — linters for a dozen languages, SAST, dependency vulnerabilities, secrets, infrastructure-as-code, licensing, schema and API compatibility, documentation — and language presets for C#, Python, Node, Go and Rust. Plugins are off until you enable them, so audit time tracks what you chose to check. See Quality gates you control.
It drives a fleet of twenty-five agent CLIs — Claude Code, OpenAI Codex, GitHub Copilot, Cursor, Devin, Gemini, opencode, Aider, Goose and more — and routes each task to whichever one is best and available, falling back automatically when a provider hits a rate limit. No coding agent ever runs on your host: every model call that touches a repository happens through an agent CLI inside a sandbox, boxed in a real VM with its own kernel (and, on Linux, behind a host-enforced firewall) — see Security: defense in depth.
Built in C#/.NET 10. Managed repos can be any stack — Python, Node, Go, Rust, C#, or your own — through config-driven auditors.

CodeyBox ships its own web admin on the host. It is two views of one fleet.
Map view is the default, and the picture above is what it is for. You file
a feature as a chain of small items with explicit dependencies; the map draws
the graph, and the orchestrator derives the execution order from it. Every
edge names the item it waits on by title rather than by id, and the (+)
beside a card files a new item that queues behind it.
Everything is placed against a time axis: landed work to the left, what is running now in the middle, the queue forecast to the right. Zoom is semantic — chains at a distance, cards close up, an item's full stage pipeline when you zoom into it. Idle time is compressed rather than scrolled through, so months of history stay on one screen, and positions are anchored to the work rather than to the clock, so nothing drifts under the cursor while you read it.
![]() The whole fleet — chains with their counts, the fan-out of everything waiting on one item, and the predicted dispatch batches to the right of now. Four hundred chains and five hundred landed items on one canvas. | ![]() Queue view — the same fleet as a list when you want one: filter by state, reorder dispatch, and act on a row without leaving it. |
Every item goes plan → work → audit → merge → landed, and audit is a gate,
not a step: its fail path returns the item to work. That loop is drawn rather
than described, and an item's record keeps the whole history — how many times
it worked, how many times audit sent it back, every finding and which auditor
raised it.
![]() What it took to land — seven work attempts, five audits, one rejection that sent it back, and twenty-three findings across the auditor panel, each named, timed and quoted against the file it came from. | ![]() The loop, live — a running item on its third attempt with a passed audit, and the returns that got it there labelled on the arcs: operator retries and interruptions, stated in plain language underneath. |
What needs a decision is pinned to the side of the canvas rather than waiting to be found, with the actions that actually apply — answer the agent's question, retry from work, retry from audit, delegate a repair turn. Suggestions raised by agents while they work appear as ghost cards beside the item that produced them, and promote to real work items in one click.
![]() Landed, failed, running — three states of one chain side by side, with the failure tethered to its entry in the rail and offering the three things you can do about it. | ![]() Suggestions — an agent noticed the work-item prompts point at a directory that does not exist, and proposed the fix. Promote it and it becomes a work item; dismiss it and it goes away. |
The remaining page screenshots in screenshots/ predate this
rework and still show the old sidebar shell; they are generated against a
deterministic seeded instance by tools/screenshots/ and
are being regenerated.
CodeyBox is an orchestrator as well as an application. It exposes a REST API, a SignalR stream and a typed CLI, and it is designed to be left running without anyone watching it.
When you want to steer it from another machine or from your phone, use Agnes — a separate product, a remote interface to coding CLIs, which ships a first-class CodeyBox client. Point it at your orchestrator and you get the screens below. Neither product requires the other.
![]() Overview — a plain-language verdict, quota-to-reset per agent, thirty days of cumulative flow, and a "needs a look" list ranked by how stuck something is rather than by age. | ![]() Work queue — now, next in dispatch order, waiting on you, and landed. A failed item explains itself and offers the three things you can actually do about it. |
sudo on your machine. Every agent runs
in a real VM with its own kernel, so a compromised agent can't reach your
host — and on Linux hosts a host-enforced firewall stops it exfiltrating past
its allowlist.
Phases 1–3 are atomic: the change lands cleanly or not at all. A clean merge is
pure git plumbing on the host — git merge-tree then git commit-tree, no VM,
no agent — and only a genuine content conflict is handed to an in-VM agent, then
checked by a deterministic host-side scope fence. Push is a separate retryable
tier, so a flaky remote never corrupts your local result.
The optional Plan phase runs first when a work item sets the plan knob:
the agent drafts a plan artifact that reviewers evaluate before any code is
written, which is worth the extra cycle on larger or higher-risk changes. The
full state machine is in docs/concepts/architecture.md.
Most agent orchestrators run the model in a container or straight on the host. CodeyBox stacks several independent layers between an agent and your machine, so a prompt-injected or actively malicious agent has to defeat all of them:
sudo in its
sandbox still can't reach your LAN, cloud-metadata endpoints, or anything off
its allowlist — it can't flush a firewall it can't see. Providers that can't
enforce this are labelled egress not enforced, and work that requires an
enforced network profile is never placed on them. One concession exists: a
provider-host filter outside the guest (Tart Softnet on a Mac) may serve
profiled work per sandbox, only after the host's own canary passes — and it
is never ranked above host enforcement.Honest caveat: this is defense in depth, not a guarantee. A determined
adversary — especially one targeting a weaker coding agent you've installed —
may still find a path, and a misconfigured egress profile or an over-broad
project setup weakens the model. Sandbox-escape and egress-bypass testing on a
live KVM host is still outstanding. Read
docs/concepts/security.md before you trust it with
anything that matters.
On a Linux host, the fastest path — it checks prerequisites, installs what is missing, offers to set up host network isolation, builds, and writes a starter config:
curl -fsSL https://raw.githubusercontent.com/AdamFrisby/CodeyBox/main/install.sh | bash
It is idempotent, prompts before anything with side effects, and refuses to continue silently if host network isolation could not be set up. It does steps 1 to 3 for you and prints where it put the config, so when it finishes go straight to step 4.
Because the script arrives on stdin, flags need bash -s --:
curl -fsSL https://raw.githubusercontent.com/AdamFrisby/CodeyBox/main/install.sh | bash -s -- --yes
curl -fsSL https://raw.githubusercontent.com/AdamFrisby/CodeyBox/main/install.sh | bash -s -- --help
Follow all four steps to set up by hand. Use them on macOS and Windows too, where the installer does not run and only the remote-executor topology is supported.
1. Install prerequisites — the .NET 10 SDK, Git, a sandbox provider, and at least one authenticated agent CLI.
2. Clone and build. Use ./build.sh on Linux and macOS — it heals an
unwritable NuGet home first (see below). On Windows use ./build.ps1, which
forwards to dotnet with the same telemetry settings.
git clone https://github.com/AdamFrisby/CodeyBox.git
cd CodeyBox
./build.sh # Windows: ./build.ps1
If restore fails with
Failed to read NuGet.Config due to unauthorized access: This applies toinstall.shtoo, since it builds the same way. NuGet probes user-level configuration under$HOME/.nuget/NuGet/regardless of what the repository pins, so it needs that directory to be writable. A home baked read-only, or owned by another user, aborts restore for every project — and a checked-in config or--configfiledoes not help, because NuGet probes the user settings directory anyway.
./build.shhandles this for you: it sourcesscripts/nuget-home-heal.sh, which is the single source of truth for the repair and is shared with the audit path. It relocates an unwritable tree aside (no root needed), preserves the populated package cache by symlink so restore stays offline-safe, and seeds a readable user config. If$HOMEitself cannot be written to — an inherited read-only mount, say — then even moving the tree aside is impossible, so it instead redirectsDOTNET_CLI_HOMEto a writable scratch directory for that process tree../build.sh # builds, healing the NuGet home first if needed . scripts/nuget-home-heal.sh # or just heal the current shell
3. Configure a project. Drop a JSON file somewhere and point
CODEYBOX_EXTRA_CONFIG at it (it hot-reloads on change):
{
"CodeyBox": {
"SandboxProvider": "multipass",
"Projects": [
{
"Id": "my-app",
"RepositoryUrl": "https://github.com/you/my-app.git",
"BaseBranch": "main",
"Agent": "claude"
}
]
}
}
4. Run:
dotnet run --project tools/CodeyBox.Cli -- queue add \
--project my-app \
--title "Add a hello file" \
--prompt "Add hello.txt containing the word hello."
dotnet run --project tools/CodeyBox.Cli -- queue watch WORK_ITEM_ID
The step-by-step version — host networking, a minimal config, the first work
item, and what to check when it fails — is in
docs/getting-started.md.
CodeyBox trades wall-clock time and tokens for review depth. Throughput is bounded by host CPU and agent quota, because each concurrent phase runs a VM. Small, dependent tasks generally converge faster than monolithic prompts.
Tune concurrency, agent classes, auditors, iteration limits, and budgets for
your workload. Watch state transitions and updated timestamps — not only
completed-item count — to tell a quota-limited queue apart from a stuck one.
Recovery procedures are in
docs/operating/running.md and
docs/operating/recovery.md.
docs/concepts/agent-classes.mddocs/operating/host-firewall.mddocs/operating/quota.mddocs/operating/recovery.mddocs/reference/api.md,
docs/reference/webhooks.mddocs/operating/remote-executors.mddocs/extending/plugins.mdAuditors stack. You choose exactly which checks gate a merge — built-in tool auditors (formatting, build, the full test suite, coverage, mutation rigor, gitleaks secret scanning, semgrep SAST) and LLM reviewers over six audit types (security, architecture, quality, completeness, cheating, tests) — plus any of the plugin catalogue, or your own. Each runs in its own capability-scoped sandbox, and the tool-only ones hold no agent credentials.
The plugin catalogue (each one disabled until you enable it):
| Category | Auditors |
|---|---|
| Linting (22) | Biome, clang-tidy, Clippy, Cppcheck, Credo, detekt, ESLint, golangci-lint, ReSharper InspectCode, Knip, mypy, Oxlint, PHPStan, PMD, Pyright, Roslynator, RuboCop, Ruff, SpotBugs, Staticcheck, SwiftLint, and a SARIF example to build your own |
| SAST (5) | Bandit, Brakeman, CodeQL, DevSkim, Semgrep |
| Dependency vulnerabilities (8) | cargo-audit, cargo-deny, OWASP Dependency-Check, govulncheck, Grype, OSV-Scanner, Socket, Trivy |
| Secrets (4) | Betterleaks, detect-secrets, Gitleaks, TruffleHog (with live credential verification) |
| Infrastructure (10) | actionlint, cfn-lint, Checkov, Conftest, Hadolint, KICS, KubeLinter, kubeconform, TFLint, zizmor |
| Schema (3) | Spectral, SQLFluff, Squawk |
| API compatibility (3) | Buf breaking, cargo-semver-checks, GraphQL Inspector |
| Architecture (3) | dependency-cruiser, Import Linter, file-size limits |
| Documentation (3) | lychee, markdownlint, Vale |
| Scripting (2) | PSScriptAnalyzer, ShellCheck |
| Licensing (2) | REUSE, ScanCode Toolkit |
Enabling one adds its tool to the sandbox baseline; disabling it takes it back
out. → docs/extending/auditor-plugins.md
Test-heavy suites can opt into regression test selection: after every
merge CodeyBox records which lines each test covers, and audits run only the
tests a change can reach. It ships shadow-first — the full suite still runs and
the would-be selection is scored — and only switches to enforcing once a
calibration window shows it never skips a test that would have failed.
→ docs/quality/test-selection.md
The gate is hard: when any auditor fails, its findings go straight back to the
agent, which reworks and resubmits — the loop repeats until every gate passes
or it hits the iteration cap, at which point the item is flagged AuditFailed
and is not merged. The auditor set, the failing-severity threshold, and the
iteration cap are all per-project config.
→ docs/quality/audit.md
CodeyBox tracks token usage and estimated spend for every work item, broken down by phase (work, each rework, each audit iteration, merge) and by agent/model. So you can answer "what did this bugfix actually cost to run?" — and build a real feel for the economics of automated work before you scale it up.
Costs are normalised to pay-per-API list prices — even on subscription plans, and accounting for cached tokens — so they're comparable across agents and over time. Query per item or per project:
curl -H "authorization: Bearer $CODEYBOX_API_KEY" \
http://localhost:5036/workitems/<id>/costs # one item, broken out by phase
curl -H "authorization: Bearer $CODEYBOX_API_KEY" \
http://localhost:5036/projects/my-app/costs # the whole project
The admin dashboard's Costs tab charts the same data.
→ docs/operating/costs.md
codeybox is a typed client for the whole API — no more curl + jq. Run it from
source (dotnet run --project tools/CodeyBox.Cli -- <command>) or publish a
self-contained binary:
dotnet publish tools/CodeyBox.Cli -c Release -r linux-x64 -o ./bin/codeybox
codeybox configure # save API URL + token to ~/.config/codeybox
Everyday use:
# Queue a task (inline, --prompt-file, or piped in) and follow it live
ID=$(codeybox queue add --project my-app --title "Add /healthz" \
--prompt "Add a /healthz endpoint returning 200." --quiet)
codeybox queue watch "$ID" # streams state transitions over SSE
codeybox queue ls --state Working,Auditing # what's in flight
codeybox queue show <id> # full detail for one item
codeybox queue retry <id> --from audit # re-drive a failed item
codeybox queue cancel <id>
queue add also takes --agent, --work-branch, --base-branch,
--auditor-profile, --push-upstream, and --depends-on (to chain dependent
items); --json / --quiet make every command pipe-friendly.
→ docs/reference/cli.md
Twenty-five agent CLIs are supported today:
claude · codex · copilot · cursor · devin · gemini · opencode ·
antigravity · crock · aider · goose · pi · prime · autohand ·
vibe · cline · kilo · omp · continue · qwen · cmd · crush ·
caveman · dotnet-opencode · unreal
Each lives in src/CodeyBox.Agents.<Name> and implements IAgentRunner — a
subclass of CliAgentRunnerBase that builds one non-interactive invocation.
Adding another is that class, a credential mapping, an install line in the
sandbox baseline, and two smoke probes.
Agents are interchangeable. A class lists members with quality scores; the router
prefers the highest-scoring one that's within quota and under its concurrency
cap. See
docs/concepts/agents.md for each agent's auth, its
sandbox install command, and its known quirks.
| Orchestrator host | incus | multipass (local) | tart (plugin) | multipass-remote | sprites | bubblewrap | process (dev-only) |
|---|---|---|---|---|---|---|---|
| Linux | ✅ VM, egress enforced on host | ✅ VM, egress enforced on host | ❌ macOS only | ✅ VM, egress enforced on executor | ✅ VM, egress enforced on executor | ⚠️ shared kernel, no egress | ⚠️ no isolation, dev only |
| macOS | ❌ | ❌ | ✅ VM (macOS or Linux guests), egress verified per sandbox via Softnet (see below) | ✅ VM, egress enforced on executor | ✅ VM, egress enforced on executor | ❌ | ❌ |
| Windows | ❌ | ❌ | ❌ | ✅ VM, egress enforced on executor | ✅ VM, egress enforced on executor | ❌ | ❌ |
On a Mac, run the orchestrator locally (./build.sh) and give agents
local VMs with the Tart plugin — a
fresh VM with its own kernel per work item, with macOS guests as well as Linux
ones, so Apple-platform work can run too. A Tart VM is VM isolation (its own
kernel — the primary boundary), but its egress is NotEnforced by default:
guest network follows the Mac's. Opt into Softnet mode plus host-owned
per-sandbox canary verification (CodeyBox:EgressVerification with tart
opted in) and each sandbox is handed over only after its own canary passes —
the verified grant (EnforcedOnProviderHostVerified) is deliberately never
stronger than host enforcement. The fail-closed and IPv6 properties are
established only by the Mac-only operator procedure
(scripts/verify-tart-softnet.sh); until it has run on real hardware the path
is documented as unverified. Use Tart for the
work you would trust with that, and a Linux host or a remote Linux executor for
the rest — work that requires an enforced network profile is placed there
automatically.
On Windows, run the orchestrator locally (./build.ps1) with VMs on a
remote Linux executor host, where the allowlist holds.
An unenforced allowlist is never described as isolation, and unsupported
provider + host combinations fail fast at startup with a message pointing at
the matrix (docs/concepts/host-platforms.md).
Pick with CodeyBox.SandboxProvider:
| Provider | Setup | Isolation |
|---|---|---|
incus | Incus 6.3+ and existing ZFS/Btrfs pool | KVM; fast, space-efficient copy-on-write baseline clones |
multipass | snap install multipass | KVM; simplest setup |
multipass-remote | Multipass on a remote host + SSH | KVM, VMs offloaded to another machine over SSH — orchestrator stays local |
sprites | a Fly.io Sprites account | Firecracker microVMs over an HTTP/WebSocket API; writable host mounts sync back at teardown, not per exec |
bubblewrap | apt install bubblewrap | namespaces, shared kernel; integration-tested |
process | none | none — testing only, never with untrusted prompts |
More backends ship as plugins (disabled by default): cloud VMs on any
OpenStack cloud (openstack, with a sample config for Infomaniak Public Cloud),
hosted sandboxes on Daytona, E2B, Modal, Runloop and Blaxel,
local microVMs with BoxLite and microsandbox, and macOS guests with
Tart. Their egress is classified not enforced — the host can't put its
firewall in front of a machine it doesn't own — so placement keeps any work that
requires an enforced network profile on a host-enforced provider, and each
plugin's doc says exactly what isolation it does and doesn't give. The one
exception is Tart in Softnet mode with host-owned canary verification (above):
a per-sandbox verified grant, never above host enforcement.
→ docs/extending/sandbox-plugins.md
Choose explicitly: prefer incus for persistent or high-throughput headless
installations, and multipass for the simplest setup. Multipass baseline clones
copy full VM images; Incus ZFS/Btrfs clones are copy-on-write, reducing launch
time, disk use, and repeated SSD writes. multipass-remote runs the same VMs on
a separate host over SSH while the orchestrator — state, git, merge, auditors —
stays local, so you can offload VM CPU without splitting the brain.
A graphical flavour (a desktop plus VNC/X display, and a computer-use bridge
exposing screenshots and input synthesis through the sandbox API) is available on
both Incus and Multipass. Turn it on per project with
"GraphicalSandbox": true, not by selecting a provider.
→ docs/concepts/sandboxes.md
docs/concepts/sandboxes.md, including Incus
storage-pool and service-identity prerequisites.scripts/setup-host-networks.sh
creates a Linux bridge per network profile and writes nftables rules that drop
anything not on the profile's allowlist. A compromised agent with sudo can't
disable this, because it lives on the host, not in the guest.
→ docs/operating/host-firewall.mddocs/concepts/security.md — the threat
model, the trust boundaries, the sharp edges, and the known gaps. This is not
optional.Credentials are tiered: tool-only audit sandboxes hold no agent secrets, and upstream remote credentials (e.g. a GitHub PAT) live only in the orchestrator process and never cross into a sandbox.
docs/ is the full reference, indexed by task. Good entry
points:
getting-started.md — clean host to merged changeconcepts/architecture.md — the system, its boundaries, the state machineconcepts/security.md — threat model (read before deploying)concepts/projects.md — project, auditor, and upstream configconcepts/agent-classes.md — routing, quotas, and fallbackextending/plugins.md — the plugin SDKreference/api.md — the full REST referenceCodeyBox is under active development and builds clean against .NET 10. Incus is
recommended for persistent, high-throughput headless deployments; Multipass is
the simpler option. The process provider is for constrained testing only and
gives no isolation. Issues and contributions are welcome.
Because CodeyBox builds itself, its roadmap is its own work queue — and most of what's described above was built that way, by agents working through this same audit panel. Recently landed: the plugin catalogue (64 auditors plus forges, work sources, notifications, credential and sandbox backends), the majordomo, remote executors, and the coverage-baseline producer for test selection. The threads currently moving: calibrating test selection toward enforcement, verifying a merge's combined result builds before it lands, counting audit sessions against per-agent concurrency caps, autonomous exploratory testing that emits replayable regression artifacts, and smarter quota drain scheduling.
100 followers · starred Sep 2026