Open-source coding agent CLI. Switch models in one session and finish explicit goals against checks you approve.
TypeScript
146
17,531 commits
updated Oct 4, 2026
Open Multi-Agent Kit
Make “done” pass a check.
A terminal coding agent with model switching and explicit, test-backed goals.
Quick start ·
Default or opt-in? ·
Evidence and limits
Choosing OMK ·
Documentation
An AI saying “done” is not an acceptance check. With /goal verify, you choose
the command that checks your goal. OMK runs it after each settled turn and
completes the goal only when it passes on the current workspace. A later edit
makes that evidence stale. The result is only as strong as the check you choose.
/goal Fix the failing test without changing its assertions.
/goal verify node --test check.test.mjs
Use a test that exists in your project, or follow the small, reproducible goal demo. This is an explicit workflow; ordinary prompts do not enable the gate.
Choose a model, work on your repository, then switch models with /model when
another one suits the next step. The conversation stays in the same session.
You can also stop and return later with /resume or omk -c.
OMK is a standalone CLI, not a plugin for Claude Code or OpenCode. It supports subscription providers, API keys, and local models. Start with one agent that reads files, edits code, and runs commands. Add subagents or explicit verification workflows when you need them; neither is required for your first task.
Requires Node.js 22.19 or newer. Start in the repository you want to work on:
npm install -g open-multi-agent-kit --ignore-scripts
omk --version
cd your-project
omk
Without a global install, run npx --ignore-scripts open-multi-agent-kit from
that directory.
/login to authenticate a supported subscription or API-key provider./model to choose an available model.Summarize this repository and identify the commands used to check it.
Read the project configuration to support your answer. Do not edit files.
After the reply, use /model to choose another configured model and ask it to
review the answer. You stay in the same session. This is manual model switching,
not parallel agents or an independent correctness check.
Ready to see a test-backed goal? Try the goal demo, then share your result or a reproducible failure in a GitHub issue. If this workflow is useful, star OMK to help other developers find it.
For a bug fix, name the failing behavior and ask for a regression test, the smallest fix, and the check commands with their exit codes. Review the diff and those results yourself; a request to run tests does not enable a verification gate.
Built-in local bash requires sandbox-exec on macOS or bwrap plus
unprivileged user namespaces on Linux. It blocks network access and fails
closed if the backend is missing. See the safety boundary
and full quickstart for setup.
A fresh install starts one agent/tool loop after provider setup. It does not turn each prompt into a multi-agent workflow or automatically certify its answer.
| Capability | Fresh-install behavior | Where to start |
|---|---|---|
| File editing, shell commands, saved sessions | Built in; tools run when called by the agent | Usage, sessions |
| Tool-call scheduling | dag-v2 schedules resource conflicts within the agent loop; it does not launch a team | Runtime algorithms |
| Subagents | Optional extension; load it and supply agent definitions | Subagent setup and examples |
| MCP servers, extra skills and extensions | Require configured servers or installed resources | MCP, skills, extensions |
| Durable goals | Built in; /goal <objective> continues after each settled turn, up to 8 rounds. With /goal verify <command> the goal completes only when that check passes on the current workspace | Acceptance checks |
| Protocol verification and advisory judging | Explicit API/workflow opt-in; not a gate on ordinary prompts | Run protocol |
Verified runs (omk run) | Opt-in CLI and SDK on Linux with bwrap; runs an approved command in an isolated copy, checks the result, and supports resume, cancel and cleanup | Verified run |
| Context budgeting | Off by default | Settings |
| AdaptOrch integration | Optional and separate; no service calls by default | OMK + AdaptOrch |
The internal lane launcher and automatic command-sharding primitives are not connected to the default CLI path. Installing their packages is not the same as enabling an orchestration workflow.
No comparative benchmark result is published here yet. We have not established that OMK solves more tasks than another harness, that multi-agent execution improves success, or how much verification reduces false completion.
OMK targets state-of-the-art quality as a CLI coding-agent harness. SOTA is not verified.
The evidence you can inspect today covers specific failure modes:
| Behavior covered | Regression evidence | Scope |
|---|---|---|
Missing test observations produce inconclusive; a required failing test produces fail | Protocol tests | Explicit protocol evaluation, without a waiver |
| Changed artifacts, wrong command bindings, or missing ledger evidence block acceptance | Evidence binding tests | Strict evidence gate and selected workspace scope |
| A relevant workspace mutation after verification makes the receipt stale | Freshness tests | Configured receipt and mutation tracking |
| A durable goal with an approved acceptance check completes only on a passing receipt, and an edit after the check makes that receipt stale | Goal acceptance tests, live loop tests | Approval held by the running OMK process; git work tree; static command lines |
| An effect that may still be live keeps its resource claims through expiry, cancellation and authority restart, until a supervisor confirms termination | Coordination broker tests | In-process broker with canonical claim keys; no OS fencing |
| A publication is refused unless its read versions, parent revision and receipt binding all still match | Publication tests | Single accepted snapshot pointer; no multi-file filesystem atomicity |
| Cancellation after dispatch is never reported as cancelled-before-dispatch; the outcome stays unknown until settled | Operation lifecycle tests | Pure state machine; does not itself stop a remote effect |
| Zero trials is reported as absent evidence rather than zero risk, and a point estimate is not an error rate | Risk bound tests | Binomial model under a fixed policy and adequately independent samples |
These tests exercise the gates, not the rate at which they catch real bugs.
An ordinary prompt finishes when its tool loop and queued work settle;
prompt_settled is not a correctness verdict.
A useful comparison must hold the model, provider configuration, tasks, budget,
and tool permissions constant, and label default versus opt-in workflows.
Report task success, cost, latency, and false completion (reported complete but
failing the declared checks), with its denominator and per-task outcomes.
The measurement protocol
defines the reproducibility and privacy requirements. omk stats shows local
turn costs and tool failures; it does not score task correctness.
If you evaluate OMK, share a sanitized report and reproduction steps in a GitHub issue. Include failed and interrupted runs, not just successful examples.
The terminal UI shows the selected model, tools, and session status. Additional signals depend on the integrations you configure.
The header reads omk v<package.version> · OMK//CONTROL; the installed package
version is the source of truth.
OMK's provider-neutral coding-agent CLI also exposes a multi-agent control plane for explicitly configured workflows. The diagram describes that design, not what every prompt automatically runs.
The v0.98.3 SDK rejects incomplete first-party judge responses and exposes deterministic ties; it is not an automatic TUI judge.
The animation changes once every 1.5 seconds and contains no flashing. The four steps above are the complete text alternative.
AgentSession built-in local bash uses OS sandbox enforcement by default:
sandbox-exec on macOS and bwrap plus unprivileged user namespaces on Linux.
Local shell spawns restrict writes to the workspace and OS temporary directory,
disable network access, and fail closed with sandbox.backend_missing when an
enforcement backend is unavailable.
This is not read-confidentiality or whole-process containment. Other file
tools, extension and custom-tool code, injected or remote BashOperations, and
the OMK process keep the permissions of the process running them. Use
containerization when the
boundary must cover more than built-in local bash. Explicit evidence workflows
cannot treat missing required evidence as a verified result.
OMK keeps routing separate from control and evidence. Codex, Claude Code,
OpenCode Zen/Go, Kimi, GLM/ZAI, native xAI/Grok, NVIDIA NIM, and local providers
can participate through omk-ai while the run contract stays stable.
Native xai keeps subscription OAuth and XAI_API_KEY billing separate. See
provider setup,
provider resilience, and
Grok integration.
| Package | Purpose |
|---|---|
open-multi-agent-kit | Interactive coding-agent CLI and control plane |
omk-agent-core | Agent runtime, tool execution, and DAG scheduling |
omk-ai | Unified multi-provider LLM API |
omk-protocol | Versioned run contracts and semantic reducers |
omk-adaptorch-wpl | Work Packet Loop runtime |
omk-book-to-skill | Optional document-to-skill compiler |
omk-tui | Differential-rendered terminal UI library |
npm install omk-agent-core
npm install omk-ai
npm install omk-protocol
omk install npm:omk-book-to-skill@0.98.3
npm install omk-tui
v0.97.0 shipped the OpenWiki policy and workflow, but no versioned corpus or
integrity checker. The following integrity/output guards shipped in v0.98.0;
the generated corpus remains optional and is not bundled:
openwiki/ — absent. The previous untracked corpus was removed after the
hardened gate proved it carried fabricated evidence: 8 frontmatter symbols
that no declared source path defines (AgentLoop, getModel, DeepWall,
loadExtensions, createExtensionRuntime, main), 45 references to @omk/*
package names this repository does not publish, and a
restatement of this README's Scope -> Route -> Verify -> Replay loop as a
strict engine state machine, which is not what the source implements.
CI regenerates the corpus; nothing is lost.scripts/check-openwiki.mjs — shipped integrity checker. An interrupted corpus
now fails unless openwiki/.manual-review.json binds a review to the exact
corpus digest, and every frontmatter symbol must bind to one of that page's
own source_paths as a whole identifier.scripts/check-openwiki-output.mjs — output gate. The scheduled workflow
may write, upload, and open a PR for openwiki/ and nothing else, so a model
reading this repository cannot reach AGENTS.md, CLAUDE.md, or the workflow
that runs it. The gate runs once before the artifact leaves the read-only
generating job and again before the PR, because the publishing job holds write
permissions the first one does not..understand-anything/ — optional local structural graph used by Pi Lens;
it is not published or injected into prompts by default. To reach a session,
attach it through OMK's MCP client like
any other server; there is no second, bespoke path for it.Source and tests remain authoritative. Shipped guards do not turn a generated index into authority: treat corpus pages as local advisory data and recheck source.
A corpus no session can read is documentation of a plan, not a feature, so the
pages are now candidates for prompt budgeting. Enable contextBudget.openwiki
alongside contextBudget.enabled
(settings)
and each page becomes a low-priority evidence item ranked against the turn's
query. Pages compete for leftover budget and can never displace instructions or
skills; most turns carry titles and declared symbols alone, and a page's text
arrives only when the query earns it.
Admission mirrors scripts/check-openwiki.mjs rather than restating it. A
complete corpus at the current HEAD offers page text; one whose HEAD has
moved offers titles only and is marked stale; an interrupted corpus is refused
unless a review binds to its exact digest. The default is off, and with the
setting off the prompt is byte-identical to one built without a corpus.
OMK is this local, MIT-licensed coding agent. AdaptOrch is a separate proprietary evidence service. Neither requires the other: installing OMK does not create an AdaptOrch account or make calls to it by default.
For an optional integration, see the WPL package and clients and MCP setup. The WPL package exposes state, client, and adjudication primitives, not an automatic verification loop for every CLI prompt.
AdaptOrch's reports carry correctness_claim=false; they are not semantic
correctness proofs or OMK harness benchmark results.
Review AdaptOrch plans
· Claim boundary
The AdaptOrch name and marks identify that separate proprietary product and appear here with permission. They are excluded from this repository's MIT grant — see LICENSE.
The design decisions behind OMK's context, routing, memory, and orchestration layers are grounded in published work rather than invented in isolation. Each row below was retrieved and read directly; claims are at abstract level, which is the evidence grade this table asserts and no more.
| Paper | Mechanism it establishes | OMK implementation or design reference |
|---|---|---|
| arXiv:2608.22752 — The Compaction Cliff in Long-Running AI Agent Memory | Uniform summarization erodes rules and episodic logs at the same rate; measured safety-rule retention falls to 53% after one compaction and 10% after five. Type-tagged deterministic operators fix it. | Type-aware compaction triage: rule-typed items survive N rounds byte-identical |
| arXiv:2608.23023 — Most of the LLM Routing Gap Is Task Type | Most routing gain is reachable with a fixed task-type table; run-to-run flips must not be credited as wins. | Frozen task-class table plus the 2-run stability rule in the promotion gate |
| arXiv:2506.16655 — Arch-Router: Aligning LLM Routing with Human Preferences | Indirection: a classifier emits a label, a policy table maps label to decision, so models change without retraining. | classifyTaskV4 plus TASK_CLASS_THINKING_LEVELS |
| arXiv:2605.09894 — Deterministic vs. LLM-Controlled Orchestration | Holding model, prompts, and tools constant and varying only execution control, deterministic orchestration matched accuracy, improved worst-case robustness, and cut tokens up to 3.5x. | Deterministic scheduler and planned lanes; execution control is never delegated to the model |
| arXiv:2608.15565 — Admission Without Answers | Label-free admission on execution success alone admits substantial contamination; an accept/abstain/escalate decision is required. | Verified-memory admission design (spec 019), abstain is not stored |
| arXiv:2608.23471 — InjecMEM: Memory Injection Attack on LLM Agent Memory Systems | Single-interaction memory injection is a reproduced attack frame against agent memory. | Retrieved memory is injected only as provenance-tagged data, never fused into instruction position |
Entries include implemented mechanisms and design proposals; check the runtime status guide for availability. The wider survey, including approaches not adopted, is working material that is not published with the repository.
npm ci --ignore-scripts
npm run build
npm run check
npm test
npm run release:local
Direct dependencies are pinned, CI installs with --ignore-scripts, and the
published CLI includes a generated npm-shrinkwrap.json. Read
CONTRIBUTING.md and the
development guide before sending a
change.
Use it for provider choice within one CLI session, or to build workflows against its public runtime and evidence APIs. For a single-provider workflow, your current agent may be sufficient. Try the read-only task above before moving existing work.
OMK is a separate runtime with its own CLI, sessions, tool scheduler, and SDK.
One reason to choose it is to build your own acceptance workflow: define
required test observations in the run protocol,
then have your automation reject fail or inconclusive results. Receipt
integrity and freshness still need their own configured checks.
For adding a tool or prompt to an existing OpenCode setup, a plugin may be the smaller change. OMK's protocol is opt-in, not proof of better performance.
Not automatically. Subagents require setup, and verification must be part of the chosen workflow. Its result covers the declared checks, not all behavior. See what runs by default and evidence and limits.
Historical correction: the immutable v0.97.0 notes below announced a versioned OpenWiki corpus, but that release still ignored
/openwiki/and did not contain the corpus or checker. See the current repository-understanding section above for the shipped guards and optional-corpus boundary.
/goal verify <command> approves an acceptance check that runs in the default bash sandbox after each settled turn. The goal completes only when the check passes on the current workspace, and the check's output never reaches the model. See acceptance checks.omk run pauses instead of failing and resumes with restart-writer, resume or retry-tasks. omk run cancel stops a run from another shell, and omk run gc prunes derived workspaces while keeping the evidence. See cancellation and artifact GC.failed with a classified error, and every other server still contributes its tools. See failure behavior.SIGINT, SIGTERM or the new omk run cancel) no longer ends it failed. The run stays paused with failure: cancelled, and its journal records a new interrupted event. Resume the phase that was cut off with restart-writer, resume or retry-tasks; a DAG attempt whose process was confirmed stopped is released instead of spent. A verification check cut off by the cancellation is no longer signed into the receipt as a failed check. OMK 1.2.4 and earlier cannot read a journal that contains interrupted, and a status consumer that treated cancellation as terminal must handle paused. See cancellation.OMK_SESSION_CONTROL=1 or session.startControl(), then use sdk session ... --live. Exact session IDs and private per-enrollment endpoints are required; live failures never become transcript writes.OMK_VERIFIED_MEMORY=1 plus Context Budget V2 and uses transient tool-result data, not instruction text. No automatic extraction or quality-improvement claim./goal verify <command> approves an acceptance check for the durable goal and runs it through receipt-bound local bash under the default bash sandbox. A passing check is attached as goal evidence. After each settled turn the check runs again: a pass completes the goal, and a failure starts the next round with the command and its exit code, never its output. /goal complete then requires a passing receipt from this session that still matches the workspace; a tracked edit, new file or HEAD move after the check makes it stale. Approvals stay in the OMK process, so approve the check again after a restart. See acceptance checks.nextDurableGoalTimestamp(goal) is exported for SDK callers of applyDurableGoalCommand() and DurableGoalStore.transition(). It returns the wall clock, raised to no earlier than the goal's last update and later than the start of its generation, which is the time the reducer accepts after the clock steps back (seen on WSL2). See durable goal lifecycle.omk run cancel ID [--wait-ms N] cancels a verified run owned by another process. It writes a request file that the owner checks every 250 ms and never signals a PID. watchRunCancelRequest() connects the same request to an SDK caller's AbortSignal, and cancelVerifiedRun() is the SDK form of the command.omk run gc [--older-than DURATION] [--execute] prunes the derived workspaces (writer*, candidate*, tasks) of verified runs that can no longer be recovered, holding each run's owner lease while it checks. It keeps journals, keys, manifests, blobs, receipts and attestations, never follows symlinks, and only reports without --execute. SDK: collectVerifiedRuns().AgentSession checks the complete request against the model window less the output reserve and safety margin (computeHardPromptInputLimit()). estimateContextInputTokens() counts the system prompt, the messages after convertToLlm() and the tool schemas, and takes the largest of the configured tokenizer count, the character heuristic and projected provider usage.session.metacognition exposes a content-free state snapshot and the latest bounded lastDiagnostic, observed at prompt preflight and settlement. It is observation-only: it never rewrites a prompt, authorizes a tool, changes termination or grants completion.McpServerConfig.inheritEnv: false keeps the parent environment out of a stdio server, maxPendingWriteBytes (default 16 MiB) bounds bytes queued on the server's stdin, and await manager.closeAndWait() joins physical transport close. manager.status() marks a server whose close is pending with retiring: true. inheritEnv is not read from mcp.json yet.planSkills() prunes its exact search and adds dominance and eviction-refill passes to its greedy search.sel-4-codeunit, so entries cached under the earlier policy are not reused; the public optimizer identifier is unchanged.omk run reports an error outside the verified-run contract as verified-run: operation_failed (<kind> <code>), for example (Error ENOTDIR), instead of a bare operation_failed. The message itself stays out of the output because it can carry absolute paths or contract text.ps, so Korean and other non-English parent locales no longer make a live process appear unavailable. The acceptance-check regression fixtures use canonical physical temporary repository paths on macOS, preserving workspace-mismatch rejection across /var and /private/var aliases.brace-expansion 5.0.12 and undici 8.10.2, plus undici 6.28.1 for the optional Gondolin example. That example still depends on node-forge 1.4.0, whose RSA signature-verification advisory has no patched npm release as of 2026-10-04; the repository production audit continues to report it.agent_end and sent the next turn while that run still owned the session; the session rejected it with Agent is already processing, so the round was spent and the goal never continued. The controller now acts once a turn settles, when no automatic retry follows it, and queues the next round as a follow-up. An attempt that is about to be retried no longer uses up a round.goal timestamps must be monotonic or goal generation timestamp must advance when the wall clock steps back, as WSL2 does when its hypervisor resyncs time. The controller dates each transition no earlier than the journal's last timestamp.operation_failed when the wall clock steps back under load. The authority store's default clock is the wall time at open plus monotonic elapsed time; a clock injected by the caller that moves backward is still refused.writer_incomplete, and omk run status suggests recovery only for runs that are running or paused.failed with a classified public error while every other server still contributes its tools. A malformed tool result rejects with mcp.invalid_tool_result instead of counting as success. Request and handshake timeouts accept only safe integers from 0 through 2,147,483,647 ms, and 0 refuses to send. manager.status() omits free-form server-reported versions.ownership.dispatch_active. Numeric environment values with trailing characters are rejected instead of parsed as a prefix, and prompt-size estimation projects each tool's name, description and parameters instead of serializing the tool object.RpcClient keeps only the last 8,192 characters of the current child's stderr and clears them on start(). waitForIdle() and collectEvents() reject when the child fails, exits or is stopped, and one waiter unsubscribing no longer makes another miss agent_end. stop() rejects pending requests at once, returns without the one-second delay when the child already exited, and rejects with RpcTerminationUncertainError while keeping ownership when termination cannot be confirmed. prompt() propagates a server rejection, so promptAndWait() no longer waits for its 60-second timeout.SessionManager.getBranch() no longer shifts the path array for every ancestor (2,096,128 element moves at depth 2,048, now none). The run journal builds its frozen record copy only when records is read, and the memory-only journal store no longer replays every earlier record on each append (8,256 hashes for 128 appends, now 128). An AgentSession listener that unsubscribes while an event is dispatched no longer makes the next listener miss it.omk-ai and omk-agent-core fixes apply: complete() and completeSimple() no longer queue every stream event until they return, a Cursor request that reaches its deadline is closed instead of left running, provider retries reject invalid options and stop early on an aborted signal, and taking queued messages one at a time no longer copies the rest of the queue.attemptId, process settlement and stream receipts instead of the first attempt's, and the README install list includes managed-process-tree.ts, subagent-stream.ts and graph-result.ts, without which the extension did not load.Context limit reached until a manual /compact. Threshold compaction fired at 90% of the context window, but prompt admission rejects above the window minus the model's output reserve and a 10% safety margin, which is lower for 1,862 of 1,876 catalogued models: opencode-go/deepseek-v4.1-flash rejected at 516,000 input tokens while compaction waited for 900,000. Compaction now triggers at compaction.maxUsageRatio of that ceiling, less pending tool-result and image reserves (464,400 for that model). A prompt still over the ceiling gets one automatic compaction and a re-check before it is rejected; that rejection reports the committed compaction as a side effect.Context limit reached while /compact answers Already compacted. Admission kept counting the provider usage reported before the compaction: one anthropic/claude-opus-5-5 session was rejected at 808,236 estimated tokens against a 772,000-token ceiling after its history had shrunk to about 13,500. Usage recorded at or before the latest compaction no longer counts, and a repeated /compact re-cuts the tail the previous compaction kept, using the 4,096-token emergency keep budget. The visible reasoning of turns a later user message closed, which providers drop, no longer counts toward the next turn's estimate.devin/swe-2, configured with a 262,000-token window, rejected even a 45-token first message because its tool schemas (328 MCP tools plus the built-ins) were estimated at 237,218 tokens, over its 219,416-token input ceiling. Requests to such a model now withhold whole MCP servers, largest schema first, until a fully compacted session fits under the compaction trigger. Withholding stops only once a recount of the remaining schemas fits. The selection is fitted again for every turn, and within a turn whenever the model, the system prompt, a tool schema or the tool-to-server mapping changes, so an MCP server that reconnects with larger schemas under the same tool names is withheld from the next request. The active tool set is unchanged, a warning names the withheld servers, and they return on a model with room. A prompt rejected because the system prompt and tool schemas alone overflow is reported as configuration.invalid, and input still too large after automatic compaction as compaction.failed, instead of provider.context_overflow.AgentSession.close() and runtime disposal retain the session owner lease until registered work and native MCP transport closure settle. Legacy busy disposal starts the same close rather than releasing ownership early.RangeError.retry.baseDelayMs without limit, and Node fires a longer timer after 1 ms, so a base of 3,000,000,000 ms retried after about 1 ms. Delays now stop at that limit, and the exported computeRetryDelayMs returns at most 2,147,483,647; below it, every result for a base that converts to a non-negative number is unchanged. A base that converts to NaN or a negative number uses the 2 s default, +Infinity takes the limit, and a retry.maxRetries that converts to NaN now means no retries; before, neither retry ran out.accounts/fireworks/models/kimi-k3, moonshotai/Kimi-K3), which keeps their transport contract. OpenCode Go defaults to deepseek-v4.1-flash: OMK has not verified the request contract of its Kimi K3 or K2.7 Code, and models.dev marks its K2.6 deprecated.Release notes live in RELEASE_NOTES_v1.3.0.md.
omk provider adopt [<id>] [--from <source>] [--dry-run] [--json] [--status] copies an existing Codex CLI or Claude Code CLI login into OMK's credential store, so a subscription does not have to be signed in twice. Sources are read-only, --from narrows to the provider's own mapping, and no token material reaches output.agent-session reports the store error instead of advising /login, and a transient lock contention is retried instead of deciding authentication for the whole process lifetime.omk provider doctor accepts engine-registered API types (devin-agent, cursor-agent) instead of rejecting them as unsupported.custom tails are rebased, and a length stop that produced no summary text fails instead of committing an empty summary. The compaction source-entry bound is 65,536.addOAuthAccount merges imported accounts inside the storage lock (no lost update across concurrent sessions) and never overwrites a usable stored refresh token with an absent one.Release notes live in RELEASE_NOTES_v1.2.4.md.
Release notes live in RELEASE_NOTES_v1.2.3.md.
OMK builds on pi — Mario Zechner's
MIT-licensed coding-agent harness — and began from the
oh-my-pi fork. The vendored tree was
removed in this release line; OMK 0.9x is OMK-native (see
specs/constitution.md), and the design debt to both
projects stands. Thank you.
MIT
31 followers · starred Jul 2026
TypeScript
94.8%
JavaScript
3.7%
Open-source coding agent CLI. Switch models in one session and finish explicit goals against checks you approve.
TypeScript
146
17,531 commits
updated Oct 4, 2026
Open Multi-Agent Kit
Make “done” pass a check.
A terminal coding agent with model switching and explicit, test-backed goals.
Quick start ·
Default or opt-in? ·
Evidence and limits
Choosing OMK ·
Documentation
An AI saying “done” is not an acceptance check. With /goal verify, you choose
the command that checks your goal. OMK runs it after each settled turn and
completes the goal only when it passes on the current workspace. A later edit
makes that evidence stale. The result is only as strong as the check you choose.
/goal Fix the failing test without changing its assertions.
/goal verify node --test check.test.mjs
Use a test that exists in your project, or follow the small, reproducible goal demo. This is an explicit workflow; ordinary prompts do not enable the gate.
Choose a model, work on your repository, then switch models with /model when
another one suits the next step. The conversation stays in the same session.
You can also stop and return later with /resume or omk -c.
OMK is a standalone CLI, not a plugin for Claude Code or OpenCode. It supports subscription providers, API keys, and local models. Start with one agent that reads files, edits code, and runs commands. Add subagents or explicit verification workflows when you need them; neither is required for your first task.
Requires Node.js 22.19 or newer. Start in the repository you want to work on:
npm install -g open-multi-agent-kit --ignore-scripts
omk --version
cd your-project
omk
Without a global install, run npx --ignore-scripts open-multi-agent-kit from
that directory.
/login to authenticate a supported subscription or API-key provider./model to choose an available model.Summarize this repository and identify the commands used to check it.
Read the project configuration to support your answer. Do not edit files.
After the reply, use /model to choose another configured model and ask it to
review the answer. You stay in the same session. This is manual model switching,
not parallel agents or an independent correctness check.
Ready to see a test-backed goal? Try the goal demo, then share your result or a reproducible failure in a GitHub issue. If this workflow is useful, star OMK to help other developers find it.
For a bug fix, name the failing behavior and ask for a regression test, the smallest fix, and the check commands with their exit codes. Review the diff and those results yourself; a request to run tests does not enable a verification gate.
Built-in local bash requires sandbox-exec on macOS or bwrap plus
unprivileged user namespaces on Linux. It blocks network access and fails
closed if the backend is missing. See the safety boundary
and full quickstart for setup.
A fresh install starts one agent/tool loop after provider setup. It does not turn each prompt into a multi-agent workflow or automatically certify its answer.
| Capability | Fresh-install behavior | Where to start |
|---|---|---|
| File editing, shell commands, saved sessions | Built in; tools run when called by the agent | Usage, sessions |
| Tool-call scheduling | dag-v2 schedules resource conflicts within the agent loop; it does not launch a team | Runtime algorithms |
| Subagents | Optional extension; load it and supply agent definitions | Subagent setup and examples |
| MCP servers, extra skills and extensions | Require configured servers or installed resources | MCP, skills, extensions |
| Durable goals | Built in; /goal <objective> continues after each settled turn, up to 8 rounds. With /goal verify <command> the goal completes only when that check passes on the current workspace | Acceptance checks |
| Protocol verification and advisory judging | Explicit API/workflow opt-in; not a gate on ordinary prompts | Run protocol |
Verified runs (omk run) | Opt-in CLI and SDK on Linux with bwrap; runs an approved command in an isolated copy, checks the result, and supports resume, cancel and cleanup | Verified run |
| Context budgeting | Off by default | Settings |
| AdaptOrch integration | Optional and separate; no service calls by default | OMK + AdaptOrch |
The internal lane launcher and automatic command-sharding primitives are not connected to the default CLI path. Installing their packages is not the same as enabling an orchestration workflow.
No comparative benchmark result is published here yet. We have not established that OMK solves more tasks than another harness, that multi-agent execution improves success, or how much verification reduces false completion.
OMK targets state-of-the-art quality as a CLI coding-agent harness. SOTA is not verified.
The evidence you can inspect today covers specific failure modes:
| Behavior covered | Regression evidence | Scope |
|---|---|---|
Missing test observations produce inconclusive; a required failing test produces fail | Protocol tests | Explicit protocol evaluation, without a waiver |
| Changed artifacts, wrong command bindings, or missing ledger evidence block acceptance | Evidence binding tests | Strict evidence gate and selected workspace scope |
| A relevant workspace mutation after verification makes the receipt stale | Freshness tests | Configured receipt and mutation tracking |
| A durable goal with an approved acceptance check completes only on a passing receipt, and an edit after the check makes that receipt stale | Goal acceptance tests, live loop tests | Approval held by the running OMK process; git work tree; static command lines |
| An effect that may still be live keeps its resource claims through expiry, cancellation and authority restart, until a supervisor confirms termination | Coordination broker tests | In-process broker with canonical claim keys; no OS fencing |
| A publication is refused unless its read versions, parent revision and receipt binding all still match | Publication tests | Single accepted snapshot pointer; no multi-file filesystem atomicity |
| Cancellation after dispatch is never reported as cancelled-before-dispatch; the outcome stays unknown until settled | Operation lifecycle tests | Pure state machine; does not itself stop a remote effect |
| Zero trials is reported as absent evidence rather than zero risk, and a point estimate is not an error rate | Risk bound tests | Binomial model under a fixed policy and adequately independent samples |
These tests exercise the gates, not the rate at which they catch real bugs.
An ordinary prompt finishes when its tool loop and queued work settle;
prompt_settled is not a correctness verdict.
A useful comparison must hold the model, provider configuration, tasks, budget,
and tool permissions constant, and label default versus opt-in workflows.
Report task success, cost, latency, and false completion (reported complete but
failing the declared checks), with its denominator and per-task outcomes.
The measurement protocol
defines the reproducibility and privacy requirements. omk stats shows local
turn costs and tool failures; it does not score task correctness.
If you evaluate OMK, share a sanitized report and reproduction steps in a GitHub issue. Include failed and interrupted runs, not just successful examples.
The terminal UI shows the selected model, tools, and session status. Additional signals depend on the integrations you configure.
The header reads omk v<package.version> · OMK//CONTROL; the installed package
version is the source of truth.
OMK's provider-neutral coding-agent CLI also exposes a multi-agent control plane for explicitly configured workflows. The diagram describes that design, not what every prompt automatically runs.
The v0.98.3 SDK rejects incomplete first-party judge responses and exposes deterministic ties; it is not an automatic TUI judge.
The animation changes once every 1.5 seconds and contains no flashing. The four steps above are the complete text alternative.
AgentSession built-in local bash uses OS sandbox enforcement by default:
sandbox-exec on macOS and bwrap plus unprivileged user namespaces on Linux.
Local shell spawns restrict writes to the workspace and OS temporary directory,
disable network access, and fail closed with sandbox.backend_missing when an
enforcement backend is unavailable.
This is not read-confidentiality or whole-process containment. Other file
tools, extension and custom-tool code, injected or remote BashOperations, and
the OMK process keep the permissions of the process running them. Use
containerization when the
boundary must cover more than built-in local bash. Explicit evidence workflows
cannot treat missing required evidence as a verified result.
OMK keeps routing separate from control and evidence. Codex, Claude Code,
OpenCode Zen/Go, Kimi, GLM/ZAI, native xAI/Grok, NVIDIA NIM, and local providers
can participate through omk-ai while the run contract stays stable.
Native xai keeps subscription OAuth and XAI_API_KEY billing separate. See
provider setup,
provider resilience, and
Grok integration.
| Package | Purpose |
|---|---|
open-multi-agent-kit | Interactive coding-agent CLI and control plane |
omk-agent-core | Agent runtime, tool execution, and DAG scheduling |
omk-ai | Unified multi-provider LLM API |
omk-protocol | Versioned run contracts and semantic reducers |
omk-adaptorch-wpl | Work Packet Loop runtime |
omk-book-to-skill | Optional document-to-skill compiler |
omk-tui | Differential-rendered terminal UI library |
npm install omk-agent-core
npm install omk-ai
npm install omk-protocol
omk install npm:omk-book-to-skill@0.98.3
npm install omk-tui
v0.97.0 shipped the OpenWiki policy and workflow, but no versioned corpus or
integrity checker. The following integrity/output guards shipped in v0.98.0;
the generated corpus remains optional and is not bundled:
openwiki/ — absent. The previous untracked corpus was removed after the
hardened gate proved it carried fabricated evidence: 8 frontmatter symbols
that no declared source path defines (AgentLoop, getModel, DeepWall,
loadExtensions, createExtensionRuntime, main), 45 references to @omk/*
package names this repository does not publish, and a
restatement of this README's Scope -> Route -> Verify -> Replay loop as a
strict engine state machine, which is not what the source implements.
CI regenerates the corpus; nothing is lost.scripts/check-openwiki.mjs — shipped integrity checker. An interrupted corpus
now fails unless openwiki/.manual-review.json binds a review to the exact
corpus digest, and every frontmatter symbol must bind to one of that page's
own source_paths as a whole identifier.scripts/check-openwiki-output.mjs — output gate. The scheduled workflow
may write, upload, and open a PR for openwiki/ and nothing else, so a model
reading this repository cannot reach AGENTS.md, CLAUDE.md, or the workflow
that runs it. The gate runs once before the artifact leaves the read-only
generating job and again before the PR, because the publishing job holds write
permissions the first one does not..understand-anything/ — optional local structural graph used by Pi Lens;
it is not published or injected into prompts by default. To reach a session,
attach it through OMK's MCP client like
any other server; there is no second, bespoke path for it.Source and tests remain authoritative. Shipped guards do not turn a generated index into authority: treat corpus pages as local advisory data and recheck source.
A corpus no session can read is documentation of a plan, not a feature, so the
pages are now candidates for prompt budgeting. Enable contextBudget.openwiki
alongside contextBudget.enabled
(settings)
and each page becomes a low-priority evidence item ranked against the turn's
query. Pages compete for leftover budget and can never displace instructions or
skills; most turns carry titles and declared symbols alone, and a page's text
arrives only when the query earns it.
Admission mirrors scripts/check-openwiki.mjs rather than restating it. A
complete corpus at the current HEAD offers page text; one whose HEAD has
moved offers titles only and is marked stale; an interrupted corpus is refused
unless a review binds to its exact digest. The default is off, and with the
setting off the prompt is byte-identical to one built without a corpus.
OMK is this local, MIT-licensed coding agent. AdaptOrch is a separate proprietary evidence service. Neither requires the other: installing OMK does not create an AdaptOrch account or make calls to it by default.
For an optional integration, see the WPL package and clients and MCP setup. The WPL package exposes state, client, and adjudication primitives, not an automatic verification loop for every CLI prompt.
AdaptOrch's reports carry correctness_claim=false; they are not semantic
correctness proofs or OMK harness benchmark results.
Review AdaptOrch plans
· Claim boundary
The AdaptOrch name and marks identify that separate proprietary product and appear here with permission. They are excluded from this repository's MIT grant — see LICENSE.
The design decisions behind OMK's context, routing, memory, and orchestration layers are grounded in published work rather than invented in isolation. Each row below was retrieved and read directly; claims are at abstract level, which is the evidence grade this table asserts and no more.
| Paper | Mechanism it establishes | OMK implementation or design reference |
|---|---|---|
| arXiv:2608.22752 — The Compaction Cliff in Long-Running AI Agent Memory | Uniform summarization erodes rules and episodic logs at the same rate; measured safety-rule retention falls to 53% after one compaction and 10% after five. Type-tagged deterministic operators fix it. | Type-aware compaction triage: rule-typed items survive N rounds byte-identical |
| arXiv:2608.23023 — Most of the LLM Routing Gap Is Task Type | Most routing gain is reachable with a fixed task-type table; run-to-run flips must not be credited as wins. | Frozen task-class table plus the 2-run stability rule in the promotion gate |
| arXiv:2506.16655 — Arch-Router: Aligning LLM Routing with Human Preferences | Indirection: a classifier emits a label, a policy table maps label to decision, so models change without retraining. | classifyTaskV4 plus TASK_CLASS_THINKING_LEVELS |
| arXiv:2605.09894 — Deterministic vs. LLM-Controlled Orchestration | Holding model, prompts, and tools constant and varying only execution control, deterministic orchestration matched accuracy, improved worst-case robustness, and cut tokens up to 3.5x. | Deterministic scheduler and planned lanes; execution control is never delegated to the model |
| arXiv:2608.15565 — Admission Without Answers | Label-free admission on execution success alone admits substantial contamination; an accept/abstain/escalate decision is required. | Verified-memory admission design (spec 019), abstain is not stored |
| arXiv:2608.23471 — InjecMEM: Memory Injection Attack on LLM Agent Memory Systems | Single-interaction memory injection is a reproduced attack frame against agent memory. | Retrieved memory is injected only as provenance-tagged data, never fused into instruction position |
Entries include implemented mechanisms and design proposals; check the runtime status guide for availability. The wider survey, including approaches not adopted, is working material that is not published with the repository.
npm ci --ignore-scripts
npm run build
npm run check
npm test
npm run release:local
Direct dependencies are pinned, CI installs with --ignore-scripts, and the
published CLI includes a generated npm-shrinkwrap.json. Read
CONTRIBUTING.md and the
development guide before sending a
change.
Use it for provider choice within one CLI session, or to build workflows against its public runtime and evidence APIs. For a single-provider workflow, your current agent may be sufficient. Try the read-only task above before moving existing work.
OMK is a separate runtime with its own CLI, sessions, tool scheduler, and SDK.
One reason to choose it is to build your own acceptance workflow: define
required test observations in the run protocol,
then have your automation reject fail or inconclusive results. Receipt
integrity and freshness still need their own configured checks.
For adding a tool or prompt to an existing OpenCode setup, a plugin may be the smaller change. OMK's protocol is opt-in, not proof of better performance.
Not automatically. Subagents require setup, and verification must be part of the chosen workflow. Its result covers the declared checks, not all behavior. See what runs by default and evidence and limits.
Historical correction: the immutable v0.97.0 notes below announced a versioned OpenWiki corpus, but that release still ignored
/openwiki/and did not contain the corpus or checker. See the current repository-understanding section above for the shipped guards and optional-corpus boundary.
/goal verify <command> approves an acceptance check that runs in the default bash sandbox after each settled turn. The goal completes only when the check passes on the current workspace, and the check's output never reaches the model. See acceptance checks.omk run pauses instead of failing and resumes with restart-writer, resume or retry-tasks. omk run cancel stops a run from another shell, and omk run gc prunes derived workspaces while keeping the evidence. See cancellation and artifact GC.failed with a classified error, and every other server still contributes its tools. See failure behavior.SIGINT, SIGTERM or the new omk run cancel) no longer ends it failed. The run stays paused with failure: cancelled, and its journal records a new interrupted event. Resume the phase that was cut off with restart-writer, resume or retry-tasks; a DAG attempt whose process was confirmed stopped is released instead of spent. A verification check cut off by the cancellation is no longer signed into the receipt as a failed check. OMK 1.2.4 and earlier cannot read a journal that contains interrupted, and a status consumer that treated cancellation as terminal must handle paused. See cancellation.OMK_SESSION_CONTROL=1 or session.startControl(), then use sdk session ... --live. Exact session IDs and private per-enrollment endpoints are required; live failures never become transcript writes.OMK_VERIFIED_MEMORY=1 plus Context Budget V2 and uses transient tool-result data, not instruction text. No automatic extraction or quality-improvement claim./goal verify <command> approves an acceptance check for the durable goal and runs it through receipt-bound local bash under the default bash sandbox. A passing check is attached as goal evidence. After each settled turn the check runs again: a pass completes the goal, and a failure starts the next round with the command and its exit code, never its output. /goal complete then requires a passing receipt from this session that still matches the workspace; a tracked edit, new file or HEAD move after the check makes it stale. Approvals stay in the OMK process, so approve the check again after a restart. See acceptance checks.nextDurableGoalTimestamp(goal) is exported for SDK callers of applyDurableGoalCommand() and DurableGoalStore.transition(). It returns the wall clock, raised to no earlier than the goal's last update and later than the start of its generation, which is the time the reducer accepts after the clock steps back (seen on WSL2). See durable goal lifecycle.omk run cancel ID [--wait-ms N] cancels a verified run owned by another process. It writes a request file that the owner checks every 250 ms and never signals a PID. watchRunCancelRequest() connects the same request to an SDK caller's AbortSignal, and cancelVerifiedRun() is the SDK form of the command.omk run gc [--older-than DURATION] [--execute] prunes the derived workspaces (writer*, candidate*, tasks) of verified runs that can no longer be recovered, holding each run's owner lease while it checks. It keeps journals, keys, manifests, blobs, receipts and attestations, never follows symlinks, and only reports without --execute. SDK: collectVerifiedRuns().AgentSession checks the complete request against the model window less the output reserve and safety margin (computeHardPromptInputLimit()). estimateContextInputTokens() counts the system prompt, the messages after convertToLlm() and the tool schemas, and takes the largest of the configured tokenizer count, the character heuristic and projected provider usage.session.metacognition exposes a content-free state snapshot and the latest bounded lastDiagnostic, observed at prompt preflight and settlement. It is observation-only: it never rewrites a prompt, authorizes a tool, changes termination or grants completion.McpServerConfig.inheritEnv: false keeps the parent environment out of a stdio server, maxPendingWriteBytes (default 16 MiB) bounds bytes queued on the server's stdin, and await manager.closeAndWait() joins physical transport close. manager.status() marks a server whose close is pending with retiring: true. inheritEnv is not read from mcp.json yet.planSkills() prunes its exact search and adds dominance and eviction-refill passes to its greedy search.sel-4-codeunit, so entries cached under the earlier policy are not reused; the public optimizer identifier is unchanged.omk run reports an error outside the verified-run contract as verified-run: operation_failed (<kind> <code>), for example (Error ENOTDIR), instead of a bare operation_failed. The message itself stays out of the output because it can carry absolute paths or contract text.ps, so Korean and other non-English parent locales no longer make a live process appear unavailable. The acceptance-check regression fixtures use canonical physical temporary repository paths on macOS, preserving workspace-mismatch rejection across /var and /private/var aliases.brace-expansion 5.0.12 and undici 8.10.2, plus undici 6.28.1 for the optional Gondolin example. That example still depends on node-forge 1.4.0, whose RSA signature-verification advisory has no patched npm release as of 2026-10-04; the repository production audit continues to report it.agent_end and sent the next turn while that run still owned the session; the session rejected it with Agent is already processing, so the round was spent and the goal never continued. The controller now acts once a turn settles, when no automatic retry follows it, and queues the next round as a follow-up. An attempt that is about to be retried no longer uses up a round.goal timestamps must be monotonic or goal generation timestamp must advance when the wall clock steps back, as WSL2 does when its hypervisor resyncs time. The controller dates each transition no earlier than the journal's last timestamp.operation_failed when the wall clock steps back under load. The authority store's default clock is the wall time at open plus monotonic elapsed time; a clock injected by the caller that moves backward is still refused.writer_incomplete, and omk run status suggests recovery only for runs that are running or paused.failed with a classified public error while every other server still contributes its tools. A malformed tool result rejects with mcp.invalid_tool_result instead of counting as success. Request and handshake timeouts accept only safe integers from 0 through 2,147,483,647 ms, and 0 refuses to send. manager.status() omits free-form server-reported versions.ownership.dispatch_active. Numeric environment values with trailing characters are rejected instead of parsed as a prefix, and prompt-size estimation projects each tool's name, description and parameters instead of serializing the tool object.RpcClient keeps only the last 8,192 characters of the current child's stderr and clears them on start(). waitForIdle() and collectEvents() reject when the child fails, exits or is stopped, and one waiter unsubscribing no longer makes another miss agent_end. stop() rejects pending requests at once, returns without the one-second delay when the child already exited, and rejects with RpcTerminationUncertainError while keeping ownership when termination cannot be confirmed. prompt() propagates a server rejection, so promptAndWait() no longer waits for its 60-second timeout.SessionManager.getBranch() no longer shifts the path array for every ancestor (2,096,128 element moves at depth 2,048, now none). The run journal builds its frozen record copy only when records is read, and the memory-only journal store no longer replays every earlier record on each append (8,256 hashes for 128 appends, now 128). An AgentSession listener that unsubscribes while an event is dispatched no longer makes the next listener miss it.omk-ai and omk-agent-core fixes apply: complete() and completeSimple() no longer queue every stream event until they return, a Cursor request that reaches its deadline is closed instead of left running, provider retries reject invalid options and stop early on an aborted signal, and taking queued messages one at a time no longer copies the rest of the queue.attemptId, process settlement and stream receipts instead of the first attempt's, and the README install list includes managed-process-tree.ts, subagent-stream.ts and graph-result.ts, without which the extension did not load.Context limit reached until a manual /compact. Threshold compaction fired at 90% of the context window, but prompt admission rejects above the window minus the model's output reserve and a 10% safety margin, which is lower for 1,862 of 1,876 catalogued models: opencode-go/deepseek-v4.1-flash rejected at 516,000 input tokens while compaction waited for 900,000. Compaction now triggers at compaction.maxUsageRatio of that ceiling, less pending tool-result and image reserves (464,400 for that model). A prompt still over the ceiling gets one automatic compaction and a re-check before it is rejected; that rejection reports the committed compaction as a side effect.Context limit reached while /compact answers Already compacted. Admission kept counting the provider usage reported before the compaction: one anthropic/claude-opus-5-5 session was rejected at 808,236 estimated tokens against a 772,000-token ceiling after its history had shrunk to about 13,500. Usage recorded at or before the latest compaction no longer counts, and a repeated /compact re-cuts the tail the previous compaction kept, using the 4,096-token emergency keep budget. The visible reasoning of turns a later user message closed, which providers drop, no longer counts toward the next turn's estimate.devin/swe-2, configured with a 262,000-token window, rejected even a 45-token first message because its tool schemas (328 MCP tools plus the built-ins) were estimated at 237,218 tokens, over its 219,416-token input ceiling. Requests to such a model now withhold whole MCP servers, largest schema first, until a fully compacted session fits under the compaction trigger. Withholding stops only once a recount of the remaining schemas fits. The selection is fitted again for every turn, and within a turn whenever the model, the system prompt, a tool schema or the tool-to-server mapping changes, so an MCP server that reconnects with larger schemas under the same tool names is withheld from the next request. The active tool set is unchanged, a warning names the withheld servers, and they return on a model with room. A prompt rejected because the system prompt and tool schemas alone overflow is reported as configuration.invalid, and input still too large after automatic compaction as compaction.failed, instead of provider.context_overflow.AgentSession.close() and runtime disposal retain the session owner lease until registered work and native MCP transport closure settle. Legacy busy disposal starts the same close rather than releasing ownership early.RangeError.retry.baseDelayMs without limit, and Node fires a longer timer after 1 ms, so a base of 3,000,000,000 ms retried after about 1 ms. Delays now stop at that limit, and the exported computeRetryDelayMs returns at most 2,147,483,647; below it, every result for a base that converts to a non-negative number is unchanged. A base that converts to NaN or a negative number uses the 2 s default, +Infinity takes the limit, and a retry.maxRetries that converts to NaN now means no retries; before, neither retry ran out.accounts/fireworks/models/kimi-k3, moonshotai/Kimi-K3), which keeps their transport contract. OpenCode Go defaults to deepseek-v4.1-flash: OMK has not verified the request contract of its Kimi K3 or K2.7 Code, and models.dev marks its K2.6 deprecated.Release notes live in RELEASE_NOTES_v1.3.0.md.
omk provider adopt [<id>] [--from <source>] [--dry-run] [--json] [--status] copies an existing Codex CLI or Claude Code CLI login into OMK's credential store, so a subscription does not have to be signed in twice. Sources are read-only, --from narrows to the provider's own mapping, and no token material reaches output.agent-session reports the store error instead of advising /login, and a transient lock contention is retried instead of deciding authentication for the whole process lifetime.omk provider doctor accepts engine-registered API types (devin-agent, cursor-agent) instead of rejecting them as unsupported.custom tails are rebased, and a length stop that produced no summary text fails instead of committing an empty summary. The compaction source-entry bound is 65,536.addOAuthAccount merges imported accounts inside the storage lock (no lost update across concurrent sessions) and never overwrites a usable stored refresh token with an absent one.Release notes live in RELEASE_NOTES_v1.2.4.md.
Release notes live in RELEASE_NOTES_v1.2.3.md.
OMK builds on pi — Mario Zechner's
MIT-licensed coding-agent harness — and began from the
oh-my-pi fork. The vendored tree was
removed in this release line; OMK 0.9x is OMK-native (see
specs/constitution.md), and the design debt to both
projects stands. Thank you.
MIT
31 followers · starred Jul 2026
TypeScript
94.8%
JavaScript
3.7%