boshu2/agentops

DevOps discipline for AI coding agents: shape the work, track it as a graph, and get each change judged by a context that didn't write it.

Go

447

5,374 commits

updated Oct 2, 2026

See the code

README

AgentOps

AgentOps

DevOps discipline for AI coding agents: shape the work, track it as a graph, and get each change judged by a fresh agent session that didn't write it.

Validate License: Apache-2.0 Latest release Skills

Install · The loop · Goals · Try it · Skills

AgentOps provides optional skills and a CLI (ao). The same SKILL.md skills work with coding agents (Claude Code, Codex, Cursor, OpenCode, Gemini CLI, Pi and others) and personal assistants (OpenClaw, Grok Bot). You state intent as behavior in your domain's words. The skills carry it through one change (Plan → Implement → Validate, an RPI) or, for bigger work, a goal made of many RPIs tracked in Beads, a dependency-aware issue tracker.

Why use AgentOps?

When the agent…AgentOps adds
Builds something different from what you meantGiven/When/Then examples shared by implementation and review
Uses three names for one conceptOne domain term per concept, in intent, code and tests
Says “done” after a green test runA fresh judge that didn't write the change
Loses the thread on work bigger than one sessionA Beads graph holding intent, dependencies and verdicts
Runs off with a half-formed goalAn interview that settles the goal before agents go autonomous
Gives one model's answer to a hard callA council: judges in fresh contexts, each with the model, effort and perspective you assign (one model family or several vendors), compare, duel (score each other's ideas) or debate to your majority, keep dissent, and can answer an interview for you; Idea Genie brainstorms options
Loses its plans, research and decisions when the session endsPlans and decisions saved on the bead or issue (Plan, Interview, Navigate); research and idea reports under .agents/; council reports where you choose
Repeats the last session's investigationMemory turns reviewed, disclosure-checked lessons into .context/ pages safe to commit

Quickstart

Pick one method per agent: a plugin plus npx on the same agent gives you every skill twice.

Claude Code

claude plugin marketplace add boshu2/agentops
claude plugin install agentops@agentops-marketplace
claude plugin details agentops@agentops-marketplace

Check that agentops appears in the plugin inventory. The bundle includes skills, four agents and tool-call guards.

Codex

codex plugin marketplace add boshu2/agentops
codex plugin add agentops@agentops-marketplace
codex plugin list --json

Check that agentops appears in the inventory. Skills use the agentops: prefix; custom roles and read limits have separate setup.

Everything else (Cursor, OpenCode, Gemini CLI, Pi, OpenClaw, Grok Bot)

With Node.js installed, run from your project directory, then pick your agents and skills:

npx skills@latest add boshu2/agentops

Add -g for a user-level install. In scripts, name the agents: npx skills@latest add boshu2/agentops -g -a cursor opencode -y (-y without -a can install into every agent the installer knows). Installer targets include cursor, opencode, gemini-cli, antigravity, pi, grok (Grok Build) and openclaw. Grok Bot has no installer target; add the same SKILL.md folders through its skill settings. Some skills need extra tools (install guide); what each host has been tested for is in host coverage and limits.

Start a new session so the skills load. Most skills need only your coding agent; Validate also needs the ao CLI. Invocation names vary by agent: this README shows Claude Code's /agentops:<skill>; Codex uses $agentops:<skill>.

The operational loop

Each change is shaped, built and judged. You (or Plan) write intent as behavior (BDD), using one word per concept (DDD's ubiquitous language). For a system that calls queued work a Job, in Gherkin:

Feature: Job redelivery is idempotent
  A Job is one unit of queued work. Delivering it again never repeats its side effect.

  Scenario: A completed Job is delivered again
    Given Job "J-42" completed and charged the customer $20
    When the worker receives Job "J-42" again
    Then it returns the completed result of "J-42"
    And the customer has been charged $20 exactly once

  Scenario: A Job that failed before charging is delivered again
    Given Job "J-43" failed before charging the customer $20
    When the worker receives Job "J-43" again
    Then Job "J-43" completes
    And the customer has been charged $20 exactly once

The feature defines the domain term once; each scenario has concrete data, one action and an observable result. Keep scenarios in the issue or conversation; no .feature file is required.

AgentOps routes: intent goes to Plan when unclear or straight to Implement when clear; Implement runs native checks, then a fresh judgment independent of the author. Accepted work finishes, failed behavior returns to Implement for repair, missing evidence is gathered and judged again. An existing change enters at fresh judgment. An optional learning loop turns results into reviewed .context/ pages that later work queries.

StepSkillWhat it does with the scenarios
ShapeplanTurns the request into scenarios for one small change. Skip it when intent is clear.
BuildimplementMakes the change and tests both scenarios.
JudgevalidateA new session that didn't write it returns PASS, FAIL or NOT_PROVEN against the same scenarios.
LearnmemoryOptional: reviewed .context/ pages that later work can query.

Enter at the step you need; an existing change goes straight to Validate. The author never approves its own work. Merging and releasing follow your repo's rules.

Goals

rpi runs Plan → Implement → Validate for one outcome without check-ins (your agent's permission prompts still apply) and stops at acceptance, a blocker or a spent limit. Bigger work becomes a goal (experimental; needs Beads: brew install beads, then bd init in your repo):

  1. Interview. One question at a time, each with a recommended answer. You settle the outcome, its examples, domain terms, non-goals, authority and budget before agents go autonomous.
  2. Craft Goal. Returns SAFE_TO_CREATE plus a prompt to paste into /goal (Claude Code or Codex), USE_RPI (small enough for rpi), or UNSAFE_GOAL plus what's undecided. It creates nothing itself.
  3. Navigate each round. Picks a few ready work items (beads); each gets one RPI and a fresh Validate. The goal ends ACHIEVED, NOT_ACHIEVED or NEEDS_OPERATOR.

A goal acts as orchestrator: it observes the Beads work graph, picks a bounded wave of ready beads, consumes verdicts, then ratchets or stops. The graph holds a root epic and child beads: A closed with PASS, B discovered from A and ready, C ready, D blocked by C. Picked beads B and C each get one RPI: when the goal delegates, a fresh worker that starts with only that bead plans if unclear and implements, and Validate runs in a separate fresh context. Verdicts and notes are written back to the bead.

Beads holds the plan. Beads (bd) keeps work as a dependency graph outside any conversation, so a goal survives compaction and restarts. The root epic holds acceptance; each child bead is one RPI with its question, scope, notes and verdict. bd ready lists what can start now.

One bead per worker. When the goal delegates, the orchestrator holds the graph and verdicts, and each worker starts with one bead instead of the orchestrator's transcript. Validators start fresh.

bd create "Job redelivery is idempotent" -t epic
bd create "Return the completed result on redelivery" --parent <epic-id>
bd dep add <later-id> <earlier-id>         # real ordering only
bd ready --parent <epic-id>                # the frontier

Navigate shows bd commands; another tracker with status, dependencies and notes works if you map them. AgentOps never builds a second work index.

Try it

Start read-only in any repo, then swap the Job example for your own change.

# First look (changes nothing)
/agentops:research how does this repo validate input? cite files and lines, change nothing

# One change
/agentops:plan make Job redelivery return the completed result without repeating the side effect
/agentops:implement
/agentops:validate     # new session: paste the scenarios, the commit, and the authoring session's ID (your name for a hand-written change)

# One outcome, end to end
/agentops:rpi make Job redelivery return the completed result without repeating the side effect

# A goal
/agentops:interview make the job worker safe under redelivery, retries and crash recovery
/agentops:craft-goal   # then paste its prompt into /goal

Validate an existing change

Pick a finished change whose accepted behavior is recorded in an issue or conversation. Run the required checks, keep the candidate unchanged, and install ao. Then open a new conversation, fill in the references and paste:

Use the AgentOps Validate skill to judge this finished change.
Original accepted behavior: [issue link or original request text]
Candidate: [commit, branch or working tree; list every changed path]
Author context ID: [task/session ID that made the change]
Checks run: [commands and results]
I opened this new conversation for fresh review. Derive the exact subject
identity at the start and end, inspect every changed path against the original
behavior, and do not modify the candidate. Report PASS, FAIL or NOT_PROVEN with
evidence for each criterion, checked, not_checked, author and reviewer context
IDs, and freshness attestation.

PASS needs evidence for every criterion and an empty not_checked. FAIL names failed behavior or an out-of-scope change. Missing proof, identity or path coverage is NOT_PROVEN.

Read-only skill-loading smoke test

Paste this in an agent conversation in your project. It needs no ao CLI.

Use the AgentOps Research skill to trace how this repository validates user
input. Follow one path from the input through its checks and tests. Cite the
files and line numbers, explain one edge case, and identify missing coverage.
Name the Research skill file you loaded. Answer here without changing files.

The reported skill path catches missing or duplicate installs. Invoke Research directly with /agentops:research in Claude Code, $agentops:research in Codex, or / and the installed Research entry in Cursor.

Skills at a glance

All skills are optional. Load one when it answers a specific question. Full catalog: docs/SKILL-ROUTER.md.

GroupSkillsWhat it covers
Operational loopplan implement validateShape, build and judge every change
AutonomousrpiOne outcome, end to end
Goals (experimental)interview craft-goal navigateShape, write and walk a goal over the bead graph
Coordinationorchestrate agent-nativeFresh workers per bead, disjoint scopes, integration
On demandresearch domain test refactor review security doc reverse-engineerReached for when a specific question comes up
LearningmemoryCurated .context/ pages safe to commit
Judgment strategiescouncil premortem postmortem reality-check idea-genieMulti-model councils (debates, idea duels, interview panels), idea brainstorms, plan challenges, postmortems and claim audits
Runtimes and factoriescodex-exec agy-native using-gcSelected executors and Gas City integration
Skill craftskill-builder skill-evalAuthor skills and measure whether they help

Where AgentOps fits

AgentOps grew from applying DevOps experience and established engineering practice to agents. The Practice Registry records the lineage; how it works covers responsibilities.

Compare libraries, trackers and agent factories
Project or toolRole alongside AgentOps
Compound EngineeringA connected development workflow and reusable solution records
Matt Pocock's skillsComposable practices for intent, domain modeling, TDD and review
Beads or your existing trackerOwns work status, dependencies and handoffs
Factories such as Gas CityOwn agent coordination and execution through their native control plane

Choose which workflow leads the task. Carry accepted behavior and evidence into independent judgment. Shared practices are not proof that every combination has been tested.

ao CLI (needed for Validate)

Most skills need only your coding agent. Validate uses ao to identify the exact change it judges.

brew tap boshu2/agentops
brew trust --tap boshu2/agentops
brew install agentops
ao version

With Go installed: go install github.com/boshu2/agentops/cli/cmd/ao@latest. ao init is optional evidence setup; ao config --show inspects configuration; ao gate check runs this repository's own gates (mainly for contributors). See the command reference and installation guide.

Updating and advanced setup

Upgrading to 3.8

Version 3.8 retains existing 3.7 command and skill names. Use the plugin update instructions or, for npx installs, npx skills@latest update (update notes). For Homebrew: brew update && brew upgrade agentops. Start a new session afterward; new installs do not silently remove obsolete copies.

Upgrading from 3.6 or earlier: read the migration guide. Version 3.7 removed commands and skill names, including learn, codebase-recon and swarm; their current owners are memory, research and agent-native. See the 3.8 release notes and 3.7 removals.

Skill dependencies

Skill installation does not install tool dependencies:

SkillNeedsWhy
rpiao, conditionaldelegates exact-subject checks to Validate; only persists verdict.v2 when requested, with the fixed-dispatch adapter optional
planao, conditionalruns ao provenance snapshot-intent with an explicit evidence root when the intent source is not durable
implementao, conditionalat an integration boundary whose changed paths affect bound evidence, runs ao provenance evidence-orphans
validateaoderives exact subject identity with the helper and uses ao provenance store-verdict when persistence is requested; Python/schema checks are developer-only
reality-checkao, conditionalinspect selected goal measurements with ao goals or evidence-store facts with ao status
using-gcaorig prep runs ao gc prepare and ao gc check
docao, optionala requested continuity handoff may use ao session handoff/rehydrate
reverse-engineerpython3Phase 1's mechanical teardown runs scripts/reverse_engineer.py
skill-builderpython3, conditionalCreate mode's build.sh runs scripts/generate-skill-mesh.py; heal/check/audit modes are bash-only
memorypython3, conditionala selected toil investigation can use the repository helper scripts/toil-mining/recent_human.py on cleared Codex sources
securitypython3, conditionalthe composable suite and offline redteam surfaces run security_suite.py when that scan type is selected
Shared project context

Memory can read reviewed .context/ notes with ordinary filesystem tools; no ao or Beads is needed. See this repo's context map.

Adding notes requires authorized sources and fresh review of factual support and disclosure before Git admission. Drafts and evidence stay in protected external storage. These procedures do not enforce access permissions or prove that saved notes improve later work. See Memory's storage rules.

Permissions, optional hooks, and removal

The Claude Code plugin includes PreToolUse guards for private tracker data in commits, manual provenance-ledger edits and installed-skill overwrites. Installing only ao does not add hooks; other paths can opt in through the native hook installer. Read-budget guards, Codex roles and trusted Codex hooks have separate setup.

Disable Claude's plugin with /plugin disable agentops. Remove it with claude plugin uninstall agentops@agentops-marketplace; for Codex, use codex plugin remove agentops@agentops-marketplace; for npx installs, use npx skills@latest remove.

Architecture and saved review evidence

AgentOps is the operations layer for agentic engineering. Its federated integration graph connects evidence while Git owns content, the tracker owns work and the coding runtime or selected factory owns execution. Your repository owns delivery.

Native execution requires zero AgentOps skills. RPI, Gas City and Agentic Coding Flywheel are optional; their completion reports do not replace independent review.

On request, Validate can save verdict.v2 with exact content, checked scope and evidence. New proof belongs in selected, protected storage outside Git; existing evidence is preserved.

Read the architecture, operating contract and storage rules.

Troubleshooting

Common problems and how to report one
SymptomWhat to check
plugin is not recognizedUpdate your agent to a version with plugin support
A skill is missingCheck its inventory or picker, then start a new session
ao is not foundInstall the CLI and check PATH; Go installs usually use $(go env GOPATH)/bin
A skill needs another toolCheck its dependencies in the installation guide
An old skill name failsCheck the migration guide and stale copies

Report a reproducible issue with your runtime version, install method, command or prompt, and observed result. Share only evidence you are authorized to disclose.

Limits

Skills guide agents; installation alone does not enforce their instructions. A green test suite or an agreeing model can still miss a defect. Missing proof stays NOT_PROVEN. Saving notes does not establish improved outcomes or automatic knowledge compounding. See product evidence and limits.

FAQ

CLI, reviewer model and storage questions

Do I need the CLI, an orchestrator or several agents?

No. Start with one coding agent and a skill such as Research, Test or Refactor. Obtain fresh, author-distinct judgment when a change is ready. The Validate skill requires ao; native independent review does not.

Must the reviewer use another model provider?

No. The default is a fresh context from the author's model family. Cross-model review is an explicit choice; independence still matters.

Must durable work live in Git?

No. Keep source in Git, handoffs in your tracker and requested proof in protected external storage, following each owner's access and retention rules.

These independent projects can extend your AgentOps setup. Get their tools and skills directly from their authors; they are not bundled with AgentOps.

DCG, CASS and MS are Jeffrey Emanuel's projects. Their upstream documentation and distribution terms govern their tools and skills.

Contributing

Contributions are welcome: documentation fixes, reproducible bug reports, tests, CLI improvements and skills. Read the contribution guide; to work on skills from a checkout, link them with ao skills link.

Licensed under Apache-2.0.

ai-agents
claude
claude-code
claude-code-plugins
claude-marketplace
claude-skills
codex
codex-plugin
cursor
devops
opencode-plugin

Significant stargazers

Joel Natividad

234 followers · starred Feb 2026

Vlad A. Ionescu

194 followers · starred Feb 2026

Riley Langbein

13 followers · starred Apr 2026

boshu2/agentops

DevOps discipline for AI coding agents: shape the work, track it as a graph, and get each change judged by a context that didn't write it.

Go

447

5,374 commits

updated Oct 2, 2026

See the code

README

AgentOps

AgentOps

DevOps discipline for AI coding agents: shape the work, track it as a graph, and get each change judged by a fresh agent session that didn't write it.

Validate License: Apache-2.0 Latest release Skills

Install · The loop · Goals · Try it · Skills

AgentOps provides optional skills and a CLI (ao). The same SKILL.md skills work with coding agents (Claude Code, Codex, Cursor, OpenCode, Gemini CLI, Pi and others) and personal assistants (OpenClaw, Grok Bot). You state intent as behavior in your domain's words. The skills carry it through one change (Plan → Implement → Validate, an RPI) or, for bigger work, a goal made of many RPIs tracked in Beads, a dependency-aware issue tracker.

Why use AgentOps?

When the agent…AgentOps adds
Builds something different from what you meantGiven/When/Then examples shared by implementation and review
Uses three names for one conceptOne domain term per concept, in intent, code and tests
Says “done” after a green test runA fresh judge that didn't write the change
Loses the thread on work bigger than one sessionA Beads graph holding intent, dependencies and verdicts
Runs off with a half-formed goalAn interview that settles the goal before agents go autonomous
Gives one model's answer to a hard callA council: judges in fresh contexts, each with the model, effort and perspective you assign (one model family or several vendors), compare, duel (score each other's ideas) or debate to your majority, keep dissent, and can answer an interview for you; Idea Genie brainstorms options
Loses its plans, research and decisions when the session endsPlans and decisions saved on the bead or issue (Plan, Interview, Navigate); research and idea reports under .agents/; council reports where you choose
Repeats the last session's investigationMemory turns reviewed, disclosure-checked lessons into .context/ pages safe to commit

Quickstart

Pick one method per agent: a plugin plus npx on the same agent gives you every skill twice.

Claude Code

claude plugin marketplace add boshu2/agentops
claude plugin install agentops@agentops-marketplace
claude plugin details agentops@agentops-marketplace

Check that agentops appears in the plugin inventory. The bundle includes skills, four agents and tool-call guards.

Codex

codex plugin marketplace add boshu2/agentops
codex plugin add agentops@agentops-marketplace
codex plugin list --json

Check that agentops appears in the inventory. Skills use the agentops: prefix; custom roles and read limits have separate setup.

Everything else (Cursor, OpenCode, Gemini CLI, Pi, OpenClaw, Grok Bot)

With Node.js installed, run from your project directory, then pick your agents and skills:

npx skills@latest add boshu2/agentops

Add -g for a user-level install. In scripts, name the agents: npx skills@latest add boshu2/agentops -g -a cursor opencode -y (-y without -a can install into every agent the installer knows). Installer targets include cursor, opencode, gemini-cli, antigravity, pi, grok (Grok Build) and openclaw. Grok Bot has no installer target; add the same SKILL.md folders through its skill settings. Some skills need extra tools (install guide); what each host has been tested for is in host coverage and limits.

Start a new session so the skills load. Most skills need only your coding agent; Validate also needs the ao CLI. Invocation names vary by agent: this README shows Claude Code's /agentops:<skill>; Codex uses $agentops:<skill>.

The operational loop

Each change is shaped, built and judged. You (or Plan) write intent as behavior (BDD), using one word per concept (DDD's ubiquitous language). For a system that calls queued work a Job, in Gherkin:

Feature: Job redelivery is idempotent
  A Job is one unit of queued work. Delivering it again never repeats its side effect.

  Scenario: A completed Job is delivered again
    Given Job "J-42" completed and charged the customer $20
    When the worker receives Job "J-42" again
    Then it returns the completed result of "J-42"
    And the customer has been charged $20 exactly once

  Scenario: A Job that failed before charging is delivered again
    Given Job "J-43" failed before charging the customer $20
    When the worker receives Job "J-43" again
    Then Job "J-43" completes
    And the customer has been charged $20 exactly once

The feature defines the domain term once; each scenario has concrete data, one action and an observable result. Keep scenarios in the issue or conversation; no .feature file is required.

AgentOps routes: intent goes to Plan when unclear or straight to Implement when clear; Implement runs native checks, then a fresh judgment independent of the author. Accepted work finishes, failed behavior returns to Implement for repair, missing evidence is gathered and judged again. An existing change enters at fresh judgment. An optional learning loop turns results into reviewed .context/ pages that later work queries.

StepSkillWhat it does with the scenarios
ShapeplanTurns the request into scenarios for one small change. Skip it when intent is clear.
BuildimplementMakes the change and tests both scenarios.
JudgevalidateA new session that didn't write it returns PASS, FAIL or NOT_PROVEN against the same scenarios.
LearnmemoryOptional: reviewed .context/ pages that later work can query.

Enter at the step you need; an existing change goes straight to Validate. The author never approves its own work. Merging and releasing follow your repo's rules.

Goals

rpi runs Plan → Implement → Validate for one outcome without check-ins (your agent's permission prompts still apply) and stops at acceptance, a blocker or a spent limit. Bigger work becomes a goal (experimental; needs Beads: brew install beads, then bd init in your repo):

  1. Interview. One question at a time, each with a recommended answer. You settle the outcome, its examples, domain terms, non-goals, authority and budget before agents go autonomous.
  2. Craft Goal. Returns SAFE_TO_CREATE plus a prompt to paste into /goal (Claude Code or Codex), USE_RPI (small enough for rpi), or UNSAFE_GOAL plus what's undecided. It creates nothing itself.
  3. Navigate each round. Picks a few ready work items (beads); each gets one RPI and a fresh Validate. The goal ends ACHIEVED, NOT_ACHIEVED or NEEDS_OPERATOR.

A goal acts as orchestrator: it observes the Beads work graph, picks a bounded wave of ready beads, consumes verdicts, then ratchets or stops. The graph holds a root epic and child beads: A closed with PASS, B discovered from A and ready, C ready, D blocked by C. Picked beads B and C each get one RPI: when the goal delegates, a fresh worker that starts with only that bead plans if unclear and implements, and Validate runs in a separate fresh context. Verdicts and notes are written back to the bead.

Beads holds the plan. Beads (bd) keeps work as a dependency graph outside any conversation, so a goal survives compaction and restarts. The root epic holds acceptance; each child bead is one RPI with its question, scope, notes and verdict. bd ready lists what can start now.

One bead per worker. When the goal delegates, the orchestrator holds the graph and verdicts, and each worker starts with one bead instead of the orchestrator's transcript. Validators start fresh.

bd create "Job redelivery is idempotent" -t epic
bd create "Return the completed result on redelivery" --parent <epic-id>
bd dep add <later-id> <earlier-id>         # real ordering only
bd ready --parent <epic-id>                # the frontier

Navigate shows bd commands; another tracker with status, dependencies and notes works if you map them. AgentOps never builds a second work index.

Try it

Start read-only in any repo, then swap the Job example for your own change.

# First look (changes nothing)
/agentops:research how does this repo validate input? cite files and lines, change nothing

# One change
/agentops:plan make Job redelivery return the completed result without repeating the side effect
/agentops:implement
/agentops:validate     # new session: paste the scenarios, the commit, and the authoring session's ID (your name for a hand-written change)

# One outcome, end to end
/agentops:rpi make Job redelivery return the completed result without repeating the side effect

# A goal
/agentops:interview make the job worker safe under redelivery, retries and crash recovery
/agentops:craft-goal   # then paste its prompt into /goal

Validate an existing change

Pick a finished change whose accepted behavior is recorded in an issue or conversation. Run the required checks, keep the candidate unchanged, and install ao. Then open a new conversation, fill in the references and paste:

Use the AgentOps Validate skill to judge this finished change.
Original accepted behavior: [issue link or original request text]
Candidate: [commit, branch or working tree; list every changed path]
Author context ID: [task/session ID that made the change]
Checks run: [commands and results]
I opened this new conversation for fresh review. Derive the exact subject
identity at the start and end, inspect every changed path against the original
behavior, and do not modify the candidate. Report PASS, FAIL or NOT_PROVEN with
evidence for each criterion, checked, not_checked, author and reviewer context
IDs, and freshness attestation.

PASS needs evidence for every criterion and an empty not_checked. FAIL names failed behavior or an out-of-scope change. Missing proof, identity or path coverage is NOT_PROVEN.

Read-only skill-loading smoke test

Paste this in an agent conversation in your project. It needs no ao CLI.

Use the AgentOps Research skill to trace how this repository validates user
input. Follow one path from the input through its checks and tests. Cite the
files and line numbers, explain one edge case, and identify missing coverage.
Name the Research skill file you loaded. Answer here without changing files.

The reported skill path catches missing or duplicate installs. Invoke Research directly with /agentops:research in Claude Code, $agentops:research in Codex, or / and the installed Research entry in Cursor.

Skills at a glance

All skills are optional. Load one when it answers a specific question. Full catalog: docs/SKILL-ROUTER.md.

GroupSkillsWhat it covers
Operational loopplan implement validateShape, build and judge every change
AutonomousrpiOne outcome, end to end
Goals (experimental)interview craft-goal navigateShape, write and walk a goal over the bead graph
Coordinationorchestrate agent-nativeFresh workers per bead, disjoint scopes, integration
On demandresearch domain test refactor review security doc reverse-engineerReached for when a specific question comes up
LearningmemoryCurated .context/ pages safe to commit
Judgment strategiescouncil premortem postmortem reality-check idea-genieMulti-model councils (debates, idea duels, interview panels), idea brainstorms, plan challenges, postmortems and claim audits
Runtimes and factoriescodex-exec agy-native using-gcSelected executors and Gas City integration
Skill craftskill-builder skill-evalAuthor skills and measure whether they help

Where AgentOps fits

AgentOps grew from applying DevOps experience and established engineering practice to agents. The Practice Registry records the lineage; how it works covers responsibilities.

Compare libraries, trackers and agent factories
Project or toolRole alongside AgentOps
Compound EngineeringA connected development workflow and reusable solution records
Matt Pocock's skillsComposable practices for intent, domain modeling, TDD and review
Beads or your existing trackerOwns work status, dependencies and handoffs
Factories such as Gas CityOwn agent coordination and execution through their native control plane

Choose which workflow leads the task. Carry accepted behavior and evidence into independent judgment. Shared practices are not proof that every combination has been tested.

ao CLI (needed for Validate)

Most skills need only your coding agent. Validate uses ao to identify the exact change it judges.

brew tap boshu2/agentops
brew trust --tap boshu2/agentops
brew install agentops
ao version

With Go installed: go install github.com/boshu2/agentops/cli/cmd/ao@latest. ao init is optional evidence setup; ao config --show inspects configuration; ao gate check runs this repository's own gates (mainly for contributors). See the command reference and installation guide.

Updating and advanced setup

Upgrading to 3.8

Version 3.8 retains existing 3.7 command and skill names. Use the plugin update instructions or, for npx installs, npx skills@latest update (update notes). For Homebrew: brew update && brew upgrade agentops. Start a new session afterward; new installs do not silently remove obsolete copies.

Upgrading from 3.6 or earlier: read the migration guide. Version 3.7 removed commands and skill names, including learn, codebase-recon and swarm; their current owners are memory, research and agent-native. See the 3.8 release notes and 3.7 removals.

Skill dependencies

Skill installation does not install tool dependencies:

SkillNeedsWhy
rpiao, conditionaldelegates exact-subject checks to Validate; only persists verdict.v2 when requested, with the fixed-dispatch adapter optional
planao, conditionalruns ao provenance snapshot-intent with an explicit evidence root when the intent source is not durable
implementao, conditionalat an integration boundary whose changed paths affect bound evidence, runs ao provenance evidence-orphans
validateaoderives exact subject identity with the helper and uses ao provenance store-verdict when persistence is requested; Python/schema checks are developer-only
reality-checkao, conditionalinspect selected goal measurements with ao goals or evidence-store facts with ao status
using-gcaorig prep runs ao gc prepare and ao gc check
docao, optionala requested continuity handoff may use ao session handoff/rehydrate
reverse-engineerpython3Phase 1's mechanical teardown runs scripts/reverse_engineer.py
skill-builderpython3, conditionalCreate mode's build.sh runs scripts/generate-skill-mesh.py; heal/check/audit modes are bash-only
memorypython3, conditionala selected toil investigation can use the repository helper scripts/toil-mining/recent_human.py on cleared Codex sources
securitypython3, conditionalthe composable suite and offline redteam surfaces run security_suite.py when that scan type is selected
Shared project context

Memory can read reviewed .context/ notes with ordinary filesystem tools; no ao or Beads is needed. See this repo's context map.

Adding notes requires authorized sources and fresh review of factual support and disclosure before Git admission. Drafts and evidence stay in protected external storage. These procedures do not enforce access permissions or prove that saved notes improve later work. See Memory's storage rules.

Permissions, optional hooks, and removal

The Claude Code plugin includes PreToolUse guards for private tracker data in commits, manual provenance-ledger edits and installed-skill overwrites. Installing only ao does not add hooks; other paths can opt in through the native hook installer. Read-budget guards, Codex roles and trusted Codex hooks have separate setup.

Disable Claude's plugin with /plugin disable agentops. Remove it with claude plugin uninstall agentops@agentops-marketplace; for Codex, use codex plugin remove agentops@agentops-marketplace; for npx installs, use npx skills@latest remove.

Architecture and saved review evidence

AgentOps is the operations layer for agentic engineering. Its federated integration graph connects evidence while Git owns content, the tracker owns work and the coding runtime or selected factory owns execution. Your repository owns delivery.

Native execution requires zero AgentOps skills. RPI, Gas City and Agentic Coding Flywheel are optional; their completion reports do not replace independent review.

On request, Validate can save verdict.v2 with exact content, checked scope and evidence. New proof belongs in selected, protected storage outside Git; existing evidence is preserved.

Read the architecture, operating contract and storage rules.

Troubleshooting

Common problems and how to report one
SymptomWhat to check
plugin is not recognizedUpdate your agent to a version with plugin support
A skill is missingCheck its inventory or picker, then start a new session
ao is not foundInstall the CLI and check PATH; Go installs usually use $(go env GOPATH)/bin
A skill needs another toolCheck its dependencies in the installation guide
An old skill name failsCheck the migration guide and stale copies

Report a reproducible issue with your runtime version, install method, command or prompt, and observed result. Share only evidence you are authorized to disclose.

Limits

Skills guide agents; installation alone does not enforce their instructions. A green test suite or an agreeing model can still miss a defect. Missing proof stays NOT_PROVEN. Saving notes does not establish improved outcomes or automatic knowledge compounding. See product evidence and limits.

FAQ

CLI, reviewer model and storage questions

Do I need the CLI, an orchestrator or several agents?

No. Start with one coding agent and a skill such as Research, Test or Refactor. Obtain fresh, author-distinct judgment when a change is ready. The Validate skill requires ao; native independent review does not.

Must the reviewer use another model provider?

No. The default is a fresh context from the author's model family. Cross-model review is an explicit choice; independence still matters.

Must durable work live in Git?

No. Keep source in Git, handoffs in your tracker and requested proof in protected external storage, following each owner's access and retention rules.

These independent projects can extend your AgentOps setup. Get their tools and skills directly from their authors; they are not bundled with AgentOps.

DCG, CASS and MS are Jeffrey Emanuel's projects. Their upstream documentation and distribution terms govern their tools and skills.

Contributing

Contributions are welcome: documentation fixes, reproducible bug reports, tests, CLI improvements and skills. Read the contribution guide; to work on skills from a checkout, link them with ao skills link.

Licensed under Apache-2.0.

ai-agents
claude
claude-code
claude-code-plugins
claude-marketplace
claude-skills
codex
codex-plugin
cursor
devops
opencode-plugin

Significant stargazers

Joel Natividad

234 followers · starred Feb 2026

Vlad A. Ionescu

194 followers · starred Feb 2026

Riley Langbein

13 followers · starred Apr 2026

Languages

Go

44.3%

Shell

37.9%

Python

15.3%