Catch silent failures in AI agents before your users do
31
stars
336
commits
Python
primary language
Sep 14, 2026
updated
Catch silent failures in AI agent pipelines before production.
Your LangGraph pipeline runs fine — no exception. But three nodes later, something crashes with a KeyError. The real cause? A node upstream silently dropped a field. ARGUS catches this.
Beta, and under active development. ARGUS is early. Expect rough edges and bugs, and expect things to move. Issues and pull requests are welcome. Contributors: join the Discord before opening a PR — that is where updates land.
1. Install
pip install argus-agents
2. Init
argus init
Writes .cursor/skills/argus-debug/ and .claude/skills/argus-debug/. Commit them. The skill already contains the setup prompt.
3. Attach
Ask your editor agent to wire ARGUS. (The skill already contains this AI setup prompt; the landing-page copy is just a fallback.)
from argus import ArgusWatcher
app = ArgusWatcher().attach(graph)
4. Run
Same as always. Failures print in the terminal; clean runs stay silent.
[argus] run 8f3a1c02 silent_failure on retrieve
missing: documents (dropped by search)
argus show last | argus ui
5. Inspect
argus show last
argus fix <id> # paste-ready prompt for the root-cause node
argus ui
Empty dashboard → wrong directory or no run yet. Check project root or $ARGUS_DIR.
Optional — argus key set for the LLM judge. Skip it and you still get heuristics.
AI-powered detection (the semantic judge, LLM investigator, learned trends) uses your own key from the provider of your choice — OpenAI, Anthropic (Claude), or Google (Gemini). Set it once and it's saved locally for every future session:
argus key set # OpenAI by default — prompts, hidden input
argus key set --provider anthropic # or Anthropic (Claude)
argus key set --provider google # or Google (Gemini)
# pass it directly instead of being prompted:
argus key set sk-... --provider openai
# or just export it (env wins over the saved key):
export OPENAI_API_KEY=sk-... # or ANTHROPIC_API_KEY / GEMINI_API_KEY
Configured more than one? Switch the active provider anytime:
argus key use anthropic # activate a provider you already have a key for
argus key show # list configured providers (masked); * marks the active one
argus doctor # reports BYOK provider / hosted / heuristic-only mode
You pick the provider; ARGUS picks a sensible balanced model for each internal call (a cheap model for the frequent per-node checks, a stronger one for root-cause reasoning). Per-provider resolution order: env var (OPENAI_API_KEY / ANTHROPIC_API_KEY / GEMINI_API_KEY) → saved key → (hosted proxy, if you're on the cloud tier) → heuristic-only.
No key? ARGUS still works — it falls back to heuristic-only detection, no crashes.
Hosted cloud sync (argus login) is optional and only applies if a hosted backend is configured.
from argus import ArgusWatcher
watcher = ArgusWatcher()
app = watcher.attach(graph) # StateGraph or already-compiled app
result = app.invoke(initial_state) # run is persisted automatically
ARGUS monitors every node, detects failures, and saves the run. No changes to your node functions.
finalize()is optional.attach()wrapsinvoke()/ainvoke()/batch()/abatch()/stream()so the run is written to.argus/runs/when the outermost call returns — including cyclic graphs. Callingwatcher.finalize()afterwards is a no-op.
Constructor form still works if you compile yourself:
watcher = ArgusWatcher(graph) # uncompiled StateGraph
app = graph.compile()
result = app.invoke(initial_state)
| Problem | Example |
|---|---|
| Silent failures | Node returns {} or drops a required field — no exception, pipeline keeps running broken |
| Semantic failures | Output structure is fine but values are wrong (placeholders, refusals, degraded text) |
| Crash root cause | Traces KeyError at node 5 back to the upstream node that actually dropped the field |
| Contract violations | Output types don't match the next node's expected input schema |
| Latency degradation | Node takes 95%+ of timeout, or suspiciously fast LLM call (likely cached/empty) |
| Conditional path confusion | Unchosen branches correctly shown as "skipped" — not false "crashed" |
Runs in order, each more expensive — only fires when needed. Every status a layer can assign, and how node statuses roll up into the run verdict, is specified in docs/STATUS.md.
Pipelines with loops (LLM -> compiler -> if fail, retry) get special treatment:
retried (not counted as failures)Fix a bug, re-run from the failing node. Skip upstream nodes entirely:
argus replay <run-id> node_7 # re-run from node_7 onward
argus replay <run-id> node_7 --only # just that one node
argus diff <rerun-id> # compare vs original
External API calls (OpenAI, etc.) are recorded by default — replays are free and deterministic.
Spotted the bad value? Fix it in the saved state and resume from there — no code change, no re-running the steps that already worked:
argus replay <run-id> node_7 --set status=OK # correct a value
argus replay <run-id> node_7 --delete stale_field # reproduce a dropped field
argus replay <run-id> node_7 --patch fix.json # a full patch document
argus replay <run-id> node_7 --set status=OK --dry-run # preview, run nothing
Upstream nodes stay frozen, so only the resumed trajectory changes. Paths are dotted with list
indices — items[0].name — and match the field_path ARGUS reports on a failing signal, so you
can paste one straight in. A patch file takes the same three ops:
{
"delete": ["broken_field"],
"set": {"query": "fixed query", "meta.retries": 0},
"merge": {"config": {"temperature": 0}}
}
Patches are strict by default: a mistyped path errors with a "did you mean" hint instead of
silently adding a field (use --create-missing to add new keys). Every patched replay records
the patch it ran with, so the run explains its own divergence from the original.
For subtle quality issues that pattern matching can't catch:
watcher = ArgusWatcher(graph, semantic_judge=True) # opt-in; default is off
LLM evaluates output quality on every node. Catches wrong tone, unhelpful responses, outdated info. Requires a provider key (OpenAI, Anthropic, or Google) — set via argus key set [--provider ...] (see BYOK).
The judge receives all prior evidence — validator failures, anomaly signals, inspection results — so it rules with full context, not just input/output. Every decision includes an audit trail:
{
"pass": false,
"reason": "Validator correctly identified missing resolution_ticket",
"confidence": 0.85,
"evidence_considered": ["validator:payment_check", "anomaly:BA-003"],
"overridden_signals": []
}
evidence_considered — which prior signals the LLM weighedoverridden_signals — which signals the LLM disagreed with (passed despite the flag)A signature that keeps flagging legitimate output — say NL-002 reading the string "none" as a
serialized null — should be silenced, not worked around by contorting your data:
argus ignore NL-002 # everywhere in this project
argus ignore RF-001 --node draft_hook # only on one node
argus ignore --list
argus ignore NL-002 --remove
Suppressions live in .argus/config.json (commit it so the team shares them). A suppressed hit
no longer changes the node's status or the CI gate, but is still recorded on the run
(suppressed_signals) so argus stats keeps counting it. argus doctor lists what's active.
watcher = ArgusWatcher(graph, validators={
"classify": lambda o: (o.get("label") in ["yes", "no"], "unexpected label"),
"*": lambda o: ("error" not in o, "error key present"), # runs on every node
})
Validator failures cannot be overridden by the LLM judge — they are hard constraints.
from argus import ArgusWatcher, ArgusConfig
config = ArgusConfig(
semantic_judge=True, # LLM judge on every node (default: False)
judge_model="gpt-4o", # model for the judge
node_timeout_ms=30000, # flag outputs at ≥95% of this
min_expected_ms=500, # flag suspiciously fast LLM nodes
sample_rate=0.5, # persist 50% of clean runs (save disk)
persist_failures=True, # always persist failed runs
)
watcher = ArgusWatcher(graph, config=config)
argus list # all recorded runs
argus show last # most recent run
argus show <id> # inspect a specific run
argus check <id> # CI gate for an exact run; prints the JSON path checked
ARGUS_RUN_ID=<id> argus check # CI-friendly selection when the id comes from an earlier step
argus check last # newest-file fallback — avoid in a shared workspace
argus check last --format json # same verdict as one JSON object (run_id, overall_status, passed, findings[])
argus check last --fail-on crashed,silent_failure # only these run statuses fail the gate
argus inspect <id> --step <node> # dump raw input/output for a node
argus fix <id> # fix prompt for the root cause, ready to paste
argus replay <id> <node> # re-run from a node
argus diff <id-a> <id-b> # compare two runs
argus stats # signature hit stats, disable/enable/dispute signatures
argus ignore <SIG-ID> [--node N] # silence a noisy signature project-wide or on one node
argus ignore --list # show active suppressions (.argus/config.json)
argus ui # web dashboard
argus doctor # check setup health + LLM mode (BYOK/hosted/heuristic)
argus key set [--provider ...] # save a provider key locally (OpenAI/Anthropic/Google) — BYOK
argus key use <provider> # switch the active provider
argus key show # list configured providers (masked); * marks active
argus key clear [--provider ...] # remove one provider's key, or all
argus login # (optional) sign in for hosted cloud sync
argus logout # clear stored credentials
argus whoami # show current login status
argus update # check for newer release
Silent failures become test failures without changing how you invoke the graph:
pytest --argus
ARGUS auto-wraps StateGraph.compile() / compiled invoke() for the test session. A clean pipeline stays a passing test; missing fields, tool failures, crashes, and semantic degradation fail that test. Tests that never invoke a graph are unchanged. After a standalone CI run, pass its exact id with argus check <id> or ARGUS_RUN_ID=<id> argus check; argus check last only means the newest file and can select a stale or unrelated run in a shared workspace.
argus ui # opens at localhost:7842
Shows all runs, node-level detail, AI analysis, replay diffs, loop iteration badges, and comparison views. No account needed for local use.
If the table is empty, the UI is serving a different .argus than the project that just ran, or there are no runs yet. The empty state shows the path ARGUS is reading and what to do (argus show last, run the graph, check cwd vs project root).
from argus import ArgusSession
session = ArgusSession()
session.set_edges({"fetch": ["classify"], "classify": ["process"]})
fetch = session.wrap("fetch", fetch_fn)
classify = session.wrap("classify", classify_fn)
process = session.wrap("process", process_fn)
state = fetch(initial_state)
state = classify(state)
state = process(state)
session.finalize()
Works with any framework — Prefect, Temporal, plain Python.
ArgusWatcher)argus key set [--provider ...] (optional; all heuristic detection works without it)For AI setup prompts and integration guides, visit arguslabs.in.
v0.11.0 — changelog
See CONTRIBUTING.md. Join the Discord before opening a PR for updates and to talk through the change.
ARGUS is open-core. The open-source core (src/argus/, the argus-agents PyPI
package) is licensed under Apache-2.0 — see LICENSE. The cloud/
directory (hosted/enterprise components) is proprietary — see cloud/LICENSE.
Python
59.9%
TypeScript
23.3%
HTML
15.9%
Catch silent failures in AI agents before your users do
31
stars
336
commits
Python
primary language
Sep 14, 2026
updated
Catch silent failures in AI agent pipelines before production.
Your LangGraph pipeline runs fine — no exception. But three nodes later, something crashes with a KeyError. The real cause? A node upstream silently dropped a field. ARGUS catches this.
Beta, and under active development. ARGUS is early. Expect rough edges and bugs, and expect things to move. Issues and pull requests are welcome. Contributors: join the Discord before opening a PR — that is where updates land.
1. Install
pip install argus-agents
2. Init
argus init
Writes .cursor/skills/argus-debug/ and .claude/skills/argus-debug/. Commit them. The skill already contains the setup prompt.
3. Attach
Ask your editor agent to wire ARGUS. (The skill already contains this AI setup prompt; the landing-page copy is just a fallback.)
from argus import ArgusWatcher
app = ArgusWatcher().attach(graph)
4. Run
Same as always. Failures print in the terminal; clean runs stay silent.
[argus] run 8f3a1c02 silent_failure on retrieve
missing: documents (dropped by search)
argus show last | argus ui
5. Inspect
argus show last
argus fix <id> # paste-ready prompt for the root-cause node
argus ui
Empty dashboard → wrong directory or no run yet. Check project root or $ARGUS_DIR.
Optional — argus key set for the LLM judge. Skip it and you still get heuristics.
AI-powered detection (the semantic judge, LLM investigator, learned trends) uses your own key from the provider of your choice — OpenAI, Anthropic (Claude), or Google (Gemini). Set it once and it's saved locally for every future session:
argus key set # OpenAI by default — prompts, hidden input
argus key set --provider anthropic # or Anthropic (Claude)
argus key set --provider google # or Google (Gemini)
# pass it directly instead of being prompted:
argus key set sk-... --provider openai
# or just export it (env wins over the saved key):
export OPENAI_API_KEY=sk-... # or ANTHROPIC_API_KEY / GEMINI_API_KEY
Configured more than one? Switch the active provider anytime:
argus key use anthropic # activate a provider you already have a key for
argus key show # list configured providers (masked); * marks the active one
argus doctor # reports BYOK provider / hosted / heuristic-only mode
You pick the provider; ARGUS picks a sensible balanced model for each internal call (a cheap model for the frequent per-node checks, a stronger one for root-cause reasoning). Per-provider resolution order: env var (OPENAI_API_KEY / ANTHROPIC_API_KEY / GEMINI_API_KEY) → saved key → (hosted proxy, if you're on the cloud tier) → heuristic-only.
No key? ARGUS still works — it falls back to heuristic-only detection, no crashes.
Hosted cloud sync (argus login) is optional and only applies if a hosted backend is configured.
from argus import ArgusWatcher
watcher = ArgusWatcher()
app = watcher.attach(graph) # StateGraph or already-compiled app
result = app.invoke(initial_state) # run is persisted automatically
ARGUS monitors every node, detects failures, and saves the run. No changes to your node functions.
finalize()is optional.attach()wrapsinvoke()/ainvoke()/batch()/abatch()/stream()so the run is written to.argus/runs/when the outermost call returns — including cyclic graphs. Callingwatcher.finalize()afterwards is a no-op.
Constructor form still works if you compile yourself:
watcher = ArgusWatcher(graph) # uncompiled StateGraph
app = graph.compile()
result = app.invoke(initial_state)
| Problem | Example |
|---|---|
| Silent failures | Node returns {} or drops a required field — no exception, pipeline keeps running broken |
| Semantic failures | Output structure is fine but values are wrong (placeholders, refusals, degraded text) |
| Crash root cause | Traces KeyError at node 5 back to the upstream node that actually dropped the field |
| Contract violations | Output types don't match the next node's expected input schema |
| Latency degradation | Node takes 95%+ of timeout, or suspiciously fast LLM call (likely cached/empty) |
| Conditional path confusion | Unchosen branches correctly shown as "skipped" — not false "crashed" |
Runs in order, each more expensive — only fires when needed. Every status a layer can assign, and how node statuses roll up into the run verdict, is specified in docs/STATUS.md.
Pipelines with loops (LLM -> compiler -> if fail, retry) get special treatment:
retried (not counted as failures)Fix a bug, re-run from the failing node. Skip upstream nodes entirely:
argus replay <run-id> node_7 # re-run from node_7 onward
argus replay <run-id> node_7 --only # just that one node
argus diff <rerun-id> # compare vs original
External API calls (OpenAI, etc.) are recorded by default — replays are free and deterministic.
Spotted the bad value? Fix it in the saved state and resume from there — no code change, no re-running the steps that already worked:
argus replay <run-id> node_7 --set status=OK # correct a value
argus replay <run-id> node_7 --delete stale_field # reproduce a dropped field
argus replay <run-id> node_7 --patch fix.json # a full patch document
argus replay <run-id> node_7 --set status=OK --dry-run # preview, run nothing
Upstream nodes stay frozen, so only the resumed trajectory changes. Paths are dotted with list
indices — items[0].name — and match the field_path ARGUS reports on a failing signal, so you
can paste one straight in. A patch file takes the same three ops:
{
"delete": ["broken_field"],
"set": {"query": "fixed query", "meta.retries": 0},
"merge": {"config": {"temperature": 0}}
}
Patches are strict by default: a mistyped path errors with a "did you mean" hint instead of
silently adding a field (use --create-missing to add new keys). Every patched replay records
the patch it ran with, so the run explains its own divergence from the original.
For subtle quality issues that pattern matching can't catch:
watcher = ArgusWatcher(graph, semantic_judge=True) # opt-in; default is off
LLM evaluates output quality on every node. Catches wrong tone, unhelpful responses, outdated info. Requires a provider key (OpenAI, Anthropic, or Google) — set via argus key set [--provider ...] (see BYOK).
The judge receives all prior evidence — validator failures, anomaly signals, inspection results — so it rules with full context, not just input/output. Every decision includes an audit trail:
{
"pass": false,
"reason": "Validator correctly identified missing resolution_ticket",
"confidence": 0.85,
"evidence_considered": ["validator:payment_check", "anomaly:BA-003"],
"overridden_signals": []
}
evidence_considered — which prior signals the LLM weighedoverridden_signals — which signals the LLM disagreed with (passed despite the flag)A signature that keeps flagging legitimate output — say NL-002 reading the string "none" as a
serialized null — should be silenced, not worked around by contorting your data:
argus ignore NL-002 # everywhere in this project
argus ignore RF-001 --node draft_hook # only on one node
argus ignore --list
argus ignore NL-002 --remove
Suppressions live in .argus/config.json (commit it so the team shares them). A suppressed hit
no longer changes the node's status or the CI gate, but is still recorded on the run
(suppressed_signals) so argus stats keeps counting it. argus doctor lists what's active.
watcher = ArgusWatcher(graph, validators={
"classify": lambda o: (o.get("label") in ["yes", "no"], "unexpected label"),
"*": lambda o: ("error" not in o, "error key present"), # runs on every node
})
Validator failures cannot be overridden by the LLM judge — they are hard constraints.
from argus import ArgusWatcher, ArgusConfig
config = ArgusConfig(
semantic_judge=True, # LLM judge on every node (default: False)
judge_model="gpt-4o", # model for the judge
node_timeout_ms=30000, # flag outputs at ≥95% of this
min_expected_ms=500, # flag suspiciously fast LLM nodes
sample_rate=0.5, # persist 50% of clean runs (save disk)
persist_failures=True, # always persist failed runs
)
watcher = ArgusWatcher(graph, config=config)
argus list # all recorded runs
argus show last # most recent run
argus show <id> # inspect a specific run
argus check <id> # CI gate for an exact run; prints the JSON path checked
ARGUS_RUN_ID=<id> argus check # CI-friendly selection when the id comes from an earlier step
argus check last # newest-file fallback — avoid in a shared workspace
argus check last --format json # same verdict as one JSON object (run_id, overall_status, passed, findings[])
argus check last --fail-on crashed,silent_failure # only these run statuses fail the gate
argus inspect <id> --step <node> # dump raw input/output for a node
argus fix <id> # fix prompt for the root cause, ready to paste
argus replay <id> <node> # re-run from a node
argus diff <id-a> <id-b> # compare two runs
argus stats # signature hit stats, disable/enable/dispute signatures
argus ignore <SIG-ID> [--node N] # silence a noisy signature project-wide or on one node
argus ignore --list # show active suppressions (.argus/config.json)
argus ui # web dashboard
argus doctor # check setup health + LLM mode (BYOK/hosted/heuristic)
argus key set [--provider ...] # save a provider key locally (OpenAI/Anthropic/Google) — BYOK
argus key use <provider> # switch the active provider
argus key show # list configured providers (masked); * marks active
argus key clear [--provider ...] # remove one provider's key, or all
argus login # (optional) sign in for hosted cloud sync
argus logout # clear stored credentials
argus whoami # show current login status
argus update # check for newer release
Silent failures become test failures without changing how you invoke the graph:
pytest --argus
ARGUS auto-wraps StateGraph.compile() / compiled invoke() for the test session. A clean pipeline stays a passing test; missing fields, tool failures, crashes, and semantic degradation fail that test. Tests that never invoke a graph are unchanged. After a standalone CI run, pass its exact id with argus check <id> or ARGUS_RUN_ID=<id> argus check; argus check last only means the newest file and can select a stale or unrelated run in a shared workspace.
argus ui # opens at localhost:7842
Shows all runs, node-level detail, AI analysis, replay diffs, loop iteration badges, and comparison views. No account needed for local use.
If the table is empty, the UI is serving a different .argus than the project that just ran, or there are no runs yet. The empty state shows the path ARGUS is reading and what to do (argus show last, run the graph, check cwd vs project root).
from argus import ArgusSession
session = ArgusSession()
session.set_edges({"fetch": ["classify"], "classify": ["process"]})
fetch = session.wrap("fetch", fetch_fn)
classify = session.wrap("classify", classify_fn)
process = session.wrap("process", process_fn)
state = fetch(initial_state)
state = classify(state)
state = process(state)
session.finalize()
Works with any framework — Prefect, Temporal, plain Python.
ArgusWatcher)argus key set [--provider ...] (optional; all heuristic detection works without it)For AI setup prompts and integration guides, visit arguslabs.in.
v0.11.0 — changelog
See CONTRIBUTING.md. Join the Discord before opening a PR for updates and to talk through the change.
ARGUS is open-core. The open-source core (src/argus/, the argus-agents PyPI
package) is licensed under Apache-2.0 — see LICENSE. The cloud/
directory (hosted/enterprise components) is proprietary — see cloud/LICENSE.
Python
59.9%
TypeScript
23.3%
HTML
15.9%