ilagutin/IncidentCompass

Governed AI incident-triage agent backend - reference implementation (.NET 10, PostgreSQL/pgvector, local-LLM friendly)

C#

0

143 commits

updated Oct 5, 2026

See the code

README

IncidentCompass

CI Latest release License

A .NET 10 backend that turns an incident signal into a reviewable, evidence-backed triage report.

The model chooses investigation steps and proposes tool calls. The backend owns credentials, tool grants, budgets, evidence checks, approvals and external actions.

Version 0.5.0 runs memory embeddings on an in-process multilingual model the Worker installs and verifies itself, admits memory matches through an in-process multilingual cross-encoder that judges whether a chunk answers the query, chunks memory documents by section, bounds each provider call by connect, first-output and stream-inactivity limits under a four-hour attempt ceiling, and streams chat completions by default. The first Worker start downloads both models unless they are placed offline, two configuration keys are deprecated, and an existing corpus keeps its whole-file chunks until memory rebuild; read the notes before upgrading. Earlier releases: 0.4.1, 0.4.0.

One investigation

  1. Accept an OTLP trace/log export or an OTel-shaped, user or tester signal; redact it and create a durable triage job.
  2. Investigate under the job's versioned configuration, using role-scoped memory, source and ticket tools.
  3. Publish a report only after the backend resolves its evidence against allowed stored artifacts.
  4. When configured, propose a Telegram notification or GitHub issue action. The backend selects the target and freezes the payload; GitHub writes require operator approval.
  5. When configured, prepare a diff against a named base tree, then freeze it into a code write, a branch, a pull request and a comment back on the cited ticket. Each is a separate approval a person grants, nothing merges, and no test command runs, so a diff is untested by construction.
  6. Inspect the report, policy decisions, model usage and action outcomes through the API and durable ledger.
flowchart TD
  A["Signal: OTLP export, user report or tester"] --> B["Redact secrets, pseudonymize user ids"]
  B --> C["Fingerprint: strong needs a real service name and errorType"]
  C --> D{"Fault grouping"}
  D -->|"attach or suppress"| E["Existing fault, no new job"]
  D -->|"new fault"| F["Pending triage job pinned to a config hash"]
  F --> G["Worker claims the job and rehydrates that exact config"]
  G --> H["Orchestrator turn: delegate or publish_report"]
  H -->|"delegate role and task"| I["Scoped worker role, only its granted tools"]
  I -->|"proposes a tool call"| K{"ToolRuleEngine decides"}
  K -->|"denied"| X["Attempt fails closed, no report published"]
  K -->|"allowed"| M["Backend executes the tool and stores artifacts"]
  M --> I
  I -->|"validated output stored as an artifact"| H
  H -->|"publish_report"| N{"Evidence resolves against this attempt's artifacts?"}
  N -->|"no"| R["Bounded reprompt"]
  R -->|"corrected"| H
  R -->|"allowance spent"| X
  N -->|"yes"| O["Report, evidence and job/fault state commit in one transaction"]
  H -.-> L[("Triage ledger: ModelCall, PolicyDecision, BudgetEvent")]
  K -.-> L

Deterministic demo run

scripts/demo.ps1 -Mock drives five scenarios end to end. This is the deterministic mock provider, not a real model, so the classifications below show the governed path running, not model quality. For numbers produced by an actual model, see Measured run below; the two are separate records and neither stands in for the other.

Scenario | FaultId | ReportId | is_mass_issue | Classification | LedgerUrl | ReportUrl | Check
--- | --- | --- | --- | --- | --- | --- | ---
1 known-timeout-runbook | db50a885-0e75-4726-9825-6c82085a37bf | 8e34e889-39e2-4f7b-8e6a-d838febf75db | false | KnownIncident | http://localhost:5198/api/v1/faults/db50a885-0e75-4726-9825-6c82085a37bf/ledger | http://localhost:5198/api/v1/triage-reports/8e34e889-39e2-4f7b-8e6a-d838febf75db | ok
2 unknown-null-reference | 9a56b670-bfc8-4e64-98dc-debf8fb8ea55 | a483f009-adc0-4f03-a7f6-9cd204e04b18 | false | Unknown | http://localhost:5198/api/v1/faults/9a56b670-bfc8-4e64-98dc-debf8fb8ea55/ledger | http://localhost:5198/api/v1/triage-reports/a483f009-adc0-4f03-a7f6-9cd204e04b18 | ok
3 provider-unavailable-flood | 144467dd-ba0f-4627-ac4d-fc033ea0a55c | 542e060d-e45e-4072-b2d2-fef1b3fd5096 | true | SimpleKnownError | http://localhost:5198/api/v1/faults/144467dd-ba0f-4627-ac4d-fc033ea0a55c/ledger | http://localhost:5198/api/v1/triage-reports/542e060d-e45e-4072-b2d2-fef1b3fd5096 | ok
4 validation-noise | 7c95878e-c56f-41db-808a-5430a15d54af | c4ab8ec5-628f-42ed-bdf4-6944b295cd2b | false | Noise | http://localhost:5198/api/v1/faults/7c95878e-c56f-41db-808a-5430a15d54af/ledger | http://localhost:5198/api/v1/triage-reports/c4ab8ec5-628f-42ed-bdf4-6944b295cd2b | ok
5 injection-disabled-action-gate | 4f705140-7ff5-4d7f-ad08-e8bc90e931b2 | e57a7035-579d-4f75-8684-d6eed5e20889 | false | SimpleKnownError | http://localhost:5198/api/v1/faults/4f705140-7ff5-4d7f-ad08-e8bc90e931b2/ledger | http://localhost:5198/api/v1/triage-reports/e57a7035-579d-4f75-8684-d6eed5e20889 | no action lifecycle events observed across 4 bounded ledger reads

All five scenarios passed. Scenario 5 prints a detail string instead of ok because it aims a prompt-injection attempt at a disabled action gate, and that detail is the assertion: DemoActionGateResult passes only when no action lifecycle event was recorded at all.

Read that row narrowly, and read the measured run's side-effect row below the same way. An observation that nothing was recorded under a disabled configuration shows the shipped default staying off while a hostile prompt is in flight. It does not show that the approval machinery is correct, because no approval was ever reached. Measured evaluation run states that caveat against its own fifteen negative observations and points at the deterministic integration test where the action chain is actually proved.

Measured run

Separately from that mock table, one opt-in evaluation ran against a local OpenAI-compatible provider on 2026-09-11, against the exact working tree 4a7feace870c4897fcfe60efd5eafcc6b1876768. That is a git tree hash rather than a commit id, because releases here land by rebase: rebasing rewrites commit ids and leaves tree hashes alone, so the tree is what a reader can still resolve inside the v0.4.0 tag. The change committed immediately after the run, and shipped in the same release, added integration tests and changed no production code. Five frozen cases, three attempts each, criteria authored before any output was observed. The per-attempt record is committed at evaluations/triage/measured-run-tree-4a7feac.json, and Measured evaluation run shows how to turn that tree hash back into a checkout.

BandMeasured
Delivery: the signal became a job that reached the provider15 of 15 attempts, 192 model calls, all usage provider-reported
Terminal completion: the job ended with a schema-valid published report15 of 15, no job error codes
Diagnostic quality: the report matched the pre-authored criteriadiagnosis 13 of 15, evidence 14 of 15, justified refusal 14 of 15
Side effects: anything written outside the systemnone, on every attempt, under empty action grants

Nine of the fifteen published reports were InsufficientEvidence, so completion does not mean the report answered anything. This is five cases on one model and one machine, with authored keyword heuristics standing in for diagnosis quality, and an earlier run of the same corpus that same day finished only 10 of 15. No external action was proposed or executed in this run.

Measured evaluation run carries the per-case numbers, the two misses attempt by attempt, the variance, what the metrics are not, and where the external-action chain is actually proved.

Try the local flow

Prerequisites: Docker Compose, the .NET 10 SDK for configuration validation, PowerShell and an OpenAI-compatible chat endpoint. On Windows and macOS, Compose uses host.docker.internal:1234 by default. The Worker runs two in-process models and downloads both from Hugging Face on its first start: the embedding model, about 123 MB, and the relevance judge memory search admits matches with, about 544 MiB.

Copy-Item .env.example .env
# Set the exact chat model id in .env.
dotnet run --project src/IncidentCompass.Api -- config validate
powershell -ExecutionPolicy Bypass -File scripts/demo.ps1

The script builds PostgreSQL, API, Worker and Tester containers, runs the local scenarios and prints report and ledger URLs. Use scripts/demo.ps1 -Mock for the deterministic provider path.

The additive production path is intentionally separate from that demo. See the single-host production runbook for loopback-only Compose startup, preflight, bounded PostgreSQL backup and fresh-volume recovery on one trusted machine.

Compose host-port overrides do not change the fixed internal API, PostgreSQL or OTLP addresses. The mock overlay replaces only model and embedding providers. GitHub and Telegram adapters retain fixed production authorities and are tested with in-process recording handlers, not Compose endpoint doubles or configurable provider URLs.

See Quickstart for setup and Local demo walkthrough for scenarios, ports and expected output. Optional source, GitHub, Telegram and API-key settings are in Integration configuration.

Guarantees and limits

  • Unknown, ungranted or denied tools fail closed through backend policy.
  • Reports cite stored evidence from the allowed job and attempt. Valid citations do not prove a correct diagnosis.
  • Backend budgets bound investigation work. A single in-flight model call can overshoot the token limit; the excess is recorded and further calls stop.
  • Approved actions use frozen payloads and at-most-once backend invocation. An uncertain external outcome is recorded for operator review and is never automatically resent.
  • Automated tests use deterministic providers by default. The mock injection scenario checks the disabled-action configuration; separate integration tests exercise configured policy and approval.
  • There is no prompt or provider-body logging switch, because no code path writes that material anywhere. Rendered prompts, provider request and response bodies, document text, credentials, embedding vectors and relevance-judge scores reach neither the application logs nor the triage ledger, at any log level.

One deployment envelope is supported: one trusted machine, one trusted operator, one host-owned monitored checkout, one configured repository and one PostgreSQL database under Docker Compose, with the API and PostgreSQL bound to loopback. The single-host production runbook is that envelope, with preflight that refuses demo credentials, bounded backup, a real fresh-volume restore and rollback. Reaching the API from another machine means putting an authenticated TLS reverse proxy and a host firewall in front of the loopback port, not rebinding it. Demo authentication and the demo Compose defaults are local-only. Source lookup and external integrations require explicit host configuration, and every external action ships disabled.

Outside that envelope, and not provided: high availability or failover, multi-team role-based access, enterprise identity, a managed secret store (the host owns a protected environment file), more than one source repository or an arbitrary one, a UI, general cross-fault incident correlation, and exactly-once external delivery. No process is started anywhere in the product, so nothing here runs a test, a build or a deployment: a prepared diff is untested by construction and approving one approves an untested change. And a grounded report is not a correct report; evidence checks prove provenance, not truth.

Explore the implementation

IncidentCompass is a layered monolith: API and Worker compose Application use cases and Infrastructure adapters around a provider-independent Domain. PostgreSQL stores jobs, configuration snapshots, incident memory, reports, approvals and the audit ledger.

Read Architecture, Security model and Trade-offs for the boundaries and their costs.

Documentation

The documentation index states a reading order in three tiers: run it, understand it, judge it. The direct routes are:

Origin

IncidentCompass was bootstrapped from dotnet-genai-starter's layered .NET structure and model/embedding gateway boundaries, then specialized into incident triage. General chat, document ingestion and MCP product surfaces are outside this repository's scope.

ai-agents
csharp
dotnet
governance
incident-response
llm
local-llm
pgvector
rag

ilagutin/IncidentCompass

Governed AI incident-triage agent backend - reference implementation (.NET 10, PostgreSQL/pgvector, local-LLM friendly)

C#

0

143 commits

updated Oct 5, 2026

See the code

README

IncidentCompass

CI Latest release License

A .NET 10 backend that turns an incident signal into a reviewable, evidence-backed triage report.

The model chooses investigation steps and proposes tool calls. The backend owns credentials, tool grants, budgets, evidence checks, approvals and external actions.

Version 0.5.0 runs memory embeddings on an in-process multilingual model the Worker installs and verifies itself, admits memory matches through an in-process multilingual cross-encoder that judges whether a chunk answers the query, chunks memory documents by section, bounds each provider call by connect, first-output and stream-inactivity limits under a four-hour attempt ceiling, and streams chat completions by default. The first Worker start downloads both models unless they are placed offline, two configuration keys are deprecated, and an existing corpus keeps its whole-file chunks until memory rebuild; read the notes before upgrading. Earlier releases: 0.4.1, 0.4.0.

One investigation

  1. Accept an OTLP trace/log export or an OTel-shaped, user or tester signal; redact it and create a durable triage job.
  2. Investigate under the job's versioned configuration, using role-scoped memory, source and ticket tools.
  3. Publish a report only after the backend resolves its evidence against allowed stored artifacts.
  4. When configured, propose a Telegram notification or GitHub issue action. The backend selects the target and freezes the payload; GitHub writes require operator approval.
  5. When configured, prepare a diff against a named base tree, then freeze it into a code write, a branch, a pull request and a comment back on the cited ticket. Each is a separate approval a person grants, nothing merges, and no test command runs, so a diff is untested by construction.
  6. Inspect the report, policy decisions, model usage and action outcomes through the API and durable ledger.
flowchart TD
  A["Signal: OTLP export, user report or tester"] --> B["Redact secrets, pseudonymize user ids"]
  B --> C["Fingerprint: strong needs a real service name and errorType"]
  C --> D{"Fault grouping"}
  D -->|"attach or suppress"| E["Existing fault, no new job"]
  D -->|"new fault"| F["Pending triage job pinned to a config hash"]
  F --> G["Worker claims the job and rehydrates that exact config"]
  G --> H["Orchestrator turn: delegate or publish_report"]
  H -->|"delegate role and task"| I["Scoped worker role, only its granted tools"]
  I -->|"proposes a tool call"| K{"ToolRuleEngine decides"}
  K -->|"denied"| X["Attempt fails closed, no report published"]
  K -->|"allowed"| M["Backend executes the tool and stores artifacts"]
  M --> I
  I -->|"validated output stored as an artifact"| H
  H -->|"publish_report"| N{"Evidence resolves against this attempt's artifacts?"}
  N -->|"no"| R["Bounded reprompt"]
  R -->|"corrected"| H
  R -->|"allowance spent"| X
  N -->|"yes"| O["Report, evidence and job/fault state commit in one transaction"]
  H -.-> L[("Triage ledger: ModelCall, PolicyDecision, BudgetEvent")]
  K -.-> L

Deterministic demo run

scripts/demo.ps1 -Mock drives five scenarios end to end. This is the deterministic mock provider, not a real model, so the classifications below show the governed path running, not model quality. For numbers produced by an actual model, see Measured run below; the two are separate records and neither stands in for the other.

Scenario | FaultId | ReportId | is_mass_issue | Classification | LedgerUrl | ReportUrl | Check
--- | --- | --- | --- | --- | --- | --- | ---
1 known-timeout-runbook | db50a885-0e75-4726-9825-6c82085a37bf | 8e34e889-39e2-4f7b-8e6a-d838febf75db | false | KnownIncident | http://localhost:5198/api/v1/faults/db50a885-0e75-4726-9825-6c82085a37bf/ledger | http://localhost:5198/api/v1/triage-reports/8e34e889-39e2-4f7b-8e6a-d838febf75db | ok
2 unknown-null-reference | 9a56b670-bfc8-4e64-98dc-debf8fb8ea55 | a483f009-adc0-4f03-a7f6-9cd204e04b18 | false | Unknown | http://localhost:5198/api/v1/faults/9a56b670-bfc8-4e64-98dc-debf8fb8ea55/ledger | http://localhost:5198/api/v1/triage-reports/a483f009-adc0-4f03-a7f6-9cd204e04b18 | ok
3 provider-unavailable-flood | 144467dd-ba0f-4627-ac4d-fc033ea0a55c | 542e060d-e45e-4072-b2d2-fef1b3fd5096 | true | SimpleKnownError | http://localhost:5198/api/v1/faults/144467dd-ba0f-4627-ac4d-fc033ea0a55c/ledger | http://localhost:5198/api/v1/triage-reports/542e060d-e45e-4072-b2d2-fef1b3fd5096 | ok
4 validation-noise | 7c95878e-c56f-41db-808a-5430a15d54af | c4ab8ec5-628f-42ed-bdf4-6944b295cd2b | false | Noise | http://localhost:5198/api/v1/faults/7c95878e-c56f-41db-808a-5430a15d54af/ledger | http://localhost:5198/api/v1/triage-reports/c4ab8ec5-628f-42ed-bdf4-6944b295cd2b | ok
5 injection-disabled-action-gate | 4f705140-7ff5-4d7f-ad08-e8bc90e931b2 | e57a7035-579d-4f75-8684-d6eed5e20889 | false | SimpleKnownError | http://localhost:5198/api/v1/faults/4f705140-7ff5-4d7f-ad08-e8bc90e931b2/ledger | http://localhost:5198/api/v1/triage-reports/e57a7035-579d-4f75-8684-d6eed5e20889 | no action lifecycle events observed across 4 bounded ledger reads

All five scenarios passed. Scenario 5 prints a detail string instead of ok because it aims a prompt-injection attempt at a disabled action gate, and that detail is the assertion: DemoActionGateResult passes only when no action lifecycle event was recorded at all.

Read that row narrowly, and read the measured run's side-effect row below the same way. An observation that nothing was recorded under a disabled configuration shows the shipped default staying off while a hostile prompt is in flight. It does not show that the approval machinery is correct, because no approval was ever reached. Measured evaluation run states that caveat against its own fifteen negative observations and points at the deterministic integration test where the action chain is actually proved.

Measured run

Separately from that mock table, one opt-in evaluation ran against a local OpenAI-compatible provider on 2026-09-11, against the exact working tree 4a7feace870c4897fcfe60efd5eafcc6b1876768. That is a git tree hash rather than a commit id, because releases here land by rebase: rebasing rewrites commit ids and leaves tree hashes alone, so the tree is what a reader can still resolve inside the v0.4.0 tag. The change committed immediately after the run, and shipped in the same release, added integration tests and changed no production code. Five frozen cases, three attempts each, criteria authored before any output was observed. The per-attempt record is committed at evaluations/triage/measured-run-tree-4a7feac.json, and Measured evaluation run shows how to turn that tree hash back into a checkout.

BandMeasured
Delivery: the signal became a job that reached the provider15 of 15 attempts, 192 model calls, all usage provider-reported
Terminal completion: the job ended with a schema-valid published report15 of 15, no job error codes
Diagnostic quality: the report matched the pre-authored criteriadiagnosis 13 of 15, evidence 14 of 15, justified refusal 14 of 15
Side effects: anything written outside the systemnone, on every attempt, under empty action grants

Nine of the fifteen published reports were InsufficientEvidence, so completion does not mean the report answered anything. This is five cases on one model and one machine, with authored keyword heuristics standing in for diagnosis quality, and an earlier run of the same corpus that same day finished only 10 of 15. No external action was proposed or executed in this run.

Measured evaluation run carries the per-case numbers, the two misses attempt by attempt, the variance, what the metrics are not, and where the external-action chain is actually proved.

Try the local flow

Prerequisites: Docker Compose, the .NET 10 SDK for configuration validation, PowerShell and an OpenAI-compatible chat endpoint. On Windows and macOS, Compose uses host.docker.internal:1234 by default. The Worker runs two in-process models and downloads both from Hugging Face on its first start: the embedding model, about 123 MB, and the relevance judge memory search admits matches with, about 544 MiB.

Copy-Item .env.example .env
# Set the exact chat model id in .env.
dotnet run --project src/IncidentCompass.Api -- config validate
powershell -ExecutionPolicy Bypass -File scripts/demo.ps1

The script builds PostgreSQL, API, Worker and Tester containers, runs the local scenarios and prints report and ledger URLs. Use scripts/demo.ps1 -Mock for the deterministic provider path.

The additive production path is intentionally separate from that demo. See the single-host production runbook for loopback-only Compose startup, preflight, bounded PostgreSQL backup and fresh-volume recovery on one trusted machine.

Compose host-port overrides do not change the fixed internal API, PostgreSQL or OTLP addresses. The mock overlay replaces only model and embedding providers. GitHub and Telegram adapters retain fixed production authorities and are tested with in-process recording handlers, not Compose endpoint doubles or configurable provider URLs.

See Quickstart for setup and Local demo walkthrough for scenarios, ports and expected output. Optional source, GitHub, Telegram and API-key settings are in Integration configuration.

Guarantees and limits

  • Unknown, ungranted or denied tools fail closed through backend policy.
  • Reports cite stored evidence from the allowed job and attempt. Valid citations do not prove a correct diagnosis.
  • Backend budgets bound investigation work. A single in-flight model call can overshoot the token limit; the excess is recorded and further calls stop.
  • Approved actions use frozen payloads and at-most-once backend invocation. An uncertain external outcome is recorded for operator review and is never automatically resent.
  • Automated tests use deterministic providers by default. The mock injection scenario checks the disabled-action configuration; separate integration tests exercise configured policy and approval.
  • There is no prompt or provider-body logging switch, because no code path writes that material anywhere. Rendered prompts, provider request and response bodies, document text, credentials, embedding vectors and relevance-judge scores reach neither the application logs nor the triage ledger, at any log level.

One deployment envelope is supported: one trusted machine, one trusted operator, one host-owned monitored checkout, one configured repository and one PostgreSQL database under Docker Compose, with the API and PostgreSQL bound to loopback. The single-host production runbook is that envelope, with preflight that refuses demo credentials, bounded backup, a real fresh-volume restore and rollback. Reaching the API from another machine means putting an authenticated TLS reverse proxy and a host firewall in front of the loopback port, not rebinding it. Demo authentication and the demo Compose defaults are local-only. Source lookup and external integrations require explicit host configuration, and every external action ships disabled.

Outside that envelope, and not provided: high availability or failover, multi-team role-based access, enterprise identity, a managed secret store (the host owns a protected environment file), more than one source repository or an arbitrary one, a UI, general cross-fault incident correlation, and exactly-once external delivery. No process is started anywhere in the product, so nothing here runs a test, a build or a deployment: a prepared diff is untested by construction and approving one approves an untested change. And a grounded report is not a correct report; evidence checks prove provenance, not truth.

Explore the implementation

IncidentCompass is a layered monolith: API and Worker compose Application use cases and Infrastructure adapters around a provider-independent Domain. PostgreSQL stores jobs, configuration snapshots, incident memory, reports, approvals and the audit ledger.

Read Architecture, Security model and Trade-offs for the boundaries and their costs.

Documentation

The documentation index states a reading order in three tiers: run it, understand it, judge it. The direct routes are:

Origin

IncidentCompass was bootstrapped from dotnet-genai-starter's layered .NET structure and model/embedding gateway boundaries, then specialized into incident triage. General chat, document ingestion and MCP product surfaces are outside this repository's scope.

ai-agents
csharp
dotnet
governance
incident-response
llm
local-llm
pgvector
rag