A sandbox for exploring the Caracal distributed resilience library
TypeScript
1
6 commits
updated Sep 18, 2026
A sandbox for exploring the distributed resilience library Caracal.
Caracal is a distributed resilience library for TypeScript. It wraps an asynchronous operation in a policy pipeline — timeout, retry, circuit breaker, bulkhead — and, through Redis, lets many replicas share those policies. Its defining idea is scope: you say where a limit belongs — a process, a fleet, a region, a tenant, a shard — so the constraint matches your failure domain instead of multiplying per replica.
Case studies that run the same workload against the same dependency, move the coordination boundary, and record what actually happened. Every headline number is read from the most independent instrument available - the dependency's own counters wherever they exist - and where a claim can only come from caracal's own events, the study says so and checks it against a witness caracal does not control.
| Case study | The boundary moves to... | What it measures |
|---|---|---|
01 · local-only | the process | 4 × bulkhead.local({ limit: 5 }) → the dependency's peak in-flight reaches 20, not 5 |
02 · distributed | the fleet | one shared budget → peak in-flight held at 5 (refused 83 → 4,816) |
03 · multi-region | the region | break eu's dependency → eu's breaker opens, us: 0 opens |
04 · tenant isolation | the tenant | 50 tenants, one noisy → exactly 1 breaker opens |
05 · postgres permit-hold | the query | a 2s pg_sleep behind a 500ms timeout holds its permit ~2,059ms |
06 · observability | the sink | a blocking sink triples p99; a throwing/slow-async one leaves it unchanged |
07 · chaos | the lease | kill a replica, freeze another → live leases never observed above 3, return to 0 |
08 · overload | the ceiling | a budget of 20 against a capacity-8 dependency → the dependency itself 503s, not caracal |
09 · local bulkhead lease | the permit | a hung bulkhead.local holder is aborted after its leaseMs, so its permit is released |
10 · dispose abandoned body | the response | retry abandons a 5xx body and the fetch adapter's dispose cancels it |
11 · breaker probe lease | the probe | a hung half-open probe releases its slot after probeLeaseTtlMs |
12 · one shared rate budget | the fleet | one shared rate budget → the dependency sustains ~100 req/s, not ~400 |
13 · rate vs concurrency | the axis | rate 100/s at 200ms latency → ~20 in flight (a rate limit is not a concurrency ceiling) |
14 · burst and retry-after | the knob | burst: 20 admits a cluster, then 50/s sheds the surplus with a retry-after hint |
15 · per-tenant rate quota | the tenant | 50 tenants, one noisy → the noisy tenant's rate is capped, the other 49 unaffected |
Each number above is read from the most independent instrument available: the dependency's own counters (01, 02, 08), caracal's events for breaker and permit claims (03–05), the client's own timings (06), and Redis's own lease count (07). Three of those instruments are not caracal at all: the dependency's counters, the client's clocks, and Redis's own lease state. Where a claim can only be reported by caracal - which breaker opened, how many permits are live - the study says so rather than implying it was independently measured.
Requires Node ≥ 20.3 and Docker. The workloads run on the host; only the stack (Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana) runs in Docker.
npm install # @gkoos/caracal + the workspace packages
npm run verify:api # the installed caracal surface matches what the sandbox expects
npm run stack:up # Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana
npm run smoke # the whole chain: pipeline → events → OTLP → dashboards
Then open Grafana at http://localhost:3000/d/caracal-overview and Jaeger at http://localhost:16686. Tear the stack down with npm run stack:down.
Run any single case study with npm run demo <id> (each study's README has the exact command), and compare two runs side by side:
npm run compare 01 02
Make sure to run the demos before comparing, because the comparison is based on the summary.json files in runs/, generated by the demos.
More things to run:
| Command | What it proves |
|---|---|
npm run smoke | local policies, healthy dependency, everything observable |
npm run smoke:distributed | same workload, one shared budget in Redis |
npm run smoke:breaker | the breaker opens and sheds traffic (retry disabled so the signal isn't diluted) |
npm run smoke:overload | the bulkhead limit is below the offered concurrency: the surplus is shed as capacity and the breaker stays closed |
npm run observability | the sink contract: a throwing/slow-async sink leaves p99 unchanged, a blocking one does not |
CARACAL_OTEL=off npm run smoke | the event → summary → compare path with no SDK and no stack |
npm run stack:up provisions a Grafana dashboard — Caracal overview, the default home at http://localhost:3000. The witness is the first panel: the dependency's own concurrency count, the number the headline claims are decided on. Below it, the caracal panels show executions, attempts, bulkhead occupancy, breaker state, and coordination cost.
To compare two studies, run both (npm run demo 01 then npm run demo 02), then multi-select them in the Demo variable — the caracal panels overlay them, color-coded by demo. Widen the time range (it defaults to the last 5 minutes) so the runs are in view.
The witness panel is live only — it shows whatever is running now, not a per-run history — so the headline 20 vs 5 still comes from npm run compare 01 02.
packages/caracal-observability is consumer code, written the way you'd write it.abort: "unsupported", so a timed-out query keeps its permit until it settles. Study 05 explains this.07 shows where that boundary actually is.Building the sandbox surfaced a few non-obvious traps: a dead collector that fails only at shutdown, metric instruments that are silently permanent no-ops, a saturated bulkhead that tripped its own outer breaker (fixed in caracal 0.5.0), and seven more. Each is recorded with the evidence that produced it in docs/findings.md.
apps/partner-api/ # THE workload. Topology is configuration, not code
packages/caracal-observability # caracal events → OTLP traces + metrics (+ ndjson)
packages/caracal-runner/ # summary schema, named checks, compare/report
stack/ # valkey, valkey-eu, postgres, otel, jaeger, prometheus, grafana
case-studies/ # one case study per folder: study.yaml + README
runs/ # artifacts (gitignored): summary.json, report.md, events.ndjson
scripts/ # stack, smoke, demo, diagrams, verify-api, link-local, diagnose-otel
Every case study has a diagram.svg drawn from one visual language: replicas in terracotta, the dependency in steel blue, coordinators in brick red, the witness as a dashed emerald line, chaos as a red bolt. npm run diagrams regenerates them all from scripts/diagrams.mjs, so they stay consistent by construction.
The case studies run against the published package. To test unreleased changes:
npm run link:local # packs ../caracal and installs the tarball
npm run link:npm # back to @gkoos/caracal@0.6.0
npm run verify:api fails if the surface drifts from what the sandbox uses.
npm run compare 01 02 is the whole argument in one table: two runs of the same workload against the same dependency, differing only in where the limit lives.
KPI 01-local-only 02-distributed-basic verdict
witness peak in-flight 20 5 02-distributed-basic (lower is better)
successful requests/s ~392 ~108 01-local-only (higher is better)
rejected requests/s ~5 ~324 01-local-only (lower is better)
witness peak / limit 20 / 5 5 / 5 -
coordinator trips/exec n/a ~4 n/a
The verdict column looks like 01 wins on the rate rows. It doesn't, read it this way:
capacity at the client instead of being pushed onto the dependency. Shedding is the budget working, not failing.The rows that read same (breaker opens, peak probes, lease lost) are the control: the dependency stayed healthy in both runs, so the only thing that moved is the coordination boundary.
Capture a study's result as its committed baseline with npm run baseline <id>, then drift-check it with node --import tsx packages/caracal-runner/bin/compare.mjs <id> --against-baseline — it fails if any KPI moved more than 10%.
6 commits
TypeScript
66.4%
JavaScript
32.9%
A sandbox for exploring the Caracal distributed resilience library
TypeScript
1
6 commits
updated Sep 18, 2026
A sandbox for exploring the distributed resilience library Caracal.
Caracal is a distributed resilience library for TypeScript. It wraps an asynchronous operation in a policy pipeline — timeout, retry, circuit breaker, bulkhead — and, through Redis, lets many replicas share those policies. Its defining idea is scope: you say where a limit belongs — a process, a fleet, a region, a tenant, a shard — so the constraint matches your failure domain instead of multiplying per replica.
Case studies that run the same workload against the same dependency, move the coordination boundary, and record what actually happened. Every headline number is read from the most independent instrument available - the dependency's own counters wherever they exist - and where a claim can only come from caracal's own events, the study says so and checks it against a witness caracal does not control.
| Case study | The boundary moves to... | What it measures |
|---|---|---|
01 · local-only | the process | 4 × bulkhead.local({ limit: 5 }) → the dependency's peak in-flight reaches 20, not 5 |
02 · distributed | the fleet | one shared budget → peak in-flight held at 5 (refused 83 → 4,816) |
03 · multi-region | the region | break eu's dependency → eu's breaker opens, us: 0 opens |
04 · tenant isolation | the tenant | 50 tenants, one noisy → exactly 1 breaker opens |
05 · postgres permit-hold | the query | a 2s pg_sleep behind a 500ms timeout holds its permit ~2,059ms |
06 · observability | the sink | a blocking sink triples p99; a throwing/slow-async one leaves it unchanged |
07 · chaos | the lease | kill a replica, freeze another → live leases never observed above 3, return to 0 |
08 · overload | the ceiling | a budget of 20 against a capacity-8 dependency → the dependency itself 503s, not caracal |
09 · local bulkhead lease | the permit | a hung bulkhead.local holder is aborted after its leaseMs, so its permit is released |
10 · dispose abandoned body | the response | retry abandons a 5xx body and the fetch adapter's dispose cancels it |
11 · breaker probe lease | the probe | a hung half-open probe releases its slot after probeLeaseTtlMs |
12 · one shared rate budget | the fleet | one shared rate budget → the dependency sustains ~100 req/s, not ~400 |
13 · rate vs concurrency | the axis | rate 100/s at 200ms latency → ~20 in flight (a rate limit is not a concurrency ceiling) |
14 · burst and retry-after | the knob | burst: 20 admits a cluster, then 50/s sheds the surplus with a retry-after hint |
15 · per-tenant rate quota | the tenant | 50 tenants, one noisy → the noisy tenant's rate is capped, the other 49 unaffected |
Each number above is read from the most independent instrument available: the dependency's own counters (01, 02, 08), caracal's events for breaker and permit claims (03–05), the client's own timings (06), and Redis's own lease count (07). Three of those instruments are not caracal at all: the dependency's counters, the client's clocks, and Redis's own lease state. Where a claim can only be reported by caracal - which breaker opened, how many permits are live - the study says so rather than implying it was independently measured.
Requires Node ≥ 20.3 and Docker. The workloads run on the host; only the stack (Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana) runs in Docker.
npm install # @gkoos/caracal + the workspace packages
npm run verify:api # the installed caracal surface matches what the sandbox expects
npm run stack:up # Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana
npm run smoke # the whole chain: pipeline → events → OTLP → dashboards
Then open Grafana at http://localhost:3000/d/caracal-overview and Jaeger at http://localhost:16686. Tear the stack down with npm run stack:down.
Run any single case study with npm run demo <id> (each study's README has the exact command), and compare two runs side by side:
npm run compare 01 02
Make sure to run the demos before comparing, because the comparison is based on the summary.json files in runs/, generated by the demos.
More things to run:
| Command | What it proves |
|---|---|
npm run smoke | local policies, healthy dependency, everything observable |
npm run smoke:distributed | same workload, one shared budget in Redis |
npm run smoke:breaker | the breaker opens and sheds traffic (retry disabled so the signal isn't diluted) |
npm run smoke:overload | the bulkhead limit is below the offered concurrency: the surplus is shed as capacity and the breaker stays closed |
npm run observability | the sink contract: a throwing/slow-async sink leaves p99 unchanged, a blocking one does not |
CARACAL_OTEL=off npm run smoke | the event → summary → compare path with no SDK and no stack |
npm run stack:up provisions a Grafana dashboard — Caracal overview, the default home at http://localhost:3000. The witness is the first panel: the dependency's own concurrency count, the number the headline claims are decided on. Below it, the caracal panels show executions, attempts, bulkhead occupancy, breaker state, and coordination cost.
To compare two studies, run both (npm run demo 01 then npm run demo 02), then multi-select them in the Demo variable — the caracal panels overlay them, color-coded by demo. Widen the time range (it defaults to the last 5 minutes) so the runs are in view.
The witness panel is live only — it shows whatever is running now, not a per-run history — so the headline 20 vs 5 still comes from npm run compare 01 02.
packages/caracal-observability is consumer code, written the way you'd write it.abort: "unsupported", so a timed-out query keeps its permit until it settles. Study 05 explains this.07 shows where that boundary actually is.Building the sandbox surfaced a few non-obvious traps: a dead collector that fails only at shutdown, metric instruments that are silently permanent no-ops, a saturated bulkhead that tripped its own outer breaker (fixed in caracal 0.5.0), and seven more. Each is recorded with the evidence that produced it in docs/findings.md.
apps/partner-api/ # THE workload. Topology is configuration, not code
packages/caracal-observability # caracal events → OTLP traces + metrics (+ ndjson)
packages/caracal-runner/ # summary schema, named checks, compare/report
stack/ # valkey, valkey-eu, postgres, otel, jaeger, prometheus, grafana
case-studies/ # one case study per folder: study.yaml + README
runs/ # artifacts (gitignored): summary.json, report.md, events.ndjson
scripts/ # stack, smoke, demo, diagrams, verify-api, link-local, diagnose-otel
Every case study has a diagram.svg drawn from one visual language: replicas in terracotta, the dependency in steel blue, coordinators in brick red, the witness as a dashed emerald line, chaos as a red bolt. npm run diagrams regenerates them all from scripts/diagrams.mjs, so they stay consistent by construction.
The case studies run against the published package. To test unreleased changes:
npm run link:local # packs ../caracal and installs the tarball
npm run link:npm # back to @gkoos/caracal@0.6.0
npm run verify:api fails if the surface drifts from what the sandbox uses.
npm run compare 01 02 is the whole argument in one table: two runs of the same workload against the same dependency, differing only in where the limit lives.
KPI 01-local-only 02-distributed-basic verdict
witness peak in-flight 20 5 02-distributed-basic (lower is better)
successful requests/s ~392 ~108 01-local-only (higher is better)
rejected requests/s ~5 ~324 01-local-only (lower is better)
witness peak / limit 20 / 5 5 / 5 -
coordinator trips/exec n/a ~4 n/a
The verdict column looks like 01 wins on the rate rows. It doesn't, read it this way:
capacity at the client instead of being pushed onto the dependency. Shedding is the budget working, not failing.The rows that read same (breaker opens, peak probes, lease lost) are the control: the dependency stayed healthy in both runs, so the only thing that moved is the coordination boundary.
Capture a study's result as its committed baseline with npm run baseline <id>, then drift-check it with node --import tsx packages/caracal-runner/bin/compare.mjs <id> --against-baseline — it fails if any KPI moved more than 10%.
6 commits
TypeScript
66.4%
JavaScript
32.9%