gkoos/caracal-sandbox

A sandbox for exploring the Caracal distributed resilience library

TypeScript

1

6 commits

updated Sep 18, 2026

See the code

See what people are saying

SourceMessageScoreDate

Caracal: Scoped distributed resilience policies for asynchronous operations (r/webdev)

Scoped distributed resilience policies for asynchronous operations. Most resilience libraries stop at the process boundary. You get a local circuit breaker, a local bulkhead, a local rate limit. Spin up 40 replicas and those limits multiply: a “concurrency limit of 20” quietly becomes 800…

2

Sep 27, 2026

README

caracal-sandbox

A sandbox for exploring the distributed resilience library Caracal.

What is caracal?

Caracal is a distributed resilience library for TypeScript. It wraps an asynchronous operation in a policy pipeline — timeout, retry, circuit breaker, bulkhead — and, through Redis, lets many replicas share those policies. Its defining idea is scope: you say where a limit belongs — a process, a fleet, a region, a tenant, a shard — so the constraint matches your failure domain instead of multiplying per replica.

What is this sandbox?

Case studies that run the same workload against the same dependency, move the coordination boundary, and record what actually happened. Every headline number is read from the most independent instrument available - the dependency's own counters wherever they exist - and where a claim can only come from caracal's own events, the study says so and checks it against a witness caracal does not control.

The case studies, at a glance

Case studyThe boundary moves to...What it measures
01 · local-onlythe process4 × bulkhead.local({ limit: 5 }) → the dependency's peak in-flight reaches 20, not 5
02 · distributedthe fleetone shared budget → peak in-flight held at 5 (refused 83 → 4,816)
03 · multi-regionthe regionbreak eu's dependency → eu's breaker opens, us: 0 opens
04 · tenant isolationthe tenant50 tenants, one noisy → exactly 1 breaker opens
05 · postgres permit-holdthe querya 2s pg_sleep behind a 500ms timeout holds its permit ~2,059ms
06 · observabilitythe sinka blocking sink triples p99; a throwing/slow-async one leaves it unchanged
07 · chaosthe leasekill a replica, freeze another → live leases never observed above 3, return to 0
08 · overloadthe ceilinga budget of 20 against a capacity-8 dependency → the dependency itself 503s, not caracal
09 · local bulkhead leasethe permita hung bulkhead.local holder is aborted after its leaseMs, so its permit is released
10 · dispose abandoned bodythe responseretry abandons a 5xx body and the fetch adapter's dispose cancels it
11 · breaker probe leasethe probea hung half-open probe releases its slot after probeLeaseTtlMs
12 · one shared rate budgetthe fleetone shared rate budget → the dependency sustains ~100 req/s, not ~400
13 · rate vs concurrencythe axisrate 100/s at 200ms latency → ~20 in flight (a rate limit is not a concurrency ceiling)
14 · burst and retry-afterthe knobburst: 20 admits a cluster, then 50/s sheds the surplus with a retry-after hint
15 · per-tenant rate quotathe tenant50 tenants, one noisy → the noisy tenant's rate is capped, the other 49 unaffected

Each number above is read from the most independent instrument available: the dependency's own counters (01, 02, 08), caracal's events for breaker and permit claims (03–05), the client's own timings (06), and Redis's own lease count (07). Three of those instruments are not caracal at all: the dependency's counters, the client's clocks, and Redis's own lease state. Where a claim can only be reported by caracal - which breaker opened, how many permits are live - the study says so rather than implying it was independently measured.

Run it

Requires Node ≥ 20.3 and Docker. The workloads run on the host; only the stack (Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana) runs in Docker.

npm install          # @gkoos/caracal + the workspace packages
npm run verify:api   # the installed caracal surface matches what the sandbox expects
npm run stack:up     # Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana
npm run smoke        # the whole chain: pipeline → events → OTLP → dashboards

Then open Grafana at http://localhost:3000/d/caracal-overview and Jaeger at http://localhost:16686. Tear the stack down with npm run stack:down.

Run any single case study with npm run demo <id> (each study's README has the exact command), and compare two runs side by side:

npm run compare 01 02

Make sure to run the demos before comparing, because the comparison is based on the summary.json files in runs/, generated by the demos.

More things to run:

CommandWhat it proves
npm run smokelocal policies, healthy dependency, everything observable
npm run smoke:distributedsame workload, one shared budget in Redis
npm run smoke:breakerthe breaker opens and sheds traffic (retry disabled so the signal isn't diluted)
npm run smoke:overloadthe bulkhead limit is below the offered concurrency: the surplus is shed as capacity and the breaker stays closed
npm run observabilitythe sink contract: a throwing/slow-async sink leaves p99 unchanged, a blocking one does not
CARACAL_OTEL=off npm run smokethe event → summary → compare path with no SDK and no stack

Watching it in Grafana

npm run stack:up provisions a Grafana dashboard — Caracal overview, the default home at http://localhost:3000. The witness is the first panel: the dependency's own concurrency count, the number the headline claims are decided on. Below it, the caracal panels show executions, attempts, bulkhead occupancy, breaker state, and coordination cost.

To compare two studies, run both (npm run demo 01 then npm run demo 02), then multi-select them in the Demo variable — the caracal panels overlay them, color-coded by demo. Widen the time range (it defaults to the last 5 minutes) so the runs are in view.

The witness panel is live only — it shows whatever is running now, not a per-run history — so the headline 20 vs 5 still comes from npm run compare 01 02.

What caracal is not

  • Caracal exposes events, not metrics or OpenTelemetry. The bridge in packages/caracal-observability is consumer code, written the way you'd write it.
  • A distributed bulkhead sheds, it does not queue. The case studies show the behavior.
  • The PostgreSQL adapter declares abort: "unsupported", so a timed-out query keeps its permit until it settles. Study 05 explains this.
  • Leases are not fencing tokens, and this is not exactly-once execution. Study 07 shows where that boundary actually is.

Findings

Building the sandbox surfaced a few non-obvious traps: a dead collector that fails only at shutdown, metric instruments that are silently permanent no-ops, a saturated bulkhead that tripped its own outer breaker (fixed in caracal 0.5.0), and seven more. Each is recorded with the evidence that produced it in docs/findings.md.

Layout

apps/partner-api/              # THE workload. Topology is configuration, not code
packages/caracal-observability # caracal events → OTLP traces + metrics (+ ndjson)
packages/caracal-runner/       # summary schema, named checks, compare/report
stack/                         # valkey, valkey-eu, postgres, otel, jaeger, prometheus, grafana
case-studies/                  # one case study per folder: study.yaml + README
runs/                          # artifacts (gitignored): summary.json, report.md, events.ndjson
scripts/                       # stack, smoke, demo, diagrams, verify-api, link-local, diagnose-otel

Diagrams

Every case study has a diagram.svg drawn from one visual language: replicas in terracotta, the dependency in steel blue, coordinators in brick red, the witness as a dashed emerald line, chaos as a red bolt. npm run diagrams regenerates them all from scripts/diagrams.mjs, so they stay consistent by construction.

Using the local caracal checkout

The case studies run against the published package. To test unreleased changes:

npm run link:local    # packs ../caracal and installs the tarball
npm run link:npm      # back to @gkoos/caracal@0.6.0

npm run verify:api fails if the surface drifts from what the sandbox uses.

Reading the comparison

npm run compare 01 02 is the whole argument in one table: two runs of the same workload against the same dependency, differing only in where the limit lives.

KPI                      01-local-only  02-distributed-basic  verdict
witness peak in-flight   20             5                     02-distributed-basic (lower is better)
successful requests/s    ~392           ~108                   01-local-only (higher is better)
rejected requests/s      ~5             ~324                   01-local-only (lower is better)
witness peak / limit     20 / 5         5 / 5                  -
coordinator trips/exec   n/a            ~4                     n/a

The verdict column looks like 01 wins on the rate rows. It doesn't, read it this way:

  • witness peak in-flight, 20 → 5, is the entire point. Four per-process limits still let the dependency's in-flight peak reach 20; one shared budget holds it at 5. This is the headline the study exists to show.
  • successful requests/s, ~392 → ~108, and rejected requests/s, ~5 → ~324, are the cost, and they are expected. The offered load far exceeds a fleet-wide budget of 5, so the surplus is rejected as capacity at the client instead of being pushed onto the dependency. Shedding is the budget working, not failing.
  • witness peak / limit, 20/5 → 5/5, is the budget made visible: 02 holds the dependency at its configured ceiling, while 01 shows four per-process 5s adding up to 20.
  • coordinator trips/exec, n/a → ~4, is the other cost of sharing: coordination is no longer free, each execution pays a few Redis roundtrips.

The rows that read same (breaker opens, peak probes, lease lost) are the control: the dependency stayed healthy in both runs, so the only thing that moved is the coordination boundary.

Capture a study's result as its committed baseline with npm run baseline <id>, then drift-check it with node --import tsx packages/caracal-runner/bin/compare.mjs <id> --against-baseline — it fails if any KPI moved more than 10%.

Contributors

gkoos

6 commits

gkoos/caracal-sandbox

A sandbox for exploring the Caracal distributed resilience library

TypeScript

1

6 commits

updated Sep 18, 2026

See the code

See what people are saying

SourceMessageScoreDate

Caracal: Scoped distributed resilience policies for asynchronous operations (r/webdev)

Scoped distributed resilience policies for asynchronous operations. Most resilience libraries stop at the process boundary. You get a local circuit breaker, a local bulkhead, a local rate limit. Spin up 40 replicas and those limits multiply: a “concurrency limit of 20” quietly becomes 800…

2

Sep 27, 2026

README

caracal-sandbox

A sandbox for exploring the distributed resilience library Caracal.

What is caracal?

Caracal is a distributed resilience library for TypeScript. It wraps an asynchronous operation in a policy pipeline — timeout, retry, circuit breaker, bulkhead — and, through Redis, lets many replicas share those policies. Its defining idea is scope: you say where a limit belongs — a process, a fleet, a region, a tenant, a shard — so the constraint matches your failure domain instead of multiplying per replica.

What is this sandbox?

Case studies that run the same workload against the same dependency, move the coordination boundary, and record what actually happened. Every headline number is read from the most independent instrument available - the dependency's own counters wherever they exist - and where a claim can only come from caracal's own events, the study says so and checks it against a witness caracal does not control.

The case studies, at a glance

Case studyThe boundary moves to...What it measures
01 · local-onlythe process4 × bulkhead.local({ limit: 5 }) → the dependency's peak in-flight reaches 20, not 5
02 · distributedthe fleetone shared budget → peak in-flight held at 5 (refused 83 → 4,816)
03 · multi-regionthe regionbreak eu's dependency → eu's breaker opens, us: 0 opens
04 · tenant isolationthe tenant50 tenants, one noisy → exactly 1 breaker opens
05 · postgres permit-holdthe querya 2s pg_sleep behind a 500ms timeout holds its permit ~2,059ms
06 · observabilitythe sinka blocking sink triples p99; a throwing/slow-async one leaves it unchanged
07 · chaosthe leasekill a replica, freeze another → live leases never observed above 3, return to 0
08 · overloadthe ceilinga budget of 20 against a capacity-8 dependency → the dependency itself 503s, not caracal
09 · local bulkhead leasethe permita hung bulkhead.local holder is aborted after its leaseMs, so its permit is released
10 · dispose abandoned bodythe responseretry abandons a 5xx body and the fetch adapter's dispose cancels it
11 · breaker probe leasethe probea hung half-open probe releases its slot after probeLeaseTtlMs
12 · one shared rate budgetthe fleetone shared rate budget → the dependency sustains ~100 req/s, not ~400
13 · rate vs concurrencythe axisrate 100/s at 200ms latency → ~20 in flight (a rate limit is not a concurrency ceiling)
14 · burst and retry-afterthe knobburst: 20 admits a cluster, then 50/s sheds the surplus with a retry-after hint
15 · per-tenant rate quotathe tenant50 tenants, one noisy → the noisy tenant's rate is capped, the other 49 unaffected

Each number above is read from the most independent instrument available: the dependency's own counters (01, 02, 08), caracal's events for breaker and permit claims (03–05), the client's own timings (06), and Redis's own lease count (07). Three of those instruments are not caracal at all: the dependency's counters, the client's clocks, and Redis's own lease state. Where a claim can only be reported by caracal - which breaker opened, how many permits are live - the study says so rather than implying it was independently measured.

Run it

Requires Node ≥ 20.3 and Docker. The workloads run on the host; only the stack (Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana) runs in Docker.

npm install          # @gkoos/caracal + the workspace packages
npm run verify:api   # the installed caracal surface matches what the sandbox expects
npm run stack:up     # Valkey, Postgres, otel-collector, Jaeger, Prometheus, Grafana
npm run smoke        # the whole chain: pipeline → events → OTLP → dashboards

Then open Grafana at http://localhost:3000/d/caracal-overview and Jaeger at http://localhost:16686. Tear the stack down with npm run stack:down.

Run any single case study with npm run demo <id> (each study's README has the exact command), and compare two runs side by side:

npm run compare 01 02

Make sure to run the demos before comparing, because the comparison is based on the summary.json files in runs/, generated by the demos.

More things to run:

CommandWhat it proves
npm run smokelocal policies, healthy dependency, everything observable
npm run smoke:distributedsame workload, one shared budget in Redis
npm run smoke:breakerthe breaker opens and sheds traffic (retry disabled so the signal isn't diluted)
npm run smoke:overloadthe bulkhead limit is below the offered concurrency: the surplus is shed as capacity and the breaker stays closed
npm run observabilitythe sink contract: a throwing/slow-async sink leaves p99 unchanged, a blocking one does not
CARACAL_OTEL=off npm run smokethe event → summary → compare path with no SDK and no stack

Watching it in Grafana

npm run stack:up provisions a Grafana dashboard — Caracal overview, the default home at http://localhost:3000. The witness is the first panel: the dependency's own concurrency count, the number the headline claims are decided on. Below it, the caracal panels show executions, attempts, bulkhead occupancy, breaker state, and coordination cost.

To compare two studies, run both (npm run demo 01 then npm run demo 02), then multi-select them in the Demo variable — the caracal panels overlay them, color-coded by demo. Widen the time range (it defaults to the last 5 minutes) so the runs are in view.

The witness panel is live only — it shows whatever is running now, not a per-run history — so the headline 20 vs 5 still comes from npm run compare 01 02.

What caracal is not

  • Caracal exposes events, not metrics or OpenTelemetry. The bridge in packages/caracal-observability is consumer code, written the way you'd write it.
  • A distributed bulkhead sheds, it does not queue. The case studies show the behavior.
  • The PostgreSQL adapter declares abort: "unsupported", so a timed-out query keeps its permit until it settles. Study 05 explains this.
  • Leases are not fencing tokens, and this is not exactly-once execution. Study 07 shows where that boundary actually is.

Findings

Building the sandbox surfaced a few non-obvious traps: a dead collector that fails only at shutdown, metric instruments that are silently permanent no-ops, a saturated bulkhead that tripped its own outer breaker (fixed in caracal 0.5.0), and seven more. Each is recorded with the evidence that produced it in docs/findings.md.

Layout

apps/partner-api/              # THE workload. Topology is configuration, not code
packages/caracal-observability # caracal events → OTLP traces + metrics (+ ndjson)
packages/caracal-runner/       # summary schema, named checks, compare/report
stack/                         # valkey, valkey-eu, postgres, otel, jaeger, prometheus, grafana
case-studies/                  # one case study per folder: study.yaml + README
runs/                          # artifacts (gitignored): summary.json, report.md, events.ndjson
scripts/                       # stack, smoke, demo, diagrams, verify-api, link-local, diagnose-otel

Diagrams

Every case study has a diagram.svg drawn from one visual language: replicas in terracotta, the dependency in steel blue, coordinators in brick red, the witness as a dashed emerald line, chaos as a red bolt. npm run diagrams regenerates them all from scripts/diagrams.mjs, so they stay consistent by construction.

Using the local caracal checkout

The case studies run against the published package. To test unreleased changes:

npm run link:local    # packs ../caracal and installs the tarball
npm run link:npm      # back to @gkoos/caracal@0.6.0

npm run verify:api fails if the surface drifts from what the sandbox uses.

Reading the comparison

npm run compare 01 02 is the whole argument in one table: two runs of the same workload against the same dependency, differing only in where the limit lives.

KPI                      01-local-only  02-distributed-basic  verdict
witness peak in-flight   20             5                     02-distributed-basic (lower is better)
successful requests/s    ~392           ~108                   01-local-only (higher is better)
rejected requests/s      ~5             ~324                   01-local-only (lower is better)
witness peak / limit     20 / 5         5 / 5                  -
coordinator trips/exec   n/a            ~4                     n/a

The verdict column looks like 01 wins on the rate rows. It doesn't, read it this way:

  • witness peak in-flight, 20 → 5, is the entire point. Four per-process limits still let the dependency's in-flight peak reach 20; one shared budget holds it at 5. This is the headline the study exists to show.
  • successful requests/s, ~392 → ~108, and rejected requests/s, ~5 → ~324, are the cost, and they are expected. The offered load far exceeds a fleet-wide budget of 5, so the surplus is rejected as capacity at the client instead of being pushed onto the dependency. Shedding is the budget working, not failing.
  • witness peak / limit, 20/5 → 5/5, is the budget made visible: 02 holds the dependency at its configured ceiling, while 01 shows four per-process 5s adding up to 20.
  • coordinator trips/exec, n/a → ~4, is the other cost of sharing: coordination is no longer free, each execution pays a few Redis roundtrips.

The rows that read same (breaker opens, peak probes, lease lost) are the control: the dependency stayed healthy in both runs, so the only thing that moved is the coordination boundary.

Capture a study's result as its committed baseline with npm run baseline <id>, then drift-check it with node --import tsx packages/caracal-runner/bin/compare.mjs <id> --against-baseline — it fails if any KPI moved more than 10%.

Contributors

gkoos

6 commits

Languages

TypeScript

66.4%

JavaScript

32.9%