Be your own indexer. One Rust binary, one command, live indexed API in under two minutes. No mandatory third-party data dependency.
See the codeTurn an EVM contract's history into a local SQL database. One Rust binary, no Postgres, no subgraph to write, and an MCP server built in.
· Website: www.nuthatch-indexer.com
curl -fsSL https://nuthatch-indexer.com/install.sh | sh # macOS Apple Silicon, Linux x86_64
nuthatch init 0xA0b86991c6218b36c1D19D4a2e9Eb0cE3606eB48 --alias usdc # USDC; the chain is detected
nuthatch dev --backfill 300 # the last 300 blocks, then keeps up
nuthatch sql "SELECT count(*) FROM usdc__transfer" # in a second terminal
The Linux binary needs glibc 2.35 or newer to run from 4.1.0, which is also what it is built on;
4.0.x ran on 2.34. No Intel Mac binary is published: there, and on other platforms, build from
source with Rust 1.95.0 (docs/install.md).
init creates a nest: a directory holding the contract's ABI, its config and, once dev runs,
its indexed data. --backfill 300 starts 300 blocks behind the tip, about an hour of mainnet, so there
are rows to query within seconds on the bundled public endpoints. Without it, dev backfills from the
contract's deployment block: for USDC that is 20 million blocks, a long backfill on free public endpoints
and a job for your own RPC (--rpc).
| Needs a subgraph | Needs handler code | Data comes from | What you run | Query with | |
|---|---|---|---|---|---|
| The Graph | yes | yes (AssemblyScript) | indexers on the network | nothing, or graph-node + Postgres + IPFS | GraphQL |
| Ponder | no | yes (TypeScript) | your RPC endpoint | Node.js, plus Postgres in production | SQL, GraphQL |
| nuthatch | no | no: tables come from the ABI | your RPC endpoint | one binary | SQL, HTTP, MCP |
Like Ponder, nuthatch reads the chain over JSON-RPC, so it needs an endpoint, and a provider may charge for one. The public endpoints it bundles are for trying it out, not for keeping it running. It complements The Graph rather than replacing it: a subgraph serves an application from a network of indexers, while a nest puts a contract's history in a database on your own machine.
Why. Getting at a contract's history usually means writing a subgraph or handler code, running a database, or renting someone else's copy. nuthatch generates the tables from the ABI, runs as one process with nothing else to install, and keeps the data on your machine: at most 2 GB of RAM per chain, no telemetry, no account. The built-in MCP server lets Claude or any MCP client query it.
init 0xAddr resolves the ABI (Sourcify, then Etherscan), generates the schema and
decoders, and scaffolds the project. You write nothing.An earlier example, now finished: Arcaidia, a speed layer over Circle's CCTP built at ETHOnline 2026, read its indexed state from two nests on Ethereum Sepolia and Arc Testnet. Its solver discovered intents there, its settlement agent tracked CCTP there, and its web console rendered from them. The nests were serving within two hours of the builder asking The Graph for a higher Studio rate limit. They were stopped on 2026-09-29.
More at nuthatch-indexer.com/stories.
Nightswatch does not run a hosted nest service.
curl -fsSL https://nuthatch-indexer.com/install.sh | sh
That downloads the prebuilt binary for your platform from the latest release, verifies its SHA-256,
and installs it to ~/.local/bin (override with NUTHATCH_INSTALL_DIR). No compiler is
involved. Prebuilt binaries cover macOS Apple Silicon and Linux x86_64 and are attached to every
release with their checksums, if you would rather fetch one by hand. No Intel Mac binary is
published; the installer says so and points at the source build below.
The Linux binary is dynamically linked and needs one thing, measured off the published
artifact with objdump -T rather than inferred:
GLIBC_2.34, so 2.34 was what you needed to run
it and 2.35 only what we compiled it on (#978);
4.1.0 references hypot at GLIBC_2.35, where libm re-versioned it, so the two numbers now agree.It links libc, libm and libgcc and no C++ runtime. Releases before 4.1 embedded DuckDB and
also needed libstdc++ from GCC 11.
Debian 12 and Ubuntu 22.04 clear it. RHEL 9 and Amazon Linux 2023 ship glibc 2.34 and ran 4.0.x; from 4.1.0 they need the source build.
Verify who built it. Every release binary carries a build provenance attestation, which a checksum cannot give you:
gh attestation verify nuthatch-x86_64-unknown-linux-gnu.tar.gz --repo nightswatchhq/nuthatch
From source, which is the only route on a platform we do not publish a binary for, Intel Macs included (the Intel build has not been verified on Intel hardware):
rustup toolchain install 1.95.0
cargo +1.95.0 install --git https://github.com/nightswatchhq/nuthatch nuthatch
The +1.95.0 is required: cargo install --git ignores the repo's toolchain pin, and a newer
default toolchain fails to compile a dependency.
Container images are published per release to ghcr.io/nightswatchhq/nuthatch - :<version> for
embedded, :<version>-scaled for the scaled build. The image ships the same binary attached to the
release, so the two cannot drift.
docs/install.md has the detail behind each of these: why the two ABI floors are
different numbers, what the attestation proves and what --repo is for, and why the toolchain pin
exists.
Chains. Ethereum, Arbitrum One, Base, BSC, Polygon, Gnosis, Optimism, Monad and Robinhood Chain are built in, with
measured public endpoints and tuned finality settings - omit --chain and nuthatch probes each for
your contract's bytecode and picks the one it lives on. Point at your own node with --rpc.
Any other EVM chain works too - World Chain, Base Sepolia, your own devnet. Name the chain and say where it lives:
nuthatch init 0xADDR --chain world-chain --rpc https://your-endpoint.example
The chain id is read from the endpoint itself, so there is no id to look up and nothing to type
wrong. A built-in chain never dials to learn its id: on one of the nine names --rpc is not consulted
for the chain id, but it is the pool - your endpoints replace the bundled public ones outright, and
nothing public is appended after them. Omit --rpc on an unregistered name and the refusal tells you
the remedy rather than just listing the built-ins.
Public endpoints are a moving target, and the ones shipped here are measured rather than assumed -
but a measurement is a snapshot, not a property. Run nuthatch doctor --rpc <url> before trusting a
long backfill to any endpoint, yours or ours: it reports the widest eth_getLogs range, the JSON-RPC
batch limit and whether the node has archive depth, and prints the largest safe --window.
Everything downstream was always chain-agnostic - dev, sql and bench, and the indexer's
unregistered-chain finality and window defaults. init's allow-list was the only thing narrower than
what nuthatch actually scaffolds, and it went in 2.4.0. See
running an unlisted EVM chain for the finality
caveat, which is the part worth reading: a chain whose finalized tag runs close to the tip needs a
depth-based policy instead, or you seal immutable Parquet that could never be corrected.
nuthatch assumes a paid RPC endpoint, or your own node, for anything you intend to keep running. That is the golden path.
Worth knowing, since we are being precise about it: most figures currently in
docs/benchmarks.md were measured against public endpoints, and that is a
known weakness of those numbers rather than a recommendation - a benchmark taken through a
rate-limited endpoint measures the endpoint. We measured the network at 99.3% of backfill wall
clock, which is why the replay rig (RFC-0039) exists and why those figures carry that caveat on the
page itself.
The free public endpoints bundled per chain exist for one job: so init → dev works with zero
setup, which is the two-minute demo, and it is deliberate. Treat them as testing and initial
validation - trying it out, checking a contract resolves, following the tip of something quiet.
They are the on-ramp, not the road.
Why they are not fine for real work, said here rather than discovered at 3am:
/ready reports stalled when that happens.eth_getLogs calls. Expect a free endpoint to throttle you long before that finishes.Check an endpoint before you trust a backfill to it. nuthatch doctor probes one and reports the
largest getLogs window it will actually serve, its batch limit, and whether it has archive history -
measured, not taken from the provider's documentation:
nuthatch doctor --rpc https://your-endpoint.example --address 0xADDR
Use your own endpoint for anything you care about - your own node, or a paid provider:
nuthatch init 0xADDR --chain arbitrum-one --rpc https://your-endpoint.example/arbitrum
nuthatch dev --rpc https://your-endpoint.example/arbitrum # or set rpc_urls in nuthatch.toml
--rpc is repeatable, and nuthatch round-robins across the pool with per-endpoint health tracking, so
listing two or three endpoints gets you failover as well as throughput. Every endpoint in a pool must be
on the same chain - nuthatch verifies this at startup and refuses to run against a mixed pool, since
indexing against the wrong chain corrupts state silently.
Every declared event becomes a table named {alias}__{event} (e.g. usdc__transfer), carrying the
event's fields plus block_number, block_hash, block_timestamp, tx_hash, log_index,
address and a _seq ordinal.
block_timestampcosts a block-header round trip per block - about 85% of backfill wall clock. A nest that will never ask a time-series question can drop the column withinit --no-timestampsand skip that entirely. It is an init-time choice: changing it later is a breaking schema change and a full re-index, so it is worth a moment's thought and is deliberately not a flag you can flip. Details.
# one-shot from the terminal (prints an aligned table; --json to pipe to jq)
nuthatch sql 'SELECT "from" AS sender, count(*) AS n FROM usdc__transfer GROUP BY 1 ORDER BY n DESC LIMIT 5'
# or over HTTP, against a running `nuthatch dev`
curl 'localhost:8288/sql?q=SELECT%20count(*)%20FROM%20usdc__transfer'
nuthatch sql queries the local store when dev is stopped, and transparently falls back to the
running instance's API when dev holds it - the same command works either way./sql returns degraded
and degraded_tables naming the affected tables, nuthatch sql prints a warning line, and the MCP
server carries the same notice. The caveat is a fact about the nest, not about the row count you
happened to get, so it appears whether or not this particular query touched the gap.bool column explains itself, because it is stored as exact text 'true'/'false' and therefore
blows up inside COALESCE, CASE, UNION and bool_and/bool_or while comparing fine on its
own. Same treatment on /sql, the MCP sql tool and the nuthatch sql REPL.uint256 values are exact text; amounts that fit in 38 digits also get a
{col}_dec DECIMAL view. SUM(value_dec) is those values: a full-width word is NULL and is not a
term. WHERE NOT {col}_overflow writes the same sum out. Ids, nonces and hashes stay on the raw
column.nuthatch mcp) - point Claude (or any
MCP client) at your indexer and ask your contract's data in plain English, fully offline.We ran someone else's benchmark rather than writing our own: Sentio's OBIB.
Case 1 indexes Transfer from LBTC across 22.2M Ethereum blocks.
| wall clock | 74.8 s |
| events | 294,278 (matches Sentio's own README exactly) |
| RPC requests | 321 |
| peak RSS | 320 MB |
Case 2 is case 1's contract with per-account balances, and OBIB's implementations get them with one
balanceOf() per account. We make none. For a plain ERC-20 the balance is the transfer history,
so we index the token's whole life instead and derive it - trading 2.5M extra blocks of cheap getLogs
for zero eth_call round trips.
| wall clock | 49.2 s (median of 3) |
| accounts | 7,634 - OBIB's published figure, exactly |
eth_call round trips | 0 |
| RPC requests | 136 |
| peak RSS | 325 MB |
Reference times for the same case: Sentio 7.78 min, Envio 8.54 min, Subsquid 46.85 min.
Two caveats, stated rather than buried. First, this is deliberately not like-for-like on range:
OBIB windows to 100,001 blocks, we index 2,611,334. On OBIB's own range we take 9.3 s - but that
run cannot produce the case's output at all, because absolute balances need history from before the
window, which is precisely why the benchmark makes the RPC calls. Second, "derived" is proven rather
than asserted: at the pinned end block, 39 sampled accounts - the ten largest, ten smallest non-zero,
ten zero-balance and ten by address order - all matched balanceOf(), including every zero-balance
account, which is the case an off-by-one in the ledger would betray.
The count is 7,634 and not 7,635 because 0x0 is the mint/burn counterparty rather than a holder. That
off-by-one was the tell that the interpretation was right.
Case 6 is the factory-template case: the Uniswap V2 factory over blocks 19,000,000-19,010,000,
discovering pairs from PairCreated and indexing Swap on every child it finds. No per-child config,
no redeploy, one rule.
| wall clock | 49.5 s (median of 5) |
| events | 35,271 = 35,039 swaps, matching OBIB's expected count exactly, plus the 232 PairCreated rows |
| children discovered | 232 |
| RPC requests | 16 |
| peak RSS | 247 MB |
For scale, OBIB's own published figures for case 6 differ between its two tables: the January 2026 results table gives Envio HyperIndex 1.92 min, Subsquid 5.34 min and Sentio 14.36 min, while the case-6 page reports Envio at 30 s from an earlier round. We are quoting both rather than the flattering one; on the second, Envio is faster than us. Note too that Envio and Subsquid serve this from their own pre-indexed networks, where nuthatch runs against plain JSON-RPC.
Both against a real provider (Alchemy), on an 11-core laptop. The artifacts are
docs/bench/obib-case1.json,
docs/bench/obib-case2.json and
docs/bench/obib-case6.json; nuthatch bench backfill re-runs any of them.
The case-2 nest is committed at obib-case2/ - keyless, so the endpoint arrives via
--rpc, and verified to rebuild from a clean checkout.
The case-6 nest is published at nightswatchhq/obib-case6
so the run can be reproduced rather than believed, and is submitted upstream as
sentioxyz/open-blockchain-indexer-benchmark#3.
Wall clock on a shared endpoint is the provider's number as much as ours. The same case-6 range on
the same commit measured anywhere from 17 s to 57 s depending on when it ran. We checked whether the
fast runs were provider caching by re-running against an adjacent, never-fetched range
(obib-case6-cold-control.json): it landed in the same
band, so caching is not the explanation. The event count and the 16 RPC requests are invariant
across every run, and they are the honest measure of range control.
Two things that number is worth knowing about:
block_timestamp - one serial round trip per block,
for a column that workload never stores. Timestamps are now demand-driven and the log window adapts
to what an endpoint will actually serve. See RFC-0029.Case 6 found a defect too, in the harness rather than the indexer: bench backfill fetched a fixed
address list, so a factory nest was measured without its children - 232 events in 2.6 s against an
expected 35,039, reported as a success. Running an outside benchmark has now found two things our own
testing did not.
Analytical queries run on Burrmill, our engine on DataFusion, over sealed Parquet. Until 4.1 they ran on DuckDB, and the change was not made for speed: measured on a production nest on 2026-10-01, Burrmill takes about 2.5× DuckDB's time for each statement and needs more memory for the same joins. What it buys is one language in the binary and exact arithmetic that refuses rather than wraps. The reasons and the log of the switch are in Replacing DuckDB, after all.
RPC ingestion → deterministic decode → redb hot store (tip)
│
past finality → content-addressed Parquet segments
│
Burrmill reads segments read-only → SQL (hot ∪ cold)
nuthatch has a Model Context Protocol server compiled in, so a coding agent can query your contract's data in plain English - offline, no phone-home. Wiring it is one step:
nuthatch dev & # the index the agent will query
nuthatch mcp --print-config # prints a copy-paste config for Claude Code / any MCP client
Or add it to Claude Code directly:
claude mcp add nuthatch -- nuthatch mcp --url http://127.0.0.1:8288
Then just ask: "what are the top USDC holders?" - the agent writes the SQL and runs it against your nest. (Making that correct on the first try is the semantic-layer work.)
Teach your agent to build nests too. Install the builder skill and an agent can drive nuthatch
itself - init, config, factories, compliance, multi-nest runtimes, troubleshooting - before you even have a nest:
cp -r skills/nuthatch-builder ~/.claude/skills/ # or your repo's .claude/skills/
Its CLI/config references are generated from the binary and CI-checked for drift, so the skill never lies about a flag (RFC-0017).
The core is "your contract → SQL." Beyond that, nuthatch has a full feature set for teams and operators who need more - none of it in the way of the happy path:
nuthatch.toml; index them together.[[calls]] block reads a contract at a fixed block and
stores the result as a table, optionally with calldata built from the row that triggered it - the
contract.balanceOf(event.params.user) a subgraph would write. Pinning the block is what keeps it
deterministic: the answer is fixed, so two operators re-running the same nest get the same bytes,
and the result is content-addressed on (chain_id, block, contract, calldata). Needs --state-rpc
pointed at an archive node.[[ipfs]] block turns a column of content addresses
into a table of resolved documents. Every body is re-hashed and checked against the CID it claims
to be, so a gateway serving the wrong bytes yields no row rather than a plausible one. The CID is
taken from whatever shape the contract stored - a bare CID, an ipfs:// URI, a gateway URL, or a
raw 32-byte digest - and the host is discarded, because that string came from a log and
honouring it would let whoever emitted the event choose what your indexer connects to.{template}__* tables - no redeploy per child.views/*.sql are named SQL evaluated at query time over hot ∪ sealed,
not IVM. And a nest can declare its own authored incremental entities in entities.toml
(RFC-0041): a SELECT that DBSP maintains as
blocks arrive, served from /derived and queryable by name from /sql, with reorgs handled as
retractions like the built-ins. On a real nest that took a panel from 2.15 s to 88 ms. A WASM
transform layer remains the imperative escape hatch./_admin/ - status, tables, view/nest inspector.
Localhost-open; off-localhost it requires a token per request.getLogs per window (N nests for
roughly one nest's RPC cost), and a runtime can span multiple chains with one isolated cursor per
chain - a Base nest and an Arbitrum nest in one runtime. Per-nest isolation, and a footprint budget
per active-chain cursor (≤2 GB). A capability, not a mandate: one chain per runtime stays the simple
default.POST /_admin/nests mounts one and
DELETE /_admin/nests/<name> unmounts one, live. A mount is admitted only if it fits the cursor's RAM
budget (refused, never warned - a budget that can be quietly exceeded is not a budget),
catches up before it joins so it never drags co-tenants back through history, and only then gets
routes. An unmount is a drain, not a route removal: the cursor finishes its window and releases
the store before anything is torn down. The set is persisted to mounts.toml, so a restart converges
on what you last asked for. An unmount keeps the dataset, so a remount is free; ?reclaim=true on the
DELETE removes it once no mount references it, and DELETE /_admin/datasets/<nid> reclaims one
unmounted earlier. A runtime may start with nothing mounted: declare its chains under [[chains]],
and the first mount onto a chain starts that chain's cursor, dialling its RPC only then.
Started with --registry, a runtime fetches a mounted NID it does not hold, verifies it as nest load
does, and installs it at data/<nid>/ first. A mount answers 202 at once, and GET /_admin/mounts/<name> reports it fetching, joining, live, or failed with the reason, across a
restart; ?wait=true answers only when it is done, with 507 for a breached budget. POST /_admin/suspend/<name> takes a mount off its cursor and answers 503 in its place, keeping its data
and record across a restart; POST /_admin/resume/<name> catches it up from where it stopped.
?dry_run=true on a mount reports its chain, backfill, per-block RPC work and projected footprint,
and the refusal a real mount would give, mounting nothing. POST /_admin/move/<name> with a new
nid catches the new nest up beside the old one, then switches the name in one step: a reader sees
the old nest, then the new, and never an error between.nuthatch worker) whose members take
cursor leases, and a query-FE tier (nuthatch serve) that serves from shared state and owns
nothing. A role flag, never a fork -
and opt-in at build time (--features postgres-store), so the published binary carries no database
driver and embedded mode stays a single file with zero services. The writer pool is safely scalable
because ownership is enforced by the store: every write carries a fence, and a stalled worker that
wakes up finds its writes refused rather than merely discouraged. Nests are added and removed
over HTTP with no restarts, versions are pinned fleet-wide so two FE nodes can never serve the same
endpoint from different schemas, and runtime secrets are injected at mount - scoped to the nests a
worker actually holds, write-only, and never baked into a content-addressed bundle. A worker pulls
the nests it is assigned from a registry, because the machine the scheduler picks may have nothing
on disk; with a bundle_hash pinned the fetch is by content address, so re-tagging a version in
a registry cannot change what a fleet runs. This is the
self-hosted distributed path for one operator's cooperating nests; per-tenant billing and authz
between untrusting paying customers stay firmly out of scope.nuthatch nest bundle packs
a nest's authored inputs into one portable, content-addressed .bundle; nest load <bundle-or-url>
verifies and installs it - regenerating the decode registry and asserting it matches - so anyone runs
your exact nest, hash-verified. Share at scale with a registry (RFC-0019): nest publish <bundle> --registry <path|s3://…> --as name@version, then nest load name@version --registry … - a filesystem
path or any S3-compatible bucket (MinIO/S3/R2, via AWS_* env), with private nests behind your
bucket's auth. Self-hosted-first: the registry is decoupled and never mandatory - a self-built bundle
and load <file|dir> need no registry at all. S3/MinIO/R2 is built in - configure it with the usual
AWS_* env (AWS_ENDPOINT for non-AWS), verified live against Hetzner Object Storage.nuthatch publish sync --target s3://bucket/prefix copies a nest's sealed Parquet segments, its catalogue and a
provenance envelope to any S3-compatible bucket or a directory, and dev --publish-target keeps
the mirror current as segments seal, uploading in streamed parts so ingestion does not wait on the
bucket. The mirror is keyed by the nest's data identity, not its NID: an edit that moves the NID but not
the data identity, which is the cosmetic case "Safe upgrades" below describes, keeps publishing to
the same dataset, and any edit that changes what is decoded forks a new one.
publish status says what is still to upload, publish verify checks every object against the
local segment (--deep re-downloads and re-hashes), and doctor --publish puts the mirror in a
health check. Reading it needs no nuthatch: DuckDB, Trino or anything that reads Parquet, as
Reading a published nest describes.--allow-breaking. Grafting does the rest: a cosmetic edit - a comment, a renamed view, a
doc change - moves the nest's identity and adopts the existing dataset, so nothing re-indexes.
Segments are content-addressed and shared across the runtime, so two nests that decode the same
contract hold one copy, not two. What a subgraph pays a full resync for, nuthatch answers with a
hash comparison.eth_call you don't need (RFC-0023). >70% of subgraphs call eth_call for
reads that are derivable from the events they already index - they fetch only because they have no
way to derive. Nuthatch does: nuthatch recipe add total_supply drops in a SQL view
that computes an ERC-20's totalSupply() as Σ minted − Σ burned from Transfer events - deterministic,
free, no archive node. That view runs at query time; it is not a DBSP circuit. It derives what a
subgraph pays an archive node to fetch. For the handful of
reads that aren't derivable but never change - decimals/symbol/name - nuthatch metadata fetch
calls once and caches forever.eth_getLogs is split and
retried, taking the provider's own suggested range when it offers one; a failure we cannot classify
is split once anyway, so an endpoint whose phrasing we have never seen still works rather than
stalling. Rate limits, transport blips and credential rejections are told apart - a rejected API key
is cooled down loudly instead of retried forever. And sealed segments now flush on a boundary derived
from the data, not from wherever a fetch window happened to stop, so two operators indexing the
same range produce byte-identical segments regardless of their RPC tuning./metrics - tip lag, rows decoded/sealed, reorgs, query counts, RSS.nuthatch is built to be fronted, not exposed raw - gateways, auth, and metering are the operator's
layer; nuthatch ships the guards (query timeout, row cap, result-byte cap, concurrency limit, a
filesystem-access denylist on /sql) and signals (/metrics) that make fronting it safe. It binds 127.0.0.1 by
default; --listen elsewhere and put a gateway in front. See docs/operators.md.
dev is the serve command - it backfills, follows the tip, and serves in one process.
Copy-paste systemd and Docker recipes are in docs/operators.md.docs/operators.md is the full operating guide, and worth reading before you
run this for real rather than after. docs/verification.md is its
counterpart: an acceptance runbook that proves a deployment works, step by falsifiable step, and says
plainly which levels we have verified ourselves and which we have not.
Still deciding whether to trust it at all? docs/kicking-the-tyres.md
is written for that: a guide to falsifying nuthatch rather than confirming it, with the cold walk,
correctness against a public subgraph, what it costs to keep running, a red-team pass on /sql, and a
section listing where we have already been wrong - including two of our own security patches and an
open finding. We would rather you found the next one than a user did.
The guide covers the questions people actually hit:
| If you're wondering | Go to |
|---|---|
| how do I tune backfill against my RPC's limits? | configuration surface - --window, --concurrency, --seal-direct |
| what do I scrape, and what should page me? | observability - metrics, alerts, health vs readiness |
| what happens when something breaks? | the failure model and the runbook |
| how do I back this up? | data lifecycle |
| how do I run an unlisted chain? | running an unlisted EVM chain |
| what isn't finished yet? | known gaps - stated plainly |
A major version is a promise about stability, not a claim of completeness.
nuthatch.toml, mounts.toml and entities.toml keep
working; a data directory upgrades drop-in, with no re-index; and the HTTP, SQL and MCP surfaces do
not break. Upgrade only: a downgrade is not promised. Off-by-default cargo features are
experimental and not covered. The full terms are the
stability contract.tests/upgrade_golden.rs).rust-toolchain.toml and the release
build all use. (Before 1.0 this file claimed 1.85, which cargo +1.85.0 check refutes in one
command. A version nobody tests is not a promise.)dev runs in production today, whether it is hosting one nest or many. Scaled mode
is built and verified across real machines, but younger - and until 0.9.3 its writer pool did not
index at all. If one process per box is enough, that is still the shape to reach for.What is deliberately not here: a hosted service, a token, telemetry, non-EVM chains, or any deployment story beyond binary + compose. Those are not backlog items; they are out of scope.
nuthatch binds 127.0.0.1 by default and is built to be fronted. Before you expose /sql to
anyone you do not trust, read SECURITY.md - and be on a current release:
/sql. DuckDB accepts a quoted function name and
the guard only matched an unquoted one, so SELECT * FROM "read_csv"('/etc/passwd') executed. Every
earlier release is affected./sql via ;-stacked COPY … TO.Both have published advisories on the repo's Security tab. The full pre-1.0 adversary pass, including
the findings we closed as not ours to fix and why, is in
docs/security-audit-2026-07-31.md.
docs/backlog.md; the running log is docs/progress-log.md.GOVERNANCE.md and the standing
design brief CLAUDE.md.Licensed under either of MIT or Apache-2.0 at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this work by you shall be dual licensed as above, without any additional terms or conditions.
be your own indexer.
Rust
93.2%
Shell
4.3%
Python
1.7%
Be your own indexer. One Rust binary, one command, live indexed API in under two minutes. No mandatory third-party data dependency.
See the codeTurn an EVM contract's history into a local SQL database. One Rust binary, no Postgres, no subgraph to write, and an MCP server built in.
· Website: www.nuthatch-indexer.com
curl -fsSL https://nuthatch-indexer.com/install.sh | sh # macOS Apple Silicon, Linux x86_64
nuthatch init 0xA0b86991c6218b36c1D19D4a2e9Eb0cE3606eB48 --alias usdc # USDC; the chain is detected
nuthatch dev --backfill 300 # the last 300 blocks, then keeps up
nuthatch sql "SELECT count(*) FROM usdc__transfer" # in a second terminal
The Linux binary needs glibc 2.35 or newer to run from 4.1.0, which is also what it is built on;
4.0.x ran on 2.34. No Intel Mac binary is published: there, and on other platforms, build from
source with Rust 1.95.0 (docs/install.md).
init creates a nest: a directory holding the contract's ABI, its config and, once dev runs,
its indexed data. --backfill 300 starts 300 blocks behind the tip, about an hour of mainnet, so there
are rows to query within seconds on the bundled public endpoints. Without it, dev backfills from the
contract's deployment block: for USDC that is 20 million blocks, a long backfill on free public endpoints
and a job for your own RPC (--rpc).
| Needs a subgraph | Needs handler code | Data comes from | What you run | Query with | |
|---|---|---|---|---|---|
| The Graph | yes | yes (AssemblyScript) | indexers on the network | nothing, or graph-node + Postgres + IPFS | GraphQL |
| Ponder | no | yes (TypeScript) | your RPC endpoint | Node.js, plus Postgres in production | SQL, GraphQL |
| nuthatch | no | no: tables come from the ABI | your RPC endpoint | one binary | SQL, HTTP, MCP |
Like Ponder, nuthatch reads the chain over JSON-RPC, so it needs an endpoint, and a provider may charge for one. The public endpoints it bundles are for trying it out, not for keeping it running. It complements The Graph rather than replacing it: a subgraph serves an application from a network of indexers, while a nest puts a contract's history in a database on your own machine.
Why. Getting at a contract's history usually means writing a subgraph or handler code, running a database, or renting someone else's copy. nuthatch generates the tables from the ABI, runs as one process with nothing else to install, and keeps the data on your machine: at most 2 GB of RAM per chain, no telemetry, no account. The built-in MCP server lets Claude or any MCP client query it.
init 0xAddr resolves the ABI (Sourcify, then Etherscan), generates the schema and
decoders, and scaffolds the project. You write nothing.An earlier example, now finished: Arcaidia, a speed layer over Circle's CCTP built at ETHOnline 2026, read its indexed state from two nests on Ethereum Sepolia and Arc Testnet. Its solver discovered intents there, its settlement agent tracked CCTP there, and its web console rendered from them. The nests were serving within two hours of the builder asking The Graph for a higher Studio rate limit. They were stopped on 2026-09-29.
More at nuthatch-indexer.com/stories.
Nightswatch does not run a hosted nest service.
curl -fsSL https://nuthatch-indexer.com/install.sh | sh
That downloads the prebuilt binary for your platform from the latest release, verifies its SHA-256,
and installs it to ~/.local/bin (override with NUTHATCH_INSTALL_DIR). No compiler is
involved. Prebuilt binaries cover macOS Apple Silicon and Linux x86_64 and are attached to every
release with their checksums, if you would rather fetch one by hand. No Intel Mac binary is
published; the installer says so and points at the source build below.
The Linux binary is dynamically linked and needs one thing, measured off the published
artifact with objdump -T rather than inferred:
GLIBC_2.34, so 2.34 was what you needed to run
it and 2.35 only what we compiled it on (#978);
4.1.0 references hypot at GLIBC_2.35, where libm re-versioned it, so the two numbers now agree.It links libc, libm and libgcc and no C++ runtime. Releases before 4.1 embedded DuckDB and
also needed libstdc++ from GCC 11.
Debian 12 and Ubuntu 22.04 clear it. RHEL 9 and Amazon Linux 2023 ship glibc 2.34 and ran 4.0.x; from 4.1.0 they need the source build.
Verify who built it. Every release binary carries a build provenance attestation, which a checksum cannot give you:
gh attestation verify nuthatch-x86_64-unknown-linux-gnu.tar.gz --repo nightswatchhq/nuthatch
From source, which is the only route on a platform we do not publish a binary for, Intel Macs included (the Intel build has not been verified on Intel hardware):
rustup toolchain install 1.95.0
cargo +1.95.0 install --git https://github.com/nightswatchhq/nuthatch nuthatch
The +1.95.0 is required: cargo install --git ignores the repo's toolchain pin, and a newer
default toolchain fails to compile a dependency.
Container images are published per release to ghcr.io/nightswatchhq/nuthatch - :<version> for
embedded, :<version>-scaled for the scaled build. The image ships the same binary attached to the
release, so the two cannot drift.
docs/install.md has the detail behind each of these: why the two ABI floors are
different numbers, what the attestation proves and what --repo is for, and why the toolchain pin
exists.
Chains. Ethereum, Arbitrum One, Base, BSC, Polygon, Gnosis, Optimism, Monad and Robinhood Chain are built in, with
measured public endpoints and tuned finality settings - omit --chain and nuthatch probes each for
your contract's bytecode and picks the one it lives on. Point at your own node with --rpc.
Any other EVM chain works too - World Chain, Base Sepolia, your own devnet. Name the chain and say where it lives:
nuthatch init 0xADDR --chain world-chain --rpc https://your-endpoint.example
The chain id is read from the endpoint itself, so there is no id to look up and nothing to type
wrong. A built-in chain never dials to learn its id: on one of the nine names --rpc is not consulted
for the chain id, but it is the pool - your endpoints replace the bundled public ones outright, and
nothing public is appended after them. Omit --rpc on an unregistered name and the refusal tells you
the remedy rather than just listing the built-ins.
Public endpoints are a moving target, and the ones shipped here are measured rather than assumed -
but a measurement is a snapshot, not a property. Run nuthatch doctor --rpc <url> before trusting a
long backfill to any endpoint, yours or ours: it reports the widest eth_getLogs range, the JSON-RPC
batch limit and whether the node has archive depth, and prints the largest safe --window.
Everything downstream was always chain-agnostic - dev, sql and bench, and the indexer's
unregistered-chain finality and window defaults. init's allow-list was the only thing narrower than
what nuthatch actually scaffolds, and it went in 2.4.0. See
running an unlisted EVM chain for the finality
caveat, which is the part worth reading: a chain whose finalized tag runs close to the tip needs a
depth-based policy instead, or you seal immutable Parquet that could never be corrected.
nuthatch assumes a paid RPC endpoint, or your own node, for anything you intend to keep running. That is the golden path.
Worth knowing, since we are being precise about it: most figures currently in
docs/benchmarks.md were measured against public endpoints, and that is a
known weakness of those numbers rather than a recommendation - a benchmark taken through a
rate-limited endpoint measures the endpoint. We measured the network at 99.3% of backfill wall
clock, which is why the replay rig (RFC-0039) exists and why those figures carry that caveat on the
page itself.
The free public endpoints bundled per chain exist for one job: so init → dev works with zero
setup, which is the two-minute demo, and it is deliberate. Treat them as testing and initial
validation - trying it out, checking a contract resolves, following the tip of something quiet.
They are the on-ramp, not the road.
Why they are not fine for real work, said here rather than discovered at 3am:
/ready reports stalled when that happens.eth_getLogs calls. Expect a free endpoint to throttle you long before that finishes.Check an endpoint before you trust a backfill to it. nuthatch doctor probes one and reports the
largest getLogs window it will actually serve, its batch limit, and whether it has archive history -
measured, not taken from the provider's documentation:
nuthatch doctor --rpc https://your-endpoint.example --address 0xADDR
Use your own endpoint for anything you care about - your own node, or a paid provider:
nuthatch init 0xADDR --chain arbitrum-one --rpc https://your-endpoint.example/arbitrum
nuthatch dev --rpc https://your-endpoint.example/arbitrum # or set rpc_urls in nuthatch.toml
--rpc is repeatable, and nuthatch round-robins across the pool with per-endpoint health tracking, so
listing two or three endpoints gets you failover as well as throughput. Every endpoint in a pool must be
on the same chain - nuthatch verifies this at startup and refuses to run against a mixed pool, since
indexing against the wrong chain corrupts state silently.
Every declared event becomes a table named {alias}__{event} (e.g. usdc__transfer), carrying the
event's fields plus block_number, block_hash, block_timestamp, tx_hash, log_index,
address and a _seq ordinal.
block_timestampcosts a block-header round trip per block - about 85% of backfill wall clock. A nest that will never ask a time-series question can drop the column withinit --no-timestampsand skip that entirely. It is an init-time choice: changing it later is a breaking schema change and a full re-index, so it is worth a moment's thought and is deliberately not a flag you can flip. Details.
# one-shot from the terminal (prints an aligned table; --json to pipe to jq)
nuthatch sql 'SELECT "from" AS sender, count(*) AS n FROM usdc__transfer GROUP BY 1 ORDER BY n DESC LIMIT 5'
# or over HTTP, against a running `nuthatch dev`
curl 'localhost:8288/sql?q=SELECT%20count(*)%20FROM%20usdc__transfer'
nuthatch sql queries the local store when dev is stopped, and transparently falls back to the
running instance's API when dev holds it - the same command works either way./sql returns degraded
and degraded_tables naming the affected tables, nuthatch sql prints a warning line, and the MCP
server carries the same notice. The caveat is a fact about the nest, not about the row count you
happened to get, so it appears whether or not this particular query touched the gap.bool column explains itself, because it is stored as exact text 'true'/'false' and therefore
blows up inside COALESCE, CASE, UNION and bool_and/bool_or while comparing fine on its
own. Same treatment on /sql, the MCP sql tool and the nuthatch sql REPL.uint256 values are exact text; amounts that fit in 38 digits also get a
{col}_dec DECIMAL view. SUM(value_dec) is those values: a full-width word is NULL and is not a
term. WHERE NOT {col}_overflow writes the same sum out. Ids, nonces and hashes stay on the raw
column.nuthatch mcp) - point Claude (or any
MCP client) at your indexer and ask your contract's data in plain English, fully offline.We ran someone else's benchmark rather than writing our own: Sentio's OBIB.
Case 1 indexes Transfer from LBTC across 22.2M Ethereum blocks.
| wall clock | 74.8 s |
| events | 294,278 (matches Sentio's own README exactly) |
| RPC requests | 321 |
| peak RSS | 320 MB |
Case 2 is case 1's contract with per-account balances, and OBIB's implementations get them with one
balanceOf() per account. We make none. For a plain ERC-20 the balance is the transfer history,
so we index the token's whole life instead and derive it - trading 2.5M extra blocks of cheap getLogs
for zero eth_call round trips.
| wall clock | 49.2 s (median of 3) |
| accounts | 7,634 - OBIB's published figure, exactly |
eth_call round trips | 0 |
| RPC requests | 136 |
| peak RSS | 325 MB |
Reference times for the same case: Sentio 7.78 min, Envio 8.54 min, Subsquid 46.85 min.
Two caveats, stated rather than buried. First, this is deliberately not like-for-like on range:
OBIB windows to 100,001 blocks, we index 2,611,334. On OBIB's own range we take 9.3 s - but that
run cannot produce the case's output at all, because absolute balances need history from before the
window, which is precisely why the benchmark makes the RPC calls. Second, "derived" is proven rather
than asserted: at the pinned end block, 39 sampled accounts - the ten largest, ten smallest non-zero,
ten zero-balance and ten by address order - all matched balanceOf(), including every zero-balance
account, which is the case an off-by-one in the ledger would betray.
The count is 7,634 and not 7,635 because 0x0 is the mint/burn counterparty rather than a holder. That
off-by-one was the tell that the interpretation was right.
Case 6 is the factory-template case: the Uniswap V2 factory over blocks 19,000,000-19,010,000,
discovering pairs from PairCreated and indexing Swap on every child it finds. No per-child config,
no redeploy, one rule.
| wall clock | 49.5 s (median of 5) |
| events | 35,271 = 35,039 swaps, matching OBIB's expected count exactly, plus the 232 PairCreated rows |
| children discovered | 232 |
| RPC requests | 16 |
| peak RSS | 247 MB |
For scale, OBIB's own published figures for case 6 differ between its two tables: the January 2026 results table gives Envio HyperIndex 1.92 min, Subsquid 5.34 min and Sentio 14.36 min, while the case-6 page reports Envio at 30 s from an earlier round. We are quoting both rather than the flattering one; on the second, Envio is faster than us. Note too that Envio and Subsquid serve this from their own pre-indexed networks, where nuthatch runs against plain JSON-RPC.
Both against a real provider (Alchemy), on an 11-core laptop. The artifacts are
docs/bench/obib-case1.json,
docs/bench/obib-case2.json and
docs/bench/obib-case6.json; nuthatch bench backfill re-runs any of them.
The case-2 nest is committed at obib-case2/ - keyless, so the endpoint arrives via
--rpc, and verified to rebuild from a clean checkout.
The case-6 nest is published at nightswatchhq/obib-case6
so the run can be reproduced rather than believed, and is submitted upstream as
sentioxyz/open-blockchain-indexer-benchmark#3.
Wall clock on a shared endpoint is the provider's number as much as ours. The same case-6 range on
the same commit measured anywhere from 17 s to 57 s depending on when it ran. We checked whether the
fast runs were provider caching by re-running against an adjacent, never-fetched range
(obib-case6-cold-control.json): it landed in the same
band, so caching is not the explanation. The event count and the 16 RPC requests are invariant
across every run, and they are the honest measure of range control.
Two things that number is worth knowing about:
block_timestamp - one serial round trip per block,
for a column that workload never stores. Timestamps are now demand-driven and the log window adapts
to what an endpoint will actually serve. See RFC-0029.Case 6 found a defect too, in the harness rather than the indexer: bench backfill fetched a fixed
address list, so a factory nest was measured without its children - 232 events in 2.6 s against an
expected 35,039, reported as a success. Running an outside benchmark has now found two things our own
testing did not.
Analytical queries run on Burrmill, our engine on DataFusion, over sealed Parquet. Until 4.1 they ran on DuckDB, and the change was not made for speed: measured on a production nest on 2026-10-01, Burrmill takes about 2.5× DuckDB's time for each statement and needs more memory for the same joins. What it buys is one language in the binary and exact arithmetic that refuses rather than wraps. The reasons and the log of the switch are in Replacing DuckDB, after all.
RPC ingestion → deterministic decode → redb hot store (tip)
│
past finality → content-addressed Parquet segments
│
Burrmill reads segments read-only → SQL (hot ∪ cold)
nuthatch has a Model Context Protocol server compiled in, so a coding agent can query your contract's data in plain English - offline, no phone-home. Wiring it is one step:
nuthatch dev & # the index the agent will query
nuthatch mcp --print-config # prints a copy-paste config for Claude Code / any MCP client
Or add it to Claude Code directly:
claude mcp add nuthatch -- nuthatch mcp --url http://127.0.0.1:8288
Then just ask: "what are the top USDC holders?" - the agent writes the SQL and runs it against your nest. (Making that correct on the first try is the semantic-layer work.)
Teach your agent to build nests too. Install the builder skill and an agent can drive nuthatch
itself - init, config, factories, compliance, multi-nest runtimes, troubleshooting - before you even have a nest:
cp -r skills/nuthatch-builder ~/.claude/skills/ # or your repo's .claude/skills/
Its CLI/config references are generated from the binary and CI-checked for drift, so the skill never lies about a flag (RFC-0017).
The core is "your contract → SQL." Beyond that, nuthatch has a full feature set for teams and operators who need more - none of it in the way of the happy path:
nuthatch.toml; index them together.[[calls]] block reads a contract at a fixed block and
stores the result as a table, optionally with calldata built from the row that triggered it - the
contract.balanceOf(event.params.user) a subgraph would write. Pinning the block is what keeps it
deterministic: the answer is fixed, so two operators re-running the same nest get the same bytes,
and the result is content-addressed on (chain_id, block, contract, calldata). Needs --state-rpc
pointed at an archive node.[[ipfs]] block turns a column of content addresses
into a table of resolved documents. Every body is re-hashed and checked against the CID it claims
to be, so a gateway serving the wrong bytes yields no row rather than a plausible one. The CID is
taken from whatever shape the contract stored - a bare CID, an ipfs:// URI, a gateway URL, or a
raw 32-byte digest - and the host is discarded, because that string came from a log and
honouring it would let whoever emitted the event choose what your indexer connects to.{template}__* tables - no redeploy per child.views/*.sql are named SQL evaluated at query time over hot ∪ sealed,
not IVM. And a nest can declare its own authored incremental entities in entities.toml
(RFC-0041): a SELECT that DBSP maintains as
blocks arrive, served from /derived and queryable by name from /sql, with reorgs handled as
retractions like the built-ins. On a real nest that took a panel from 2.15 s to 88 ms. A WASM
transform layer remains the imperative escape hatch./_admin/ - status, tables, view/nest inspector.
Localhost-open; off-localhost it requires a token per request.getLogs per window (N nests for
roughly one nest's RPC cost), and a runtime can span multiple chains with one isolated cursor per
chain - a Base nest and an Arbitrum nest in one runtime. Per-nest isolation, and a footprint budget
per active-chain cursor (≤2 GB). A capability, not a mandate: one chain per runtime stays the simple
default.POST /_admin/nests mounts one and
DELETE /_admin/nests/<name> unmounts one, live. A mount is admitted only if it fits the cursor's RAM
budget (refused, never warned - a budget that can be quietly exceeded is not a budget),
catches up before it joins so it never drags co-tenants back through history, and only then gets
routes. An unmount is a drain, not a route removal: the cursor finishes its window and releases
the store before anything is torn down. The set is persisted to mounts.toml, so a restart converges
on what you last asked for. An unmount keeps the dataset, so a remount is free; ?reclaim=true on the
DELETE removes it once no mount references it, and DELETE /_admin/datasets/<nid> reclaims one
unmounted earlier. A runtime may start with nothing mounted: declare its chains under [[chains]],
and the first mount onto a chain starts that chain's cursor, dialling its RPC only then.
Started with --registry, a runtime fetches a mounted NID it does not hold, verifies it as nest load
does, and installs it at data/<nid>/ first. A mount answers 202 at once, and GET /_admin/mounts/<name> reports it fetching, joining, live, or failed with the reason, across a
restart; ?wait=true answers only when it is done, with 507 for a breached budget. POST /_admin/suspend/<name> takes a mount off its cursor and answers 503 in its place, keeping its data
and record across a restart; POST /_admin/resume/<name> catches it up from where it stopped.
?dry_run=true on a mount reports its chain, backfill, per-block RPC work and projected footprint,
and the refusal a real mount would give, mounting nothing. POST /_admin/move/<name> with a new
nid catches the new nest up beside the old one, then switches the name in one step: a reader sees
the old nest, then the new, and never an error between.nuthatch worker) whose members take
cursor leases, and a query-FE tier (nuthatch serve) that serves from shared state and owns
nothing. A role flag, never a fork -
and opt-in at build time (--features postgres-store), so the published binary carries no database
driver and embedded mode stays a single file with zero services. The writer pool is safely scalable
because ownership is enforced by the store: every write carries a fence, and a stalled worker that
wakes up finds its writes refused rather than merely discouraged. Nests are added and removed
over HTTP with no restarts, versions are pinned fleet-wide so two FE nodes can never serve the same
endpoint from different schemas, and runtime secrets are injected at mount - scoped to the nests a
worker actually holds, write-only, and never baked into a content-addressed bundle. A worker pulls
the nests it is assigned from a registry, because the machine the scheduler picks may have nothing
on disk; with a bundle_hash pinned the fetch is by content address, so re-tagging a version in
a registry cannot change what a fleet runs. This is the
self-hosted distributed path for one operator's cooperating nests; per-tenant billing and authz
between untrusting paying customers stay firmly out of scope.nuthatch nest bundle packs
a nest's authored inputs into one portable, content-addressed .bundle; nest load <bundle-or-url>
verifies and installs it - regenerating the decode registry and asserting it matches - so anyone runs
your exact nest, hash-verified. Share at scale with a registry (RFC-0019): nest publish <bundle> --registry <path|s3://…> --as name@version, then nest load name@version --registry … - a filesystem
path or any S3-compatible bucket (MinIO/S3/R2, via AWS_* env), with private nests behind your
bucket's auth. Self-hosted-first: the registry is decoupled and never mandatory - a self-built bundle
and load <file|dir> need no registry at all. S3/MinIO/R2 is built in - configure it with the usual
AWS_* env (AWS_ENDPOINT for non-AWS), verified live against Hetzner Object Storage.nuthatch publish sync --target s3://bucket/prefix copies a nest's sealed Parquet segments, its catalogue and a
provenance envelope to any S3-compatible bucket or a directory, and dev --publish-target keeps
the mirror current as segments seal, uploading in streamed parts so ingestion does not wait on the
bucket. The mirror is keyed by the nest's data identity, not its NID: an edit that moves the NID but not
the data identity, which is the cosmetic case "Safe upgrades" below describes, keeps publishing to
the same dataset, and any edit that changes what is decoded forks a new one.
publish status says what is still to upload, publish verify checks every object against the
local segment (--deep re-downloads and re-hashes), and doctor --publish puts the mirror in a
health check. Reading it needs no nuthatch: DuckDB, Trino or anything that reads Parquet, as
Reading a published nest describes.--allow-breaking. Grafting does the rest: a cosmetic edit - a comment, a renamed view, a
doc change - moves the nest's identity and adopts the existing dataset, so nothing re-indexes.
Segments are content-addressed and shared across the runtime, so two nests that decode the same
contract hold one copy, not two. What a subgraph pays a full resync for, nuthatch answers with a
hash comparison.eth_call you don't need (RFC-0023). >70% of subgraphs call eth_call for
reads that are derivable from the events they already index - they fetch only because they have no
way to derive. Nuthatch does: nuthatch recipe add total_supply drops in a SQL view
that computes an ERC-20's totalSupply() as Σ minted − Σ burned from Transfer events - deterministic,
free, no archive node. That view runs at query time; it is not a DBSP circuit. It derives what a
subgraph pays an archive node to fetch. For the handful of
reads that aren't derivable but never change - decimals/symbol/name - nuthatch metadata fetch
calls once and caches forever.eth_getLogs is split and
retried, taking the provider's own suggested range when it offers one; a failure we cannot classify
is split once anyway, so an endpoint whose phrasing we have never seen still works rather than
stalling. Rate limits, transport blips and credential rejections are told apart - a rejected API key
is cooled down loudly instead of retried forever. And sealed segments now flush on a boundary derived
from the data, not from wherever a fetch window happened to stop, so two operators indexing the
same range produce byte-identical segments regardless of their RPC tuning./metrics - tip lag, rows decoded/sealed, reorgs, query counts, RSS.nuthatch is built to be fronted, not exposed raw - gateways, auth, and metering are the operator's
layer; nuthatch ships the guards (query timeout, row cap, result-byte cap, concurrency limit, a
filesystem-access denylist on /sql) and signals (/metrics) that make fronting it safe. It binds 127.0.0.1 by
default; --listen elsewhere and put a gateway in front. See docs/operators.md.
dev is the serve command - it backfills, follows the tip, and serves in one process.
Copy-paste systemd and Docker recipes are in docs/operators.md.docs/operators.md is the full operating guide, and worth reading before you
run this for real rather than after. docs/verification.md is its
counterpart: an acceptance runbook that proves a deployment works, step by falsifiable step, and says
plainly which levels we have verified ourselves and which we have not.
Still deciding whether to trust it at all? docs/kicking-the-tyres.md
is written for that: a guide to falsifying nuthatch rather than confirming it, with the cold walk,
correctness against a public subgraph, what it costs to keep running, a red-team pass on /sql, and a
section listing where we have already been wrong - including two of our own security patches and an
open finding. We would rather you found the next one than a user did.
The guide covers the questions people actually hit:
| If you're wondering | Go to |
|---|---|
| how do I tune backfill against my RPC's limits? | configuration surface - --window, --concurrency, --seal-direct |
| what do I scrape, and what should page me? | observability - metrics, alerts, health vs readiness |
| what happens when something breaks? | the failure model and the runbook |
| how do I back this up? | data lifecycle |
| how do I run an unlisted chain? | running an unlisted EVM chain |
| what isn't finished yet? | known gaps - stated plainly |
A major version is a promise about stability, not a claim of completeness.
nuthatch.toml, mounts.toml and entities.toml keep
working; a data directory upgrades drop-in, with no re-index; and the HTTP, SQL and MCP surfaces do
not break. Upgrade only: a downgrade is not promised. Off-by-default cargo features are
experimental and not covered. The full terms are the
stability contract.tests/upgrade_golden.rs).rust-toolchain.toml and the release
build all use. (Before 1.0 this file claimed 1.85, which cargo +1.85.0 check refutes in one
command. A version nobody tests is not a promise.)dev runs in production today, whether it is hosting one nest or many. Scaled mode
is built and verified across real machines, but younger - and until 0.9.3 its writer pool did not
index at all. If one process per box is enough, that is still the shape to reach for.What is deliberately not here: a hosted service, a token, telemetry, non-EVM chains, or any deployment story beyond binary + compose. Those are not backlog items; they are out of scope.
nuthatch binds 127.0.0.1 by default and is built to be fronted. Before you expose /sql to
anyone you do not trust, read SECURITY.md - and be on a current release:
/sql. DuckDB accepts a quoted function name and
the guard only matched an unquoted one, so SELECT * FROM "read_csv"('/etc/passwd') executed. Every
earlier release is affected./sql via ;-stacked COPY … TO.Both have published advisories on the repo's Security tab. The full pre-1.0 adversary pass, including
the findings we closed as not ours to fix and why, is in
docs/security-audit-2026-07-31.md.
docs/backlog.md; the running log is docs/progress-log.md.GOVERNANCE.md and the standing
design brief CLAUDE.md.Licensed under either of MIT or Apache-2.0 at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this work by you shall be dual licensed as above, without any additional terms or conditions.
be your own indexer.
Rust
93.2%
Shell
4.3%
Python
1.7%