
Documentation: https://cyb3rdudu.github.io/shardr/
shardr is a decentralized LLM repository with sync-based distribution. It keeps large language models available and digitally sovereign: artifacts live in content-addressed storage on your machines, synchronize over a BitTorrent-based peer network, and serve through an OpenAI-compatible runtime. A global system operated by its users — independent of any single provider, with availability that grows with every participating node.
A local content-addressed model repository. All models on a machine
live in a single store and are addressed by reference (ns/name:quant).
Identical files are stored once; every write is verified against its
digest.
Runtime-independent artifacts. An artifact contains weights, tokenizer, chat template, and metadata. Runtime configuration is applied separately at serve time and is not part of the stored model — so the repository is not tied to a specific inference runtime. llama-server is supported today; further runtimes can be added without re-importing existing models.
Collection management. shardr models lists the inventory with sizes
and quants, shardr verify --all re-hashes the store on demand, and
shardr pull retrieves missing artifacts when needed.
Distribution. Artifacts synchronize between shardhive instances over a BitTorrent-based peer network; availability grows with every participating node (see Synchronization & integrity).
| Component | Kind | Status |
|---|---|---|
shardhive | Storage daemon — content-addressed store (CAS), imports (local / Hugging Face / BitTorrent), sync client, API v1 over a 0600-mode Unix socket | working |
shardr | Model runtime & CLI — run/serve/stop lifecycle, layered runtime configuration, zero-copy serving via llama-server | working |
shardrbay | Discovery index over the peer network | planned |
shardr build | Modelfile composition — adapters and templates into distributable model images | planned |
go build -o sh-bin/shardr ./cmd/shardr
go build -o sh-bin/shardhive ./cmd/shardhive
make llama # fetches the prebuilt llama-server pinned in runtime/llama.lock into bin/
# or: any llama-server in $PATH, or $SHARDR_LLAMA_SERVER
Start the storage daemon (all clients communicate with it over a
mode-0600 Unix socket; $SHARDR_SOCKET overrides the path):
shardhive serve
On start, shardhive begins sharing every complete artifact it holds (startup sync).
# local files (regular files only — symlinks are refused; the quantization
# is derived from the filename, the model family from the stem)
shardr import local ~/Models/qwen3.5-9b-q8_0.gguf --as qwen/test
# Hugging Face (the commit SHA is pinned as provenance; identical
# classification rules apply)
shardr import hf Qwen/Qwen3-4B-Instruct-GGUF
# BitTorrent (the manifest pin is mandatory — a bare magnet link is never
# accepted)
shardr import bt "magnet:?xt=…" --manifest sha256:ab…
# Catalog (pirateface.co): search the listing, pull anchored — bytes
# verify against Hugging Face at the pinned revision, then keep seeding
# the listed swarm (good-citizen mode). A bare owner/repo (no :quant)
# is a catalog pull; rescued models need --trust-catalog.
shardr catalog search qwen 0.5b gguf
shardr pull bartowski/Qwen2.5-0.5B-Instruct-GGUF --quant raw
shardr models # inventory: namespaces, quants, sizes
shardr status # job progress / recent jobs
shardr run qwen/test:q8_0 # foreground, Ctrl-C = clean shutdown
shardr serve qwen/test:q8_0 --id mainllm # background instance
curl http://127.0.0.1:<port>/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"shardr:///qwen/test:q8_0","messages":[{"role":"user","content":"Hi"}]}'
shardr stop mainllm # SIGTERM within 30 s, then SIGKILL
The short reference is canonicalized, resolved and ensured against the daemon; llama-server then maps the weights directly from the CAS — no additional copies are written to disk. The served model id is the canonical reference.
Lowest to highest precedence: advisory defaults from the artifact →
~/.config/shardr/config.toml ([runtimes.llama], per-model
[models."ns/name:quant"]) → --config file.toml → --set key=value.
shardr run qwen/test:q8_0 --set llama.n_gpu_layers=40 --set llama.ctx_size=32768
Keys are validated against the 002 §7.1 allowlist (n_gpu_layers,
ctx_size, n_threads, flash_attn, mlock, kv_cache_type,
batch_size, ubatch_size, n_parallel, jinja, mmproj_variant).
Unknown keys fail loudly, naming the layer they came from. Bool keys are
tri-state: absent = inherit, true = pass flag, false = omit flag
(the runtime default applies — not "off").
Sharing is enabled by default: every complete artifact contributes to global
availability. Configure [swarm] in config.toml (seed, upload_limit,
dht). Missing content is filled automatically — CAS hit → import → peer
network (shardr pull <ref> fills without serving).
shardr verify --all # integrity re-hash (exit 0 clean, 1 mismatch, 2 missing)
Trust derives from digests, never from transports: every byte is verified against its content address on write, and BitTorrent imports require a pinned manifest digest.
The llama.cpp runtime shardr ships is pinned in
runtime/llama.lock — the single version truth:
an upstream prebuilt b-release (bNNNN) with the full commit SHA and
per-platform binary-archive SHA-256s, parsed fail-closed by
internal/llamalock. shardr never compiles llama.cpp (project decision
2026-09-05): make llama fetches the digest-verified prebuilt binaries.
Two strictly separated channels:
llama-upstream-check) detects newer b-releases that are at least 7
days old (community soak filter) and carry the full platform asset
matrix, then opens a lockfile update PR (old/new pin, both SHAs,
upstream compare link, per-platform digests). Merging that PR — after
the digest-verified fetch + real runner E2E matrix (Ubuntu x86-64,
macOS Apple Silicon) ran on it — triggers release-runner, which
publishes shardr-runner-<shardr-version>-llama-<ref> bundles: the
whole prebuilt runtime dir (dylibs are rpath-relative), SHA256SUMS,
BUILDINFO.json and licenses. Bundles are reproducible; existing
releases are never overwritten — same identity with different bytes is
a hard error. A moving latest is never published.llama-nightly-canary tests the newest
b-release weekly through the same matrix (no age filter — canary
catches breakage before the soak ends), never touches the stable
lockfile and never publishes a release.Manual update: go run ./cmd/llama-lock check-update --write --ref bNNNN
(exact b-releases, ≥ 7 days old, only), then open a PR against main.
Rolling the pin back to an older b-release needs the explicit
--allow-downgrade flag. Verify a lockfile with
go run ./cmd/llama-lock validate and prove its
provenance (tag → commit, release-API digests) with
go run ./cmd/llama-lock verify.
The documentation site is the user and developer documentation:
https://cyb3rdudu.github.io/shardr/ — built from this repository's
docs/ directory.
Design contracts live in docs/specs/ — see the
spec index for surface statuses. Reference scheme &
URI grammar (000), artifact format (001), runtime & configuration (002),
CAS (003), peer synchronization (004), interface & resolution (005).
Operational end-to-end: import → CAS → synchronization → shardr run with
a real GGUF and a real llama-server (verified with a 7.7 GB model). Not yet
built: shardr build (Modelfile composition), shardrbay, mlx/vllm
runtimes, CAS garbage collection.
Go
97.6%
Shell
2.2%

Documentation: https://cyb3rdudu.github.io/shardr/
shardr is a decentralized LLM repository with sync-based distribution. It keeps large language models available and digitally sovereign: artifacts live in content-addressed storage on your machines, synchronize over a BitTorrent-based peer network, and serve through an OpenAI-compatible runtime. A global system operated by its users — independent of any single provider, with availability that grows with every participating node.
A local content-addressed model repository. All models on a machine
live in a single store and are addressed by reference (ns/name:quant).
Identical files are stored once; every write is verified against its
digest.
Runtime-independent artifacts. An artifact contains weights, tokenizer, chat template, and metadata. Runtime configuration is applied separately at serve time and is not part of the stored model — so the repository is not tied to a specific inference runtime. llama-server is supported today; further runtimes can be added without re-importing existing models.
Collection management. shardr models lists the inventory with sizes
and quants, shardr verify --all re-hashes the store on demand, and
shardr pull retrieves missing artifacts when needed.
Distribution. Artifacts synchronize between shardhive instances over a BitTorrent-based peer network; availability grows with every participating node (see Synchronization & integrity).
| Component | Kind | Status |
|---|---|---|
shardhive | Storage daemon — content-addressed store (CAS), imports (local / Hugging Face / BitTorrent), sync client, API v1 over a 0600-mode Unix socket | working |
shardr | Model runtime & CLI — run/serve/stop lifecycle, layered runtime configuration, zero-copy serving via llama-server | working |
shardrbay | Discovery index over the peer network | planned |
shardr build | Modelfile composition — adapters and templates into distributable model images | planned |
go build -o sh-bin/shardr ./cmd/shardr
go build -o sh-bin/shardhive ./cmd/shardhive
make llama # fetches the prebuilt llama-server pinned in runtime/llama.lock into bin/
# or: any llama-server in $PATH, or $SHARDR_LLAMA_SERVER
Start the storage daemon (all clients communicate with it over a
mode-0600 Unix socket; $SHARDR_SOCKET overrides the path):
shardhive serve
On start, shardhive begins sharing every complete artifact it holds (startup sync).
# local files (regular files only — symlinks are refused; the quantization
# is derived from the filename, the model family from the stem)
shardr import local ~/Models/qwen3.5-9b-q8_0.gguf --as qwen/test
# Hugging Face (the commit SHA is pinned as provenance; identical
# classification rules apply)
shardr import hf Qwen/Qwen3-4B-Instruct-GGUF
# BitTorrent (the manifest pin is mandatory — a bare magnet link is never
# accepted)
shardr import bt "magnet:?xt=…" --manifest sha256:ab…
# Catalog (pirateface.co): search the listing, pull anchored — bytes
# verify against Hugging Face at the pinned revision, then keep seeding
# the listed swarm (good-citizen mode). A bare owner/repo (no :quant)
# is a catalog pull; rescued models need --trust-catalog.
shardr catalog search qwen 0.5b gguf
shardr pull bartowski/Qwen2.5-0.5B-Instruct-GGUF --quant raw
shardr models # inventory: namespaces, quants, sizes
shardr status # job progress / recent jobs
shardr run qwen/test:q8_0 # foreground, Ctrl-C = clean shutdown
shardr serve qwen/test:q8_0 --id mainllm # background instance
curl http://127.0.0.1:<port>/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"shardr:///qwen/test:q8_0","messages":[{"role":"user","content":"Hi"}]}'
shardr stop mainllm # SIGTERM within 30 s, then SIGKILL
The short reference is canonicalized, resolved and ensured against the daemon; llama-server then maps the weights directly from the CAS — no additional copies are written to disk. The served model id is the canonical reference.
Lowest to highest precedence: advisory defaults from the artifact →
~/.config/shardr/config.toml ([runtimes.llama], per-model
[models."ns/name:quant"]) → --config file.toml → --set key=value.
shardr run qwen/test:q8_0 --set llama.n_gpu_layers=40 --set llama.ctx_size=32768
Keys are validated against the 002 §7.1 allowlist (n_gpu_layers,
ctx_size, n_threads, flash_attn, mlock, kv_cache_type,
batch_size, ubatch_size, n_parallel, jinja, mmproj_variant).
Unknown keys fail loudly, naming the layer they came from. Bool keys are
tri-state: absent = inherit, true = pass flag, false = omit flag
(the runtime default applies — not "off").
Sharing is enabled by default: every complete artifact contributes to global
availability. Configure [swarm] in config.toml (seed, upload_limit,
dht). Missing content is filled automatically — CAS hit → import → peer
network (shardr pull <ref> fills without serving).
shardr verify --all # integrity re-hash (exit 0 clean, 1 mismatch, 2 missing)
Trust derives from digests, never from transports: every byte is verified against its content address on write, and BitTorrent imports require a pinned manifest digest.
The llama.cpp runtime shardr ships is pinned in
runtime/llama.lock — the single version truth:
an upstream prebuilt b-release (bNNNN) with the full commit SHA and
per-platform binary-archive SHA-256s, parsed fail-closed by
internal/llamalock. shardr never compiles llama.cpp (project decision
2026-09-05): make llama fetches the digest-verified prebuilt binaries.
Two strictly separated channels:
llama-upstream-check) detects newer b-releases that are at least 7
days old (community soak filter) and carry the full platform asset
matrix, then opens a lockfile update PR (old/new pin, both SHAs,
upstream compare link, per-platform digests). Merging that PR — after
the digest-verified fetch + real runner E2E matrix (Ubuntu x86-64,
macOS Apple Silicon) ran on it — triggers release-runner, which
publishes shardr-runner-<shardr-version>-llama-<ref> bundles: the
whole prebuilt runtime dir (dylibs are rpath-relative), SHA256SUMS,
BUILDINFO.json and licenses. Bundles are reproducible; existing
releases are never overwritten — same identity with different bytes is
a hard error. A moving latest is never published.llama-nightly-canary tests the newest
b-release weekly through the same matrix (no age filter — canary
catches breakage before the soak ends), never touches the stable
lockfile and never publishes a release.Manual update: go run ./cmd/llama-lock check-update --write --ref bNNNN
(exact b-releases, ≥ 7 days old, only), then open a PR against main.
Rolling the pin back to an older b-release needs the explicit
--allow-downgrade flag. Verify a lockfile with
go run ./cmd/llama-lock validate and prove its
provenance (tag → commit, release-API digests) with
go run ./cmd/llama-lock verify.
The documentation site is the user and developer documentation:
https://cyb3rdudu.github.io/shardr/ — built from this repository's
docs/ directory.
Design contracts live in docs/specs/ — see the
spec index for surface statuses. Reference scheme &
URI grammar (000), artifact format (001), runtime & configuration (002),
CAS (003), peer synchronization (004), interface & resolution (005).
Operational end-to-end: import → CAS → synchronization → shardr run with
a real GGUF and a real llama-server (verified with a 7.7 GB model). Not yet
built: shardr build (Modelfile composition), shardrbay, mlx/vllm
runtimes, CAS garbage collection.
Go
97.6%
Shell
2.2%