Cyb3rDudu/shardr

Go

0

94 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

shardr — like docker for models, BitTorrent sync, OpenAI-compatible serving, inference engines from upstream. (r/LocalLLaMA)

I kept running into the same three problems: the same 40 GB quant downloaded twice because it lived in some folder I forgot about, models quietly disappearing from Hugging Face, and every tool keeping its own copy of the weights on disk. So I've been building…

0

Oct 2, 2026

README

shardr logo

shardr

Documentation: https://cyb3rdudu.github.io/shardr/

shardr is a decentralized LLM repository with sync-based distribution. It keeps large language models available and digitally sovereign: artifacts live in content-addressed storage on your machines, synchronize over a BitTorrent-based peer network, and serve through an OpenAI-compatible runtime. A global system operated by its users — independent of any single provider, with availability that grows with every participating node.

What shardr provides

A local content-addressed model repository. All models on a machine live in a single store and are addressed by reference (ns/name:quant). Identical files are stored once; every write is verified against its digest.

Runtime-independent artifacts. An artifact contains weights, tokenizer, chat template, and metadata. Runtime configuration is applied separately at serve time and is not part of the stored model — so the repository is not tied to a specific inference runtime. llama-server is supported today; further runtimes can be added without re-importing existing models.

Collection management. shardr models lists the inventory with sizes and quants, shardr verify --all re-hashes the store on demand, and shardr pull retrieves missing artifacts when needed.

Distribution. Artifacts synchronize between shardhive instances over a BitTorrent-based peer network; availability grows with every participating node (see Synchronization & integrity).

Components

ComponentKindStatus
shardhiveStorage daemon — content-addressed store (CAS), imports (local / Hugging Face / BitTorrent), sync client, API v1 over a 0600-mode Unix socketworking
shardrModel runtime & CLI — run/serve/stop lifecycle, layered runtime configuration, zero-copy serving via llama-serverworking
shardrbayDiscovery index over the peer networkplanned
shardr buildModelfile composition — adapters and templates into distributable model imagesplanned

Setup

go build -o sh-bin/shardr ./cmd/shardr
go build -o sh-bin/shardhive ./cmd/shardhive
make llama        # fetches the prebuilt llama-server pinned in runtime/llama.lock into bin/
                  # or: any llama-server in $PATH, or $SHARDR_LLAMA_SERVER

Start the storage daemon (all clients communicate with it over a mode-0600 Unix socket; $SHARDR_SOCKET overrides the path):

shardhive serve

On start, shardhive begins sharing every complete artifact it holds (startup sync).

Adding models

# local files (regular files only — symlinks are refused; the quantization
# is derived from the filename, the model family from the stem)
shardr import local ~/Models/qwen3.5-9b-q8_0.gguf --as qwen/test

# Hugging Face (the commit SHA is pinned as provenance; identical
# classification rules apply)
shardr import hf Qwen/Qwen3-4B-Instruct-GGUF

# BitTorrent (the manifest pin is mandatory — a bare magnet link is never
# accepted)
shardr import bt "magnet:?xt=…" --manifest sha256:ab…

# Catalog (pirateface.co): search the listing, pull anchored — bytes
# verify against Hugging Face at the pinned revision, then keep seeding
# the listed swarm (good-citizen mode). A bare owner/repo (no :quant)
# is a catalog pull; rescued models need --trust-catalog.
shardr catalog search qwen 0.5b gguf
shardr pull bartowski/Qwen2.5-0.5B-Instruct-GGUF --quant raw

shardr models          # inventory: namespaces, quants, sizes
shardr status          # job progress / recent jobs

Serving models

shardr run qwen/test:q8_0              # foreground, Ctrl-C = clean shutdown

shardr serve qwen/test:q8_0 --id mainllm   # background instance
curl http://127.0.0.1:<port>/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"shardr:///qwen/test:q8_0","messages":[{"role":"user","content":"Hi"}]}'
shardr stop mainllm                    # SIGTERM within 30 s, then SIGKILL

The short reference is canonicalized, resolved and ensured against the daemon; llama-server then maps the weights directly from the CAS — no additional copies are written to disk. The served model id is the canonical reference.

Runtime configuration — four layers

Lowest to highest precedence: advisory defaults from the artifact → ~/.config/shardr/config.toml ([runtimes.llama], per-model [models."ns/name:quant"]) → --config file.toml → --set key=value.

shardr run qwen/test:q8_0 --set llama.n_gpu_layers=40 --set llama.ctx_size=32768

Keys are validated against the 002 §7.1 allowlist (n_gpu_layers, ctx_size, n_threads, flash_attn, mlock, kv_cache_type, batch_size, ubatch_size, n_parallel, jinja, mmproj_variant). Unknown keys fail loudly, naming the layer they came from. Bool keys are tri-state: absent = inherit, true = pass flag, false = omit flag (the runtime default applies — not "off").

Synchronization & integrity

Sharing is enabled by default: every complete artifact contributes to global availability. Configure [swarm] in config.toml (seed, upload_limit, dht). Missing content is filled automatically — CAS hit → import → peer network (shardr pull <ref> fills without serving).

shardr verify --all     # integrity re-hash (exit 0 clean, 1 mismatch, 2 missing)

Trust derives from digests, never from transports: every byte is verified against its content address on write, and BitTorrent imports require a pinned manifest digest.

llama.cpp versioning & runner releases

The llama.cpp runtime shardr ships is pinned in runtime/llama.lock — the single version truth: an upstream prebuilt b-release (bNNNN) with the full commit SHA and per-platform binary-archive SHA-256s, parsed fail-closed by internal/llamalock. shardr never compiles llama.cpp (project decision 2026-09-05): make llama fetches the digest-verified prebuilt binaries.

Two strictly separated channels:

  • Pin/release channel — a daily check (workflow llama-upstream-check) detects newer b-releases that are at least 7 days old (community soak filter) and carry the full platform asset matrix, then opens a lockfile update PR (old/new pin, both SHAs, upstream compare link, per-platform digests). Merging that PR — after the digest-verified fetch + real runner E2E matrix (Ubuntu x86-64, macOS Apple Silicon) ran on it — triggers release-runner, which publishes shardr-runner-<shardr-version>-llama-<ref> bundles: the whole prebuilt runtime dir (dylibs are rpath-relative), SHA256SUMS, BUILDINFO.json and licenses. Bundles are reproducible; existing releases are never overwritten — same identity with different bytes is a hard error. A moving latest is never published.
  • Canary channel — llama-nightly-canary tests the newest b-release weekly through the same matrix (no age filter — canary catches breakage before the soak ends), never touches the stable lockfile and never publishes a release.

Manual update: go run ./cmd/llama-lock check-update --write --ref bNNNN (exact b-releases, ≥ 7 days old, only), then open a PR against main. Rolling the pin back to an older b-release needs the explicit --allow-downgrade flag. Verify a lockfile with go run ./cmd/llama-lock validate and prove its provenance (tag → commit, release-API digests) with go run ./cmd/llama-lock verify.

Specifications & documentation

The documentation site is the user and developer documentation: https://cyb3rdudu.github.io/shardr/ — built from this repository's docs/ directory.

Design contracts live in docs/specs/ — see the spec index for surface statuses. Reference scheme & URI grammar (000), artifact format (001), runtime & configuration (002), CAS (003), peer synchronization (004), interface & resolution (005).

Status

Operational end-to-end: import → CAS → synchronization → shardr run with a real GGUF and a real llama-server (verified with a 7.7 GB model). Not yet built: shardr build (Modelfile composition), shardrbay, mlx/vllm runtimes, CAS garbage collection.

Cyb3rDudu/shardr

Go

0

94 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

shardr — like docker for models, BitTorrent sync, OpenAI-compatible serving, inference engines from upstream. (r/LocalLLaMA)

I kept running into the same three problems: the same 40 GB quant downloaded twice because it lived in some folder I forgot about, models quietly disappearing from Hugging Face, and every tool keeping its own copy of the weights on disk. So I've been building…

0

Oct 2, 2026

README

shardr logo

shardr

Documentation: https://cyb3rdudu.github.io/shardr/

shardr is a decentralized LLM repository with sync-based distribution. It keeps large language models available and digitally sovereign: artifacts live in content-addressed storage on your machines, synchronize over a BitTorrent-based peer network, and serve through an OpenAI-compatible runtime. A global system operated by its users — independent of any single provider, with availability that grows with every participating node.

What shardr provides

A local content-addressed model repository. All models on a machine live in a single store and are addressed by reference (ns/name:quant). Identical files are stored once; every write is verified against its digest.

Runtime-independent artifacts. An artifact contains weights, tokenizer, chat template, and metadata. Runtime configuration is applied separately at serve time and is not part of the stored model — so the repository is not tied to a specific inference runtime. llama-server is supported today; further runtimes can be added without re-importing existing models.

Collection management. shardr models lists the inventory with sizes and quants, shardr verify --all re-hashes the store on demand, and shardr pull retrieves missing artifacts when needed.

Distribution. Artifacts synchronize between shardhive instances over a BitTorrent-based peer network; availability grows with every participating node (see Synchronization & integrity).

Components

ComponentKindStatus
shardhiveStorage daemon — content-addressed store (CAS), imports (local / Hugging Face / BitTorrent), sync client, API v1 over a 0600-mode Unix socketworking
shardrModel runtime & CLI — run/serve/stop lifecycle, layered runtime configuration, zero-copy serving via llama-serverworking
shardrbayDiscovery index over the peer networkplanned
shardr buildModelfile composition — adapters and templates into distributable model imagesplanned

Setup

go build -o sh-bin/shardr ./cmd/shardr
go build -o sh-bin/shardhive ./cmd/shardhive
make llama        # fetches the prebuilt llama-server pinned in runtime/llama.lock into bin/
                  # or: any llama-server in $PATH, or $SHARDR_LLAMA_SERVER

Start the storage daemon (all clients communicate with it over a mode-0600 Unix socket; $SHARDR_SOCKET overrides the path):

shardhive serve

On start, shardhive begins sharing every complete artifact it holds (startup sync).

Adding models

# local files (regular files only — symlinks are refused; the quantization
# is derived from the filename, the model family from the stem)
shardr import local ~/Models/qwen3.5-9b-q8_0.gguf --as qwen/test

# Hugging Face (the commit SHA is pinned as provenance; identical
# classification rules apply)
shardr import hf Qwen/Qwen3-4B-Instruct-GGUF

# BitTorrent (the manifest pin is mandatory — a bare magnet link is never
# accepted)
shardr import bt "magnet:?xt=…" --manifest sha256:ab…

# Catalog (pirateface.co): search the listing, pull anchored — bytes
# verify against Hugging Face at the pinned revision, then keep seeding
# the listed swarm (good-citizen mode). A bare owner/repo (no :quant)
# is a catalog pull; rescued models need --trust-catalog.
shardr catalog search qwen 0.5b gguf
shardr pull bartowski/Qwen2.5-0.5B-Instruct-GGUF --quant raw

shardr models          # inventory: namespaces, quants, sizes
shardr status          # job progress / recent jobs

Serving models

shardr run qwen/test:q8_0              # foreground, Ctrl-C = clean shutdown

shardr serve qwen/test:q8_0 --id mainllm   # background instance
curl http://127.0.0.1:<port>/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"shardr:///qwen/test:q8_0","messages":[{"role":"user","content":"Hi"}]}'
shardr stop mainllm                    # SIGTERM within 30 s, then SIGKILL

The short reference is canonicalized, resolved and ensured against the daemon; llama-server then maps the weights directly from the CAS — no additional copies are written to disk. The served model id is the canonical reference.

Runtime configuration — four layers

Lowest to highest precedence: advisory defaults from the artifact → ~/.config/shardr/config.toml ([runtimes.llama], per-model [models."ns/name:quant"]) → --config file.toml → --set key=value.

shardr run qwen/test:q8_0 --set llama.n_gpu_layers=40 --set llama.ctx_size=32768

Keys are validated against the 002 §7.1 allowlist (n_gpu_layers, ctx_size, n_threads, flash_attn, mlock, kv_cache_type, batch_size, ubatch_size, n_parallel, jinja, mmproj_variant). Unknown keys fail loudly, naming the layer they came from. Bool keys are tri-state: absent = inherit, true = pass flag, false = omit flag (the runtime default applies — not "off").

Synchronization & integrity

Sharing is enabled by default: every complete artifact contributes to global availability. Configure [swarm] in config.toml (seed, upload_limit, dht). Missing content is filled automatically — CAS hit → import → peer network (shardr pull <ref> fills without serving).

shardr verify --all     # integrity re-hash (exit 0 clean, 1 mismatch, 2 missing)

Trust derives from digests, never from transports: every byte is verified against its content address on write, and BitTorrent imports require a pinned manifest digest.

llama.cpp versioning & runner releases

The llama.cpp runtime shardr ships is pinned in runtime/llama.lock — the single version truth: an upstream prebuilt b-release (bNNNN) with the full commit SHA and per-platform binary-archive SHA-256s, parsed fail-closed by internal/llamalock. shardr never compiles llama.cpp (project decision 2026-09-05): make llama fetches the digest-verified prebuilt binaries.

Two strictly separated channels:

  • Pin/release channel — a daily check (workflow llama-upstream-check) detects newer b-releases that are at least 7 days old (community soak filter) and carry the full platform asset matrix, then opens a lockfile update PR (old/new pin, both SHAs, upstream compare link, per-platform digests). Merging that PR — after the digest-verified fetch + real runner E2E matrix (Ubuntu x86-64, macOS Apple Silicon) ran on it — triggers release-runner, which publishes shardr-runner-<shardr-version>-llama-<ref> bundles: the whole prebuilt runtime dir (dylibs are rpath-relative), SHA256SUMS, BUILDINFO.json and licenses. Bundles are reproducible; existing releases are never overwritten — same identity with different bytes is a hard error. A moving latest is never published.
  • Canary channel — llama-nightly-canary tests the newest b-release weekly through the same matrix (no age filter — canary catches breakage before the soak ends), never touches the stable lockfile and never publishes a release.

Manual update: go run ./cmd/llama-lock check-update --write --ref bNNNN (exact b-releases, ≥ 7 days old, only), then open a PR against main. Rolling the pin back to an older b-release needs the explicit --allow-downgrade flag. Verify a lockfile with go run ./cmd/llama-lock validate and prove its provenance (tag → commit, release-API digests) with go run ./cmd/llama-lock verify.

Specifications & documentation

The documentation site is the user and developer documentation: https://cyb3rdudu.github.io/shardr/ — built from this repository's docs/ directory.

Design contracts live in docs/specs/ — see the spec index for surface statuses. Reference scheme & URI grammar (000), artifact format (001), runtime & configuration (002), CAS (003), peer synchronization (004), interface & resolution (005).

Status

Operational end-to-end: import → CAS → synchronization → shardr run with a real GGUF and a real llama-server (verified with a 7.7 GB model). Not yet built: shardr build (Modelfile composition), shardrbay, mlx/vllm runtimes, CAS garbage collection.

Languages

Go

97.6%

Shell

2.2%