Run any model on Intel silicon
See the code
Run any model on Intel hardware.
Cascadia distributes LLM inference across Intel laptops, desktops, and AI PCs. Shard a model across the machines you already have and serve it through an OpenAI-compatible API. No cloud or NVIDIA GPUs required.
Frontier models don't fit on a single laptop. Cloud APIs are expensive, opaque, and require sending your data offsite. Cascadia lets you point a few Intel machines at each other and run models that none of them could handle alone.
/v1/chat/completions with SSE streaming; point existing clients at it unchangedcascadia shard cuts a HuggingFace model into INT4 per-stage shards; no external toolingmock, ov-genai, ov-runtime, ov-dist-spec (distributed speculative decoding), gemma4, a CPU-targeted sparse-moe engine for large mixture-of-experts models (Kimi K2.6, MiniMax-M2, GLM-5, DeepSeek-V4, Inkling), and qwen35 for the Qwen3.5 hybrid family (Qwen3.6 MoE, dense Qwen3.8)cascadia discover finds LAN peers over mDNScascadia doctor: diagnoses the one failure everyone hits: OpenVINO silently not seeing your GPU[!NOTE] Cascadia is in alpha status. It works on Intel AI PCs (Lunar Lake / Arrow Lake / Panther Lake) and Arc B-series (Battlemage) discrete GPUs. Intel Arc A-series discrete GPUs and Xeon CPU-only servers are on the roadmap.
On an Intel machine? Grab a self-contained bundle from Releases: OpenVINO runtime included, no build, no SDK. Unpack it and run the binary from inside — it is not installed on your PATH:
# Linux
tar -xzf cascadia-<ver>-linux-x86_64.tar.gz && cd cascadia-<ver>-linux-x86_64
./cascadia doctor
# Windows
Expand-Archive cascadia-<ver>-windows-x86_64.zip -DestinationPath .
cd cascadia-<ver>-windows-x86_64
.\cascadia.exe doctor
That's the whole install: no Rust, no OpenVINO SDK, no INTEL_OPENVINO_DIR. Commands below are written as cascadia …; from a bundle use ./cascadia (.\cascadia.exe on Windows), or add the directory to your PATH.
Or build from source. You must have Rust installed; OpenVINO isn't required, and Cascadia will mock responses:
cargo build --release -p cascadia
./target/release/cascadia doctor
./target/release/cascadia run mock-model --engine mock
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "mock-model",
"messages": [{"role": "user", "content": "Capital of France?"}]
}'
If you receive a JSON chat-completion back, it means the full path (API → engine → streaming) works.
There are follow-up steps for performing real inference in QUICKSTART.md.
Prebuilt bundles for Linux and Windows x86_64 are on the Releases page: binary + OpenVINO runtime libraries, ready to run. Building from source instead has two modes:
# Stub mode. Rust only. Good for dev / CI on macOS / Linux / Windows.
# Engines that need OpenVINO return a clean runtime error.
cargo build --release -p cascadia
# Real OpenVINO mode. Links against openvino-genai 2026.2.0+. Required
# for inference on real Intel hardware.
INTEL_OPENVINO_DIR=/path/to/openvino_genai_<platform>_2026.2.0.0 \
cargo build --release -p cascadia --features openvino
Prerequisites:
--features openvinog++ ≥ 12 on Linux) and the OpenVINO GenAI SDKINSTALL.md has download links, the Linux GPU-runtime steps (scripts/setup-openvino.sh automates them from a source checkout), and the Docker image.
[!IMPORTANT] After building, run
cascadia doctor. On Intel AI PCs, the GPU can be invisible to OpenVINO even with a working driver. That failure is otherwise silent (you would just get slow CPU inference).doctordetects the problem and tells you how to fix it.
Every command and flag is catalogued in docs/CLI.md.
Cascadia serves models from a local directory — it does not download or convert at run time. Only cascadia shard fetches from HuggingFace (caching under ~/.cache/cascadia/models/). Export once, then serve:
# Export to a 1-stage INT4 shard (export deps: see Installation above).
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \
--output-dir ~/cascadia/llama-8b-1stage \
--num-stages 1 --quantization int4
# `run` is single-machine sugar: one stage, OpenAI API on :8000.
cascadia run ~/cascadia/llama-8b-1stage --engine ov-runtime --device GPU
Or serve a whole-model OpenVINO IR through the ov-genai engine, which adds FastDraft speculative decode and prompt-lookup. That layout comes from Intel's exporter, not cascadia shard — download a pre-exported INT4 IR (Intel publishes many under the OpenVINO org) or build one with optimum-cli, see docs/engines/ov-genai.md:
cascadia run ~/models/llama-3.1-8b-int4-ov # defaults to --engine ov-genai --device GPU
For full control over engine, device, ports, and the speculative / sparsity knobs, use cascadia worker (cascadia worker --help):
cascadia worker --rank 0 --total 1 --engine ov-genai --device GPU \
--model ~/models/llama-3.1-8b-int4-ov \
--api :8000
Shard once, on whichever machine has the RAM and a Python install:
# Export-time deps (~3 GB; not needed at runtime). From a source checkout:
pip install -r tools/requirements.txt
# From a release bundle there is no tools/ — `cascadia doctor` prints the pinned line.
# Shard a HuggingFace model into 2 stages with INT4 weights:
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \
--output-dir ~/cascadia/llama-8b-2stage \
--num-stages 2 --quantization int4
Copy the output directory to each node (scp -r / rsync), or re-shard separately on each node, whichever is faster on your network. Then run one worker per node: start the last stage first so the first stage finds it (if it isn't up yet, the first stage prints a clear "waiting for downstream peer" line and retries):
# Node B (last stage, listens for activations):
cascadia worker --rank 1 --total 2 --engine ov-runtime --device GPU \
--model ~/cascadia/llama-8b-2stage \
--listen :9100
# Node A (first stage, serves the API):
cascadia worker --rank 0 --total 2 --engine ov-runtime --device GPU \
--model ~/cascadia/llama-8b-2stage \
--next 10.0.0.2:9100 --api :8000
Not sure of a node's address? cascadia discover lists Cascadia peers on the LAN and the host:port to pass to --next.
For distributed speculative decoding, run every rank with --engine ov-dist-spec (they share a wire protocol) and give rank 0 --draft-model ~/models/llama-3.2-1b-int4-ov --spec-k 4 — the draft is a local OpenVINO IR directory, not an HF id. See (docs/engines/ov-dist-spec.md).
Workers started with --api also serve a browser dashboard at / — cluster topology with per-link latency/bandwidth, live request/token counters, and a chat surface. Release bundles newer than v0.1.8 include it. Source builds don't by default: the UI is a Vite SPA embedded into the binary behind the dashboard-embed cargo feature, and cargo can't run npm for you, so a plain cargo build serves the API plus a pointer page at / instead. To embed it (needs Node 20+):
cd crates/cascadia-dashboard/web
npm ci && npm run build
cd ../../..
cargo build --release -p cascadia --features dashboard-embed # add openvino for real inference
Use --release if you want the single-static-binary property: rust-embed
bakes the assets in for release builds, but reads web/dist from disk at
request time in debug ones, so a debug binary stops serving the UI if that
directory moves or is rebuilt.
Either way the JSON endpoints the UI reads (/api/topology, /api/stats) are always served, so during UI development npm run dev in crates/cascadia-dashboard/web hosts the SPA on :5173 and proxies API calls to a running worker (VITE_API_PROXY points it at a non-default host).
$ cascadia engines
mock deterministic word-echo engine for tests
ov-genai single-stage openvino_genai.LLMPipeline; FastDraft + Prompt Lookup
ov-runtime multi-stage stateful KV cache; pre-exported per-stage v3+ shards
ov-dist-spec multi-stage spec decode (mask-based KV rewind); v5 shards
gemma4 Gemma 4 multi-stage (per-layer-type attn, KV-sharing, PLI, sliding window); gemma4_cached_v1.x shards
sparse-moe sparse mixture-of-experts on CPU: Kimi K2.6 (AVX-512 int4 GEMM), MiniMax-M2 (OV-IR shells), GLM-5 / DeepSeek-V4 / Inkling (Rust shells + int4 mmap experts, N-rank pipeline)
qwen35 Qwen3.5-family staged chain (GatedDeltaNet; 3.5/3.6 MoE or 3.8 dense); qwen3_5* IR-surgery shards (alias: qwen36-moe)
sparse-moe consumes a manifest.json + per-expert artefact tree, not cascadia shard output, see docs/architectures/moe.md. In-repo exporters: tools/export_minimax_m2.py, export_glm5.py, export_deepseek_v4.py, export_inkling.py (per-family pages under docs/architectures/); the Kimi K2.6 artefacts come from an external pipeline that is not part of this repo. Tuning: docs/perf/A3_TOPK_REDUCTION.md, docs/perf/CHESS_PER_CHANNEL.md.
cascadia shard works today with Llama (1–3.3), Mistral (7B, NeMo, Small 3.x text), Qwen2 / Qwen2.5, Qwen3 dense, DeepSeek R1 Distills (Qwen and Llama variants), Phi-3, Phi-4 / Phi-4-mini (partial rotary), and Gemma 1 / Gemma 2 (logit softcapping + the 4-norm structure; sliding-window attention is treated as full-causal, so output is exact within the window). Gemma 4 (E2B / E4B / 31B) exports through a dedicated path (tools/export_gemma4.py, auto-dispatched by cascadia shard).
Qwen3.5/3.6 hybrid MoE (model_type: qwen3_5_moe) is special-cased: cascadia shard dispatches it to a dedicated IR-surgery exporter and it serves through the qwen35 engine (alias qwen36-moe) (docs/architectures/qwen36-moe-support.md). Other mixture-of-experts and architecturally-incompatible families like Llama 4, Qwen3-MoE, Mixtral, gpt-oss, full DeepSeek-V2/V3, Gemma 3, the Gemma 4 26B-A4B MoE variant, and Mamba hybrids are detected and rejected up front with a clear error. See docs/SHARDING.md and docs/architectures/ for the full per-family status table and deep-dives.
Cascadia is a Cargo workspace; one concern per crate. The Engine + Builder traits in cascadia-engine are the plugin seam. See docs/ARCHITECTURE.md for design rationale and per-crate responsibilities. Key crates:
cascadia-api/: OpenAI-compatible HTTP (axum)cascadia-metrics/: Prometheus metric registry shared by the API, runner, and transportcascadia-runner/: Per-stage runner; concurrent-safe chunk streamingcascadia-engine/: Engine + Builder traits (the plugin seam)cascadia-engine-openvino/: Five OV engines (ov-genai, ov-runtime, ov-dist-spec, gemma4, qwen35)cascadia-engine-sparse-moe/: Sparse-MoE engine; routes only the top-k experts per tokencascadia-int4-gemm/: hand-rolled AVX-512 INT4 GEMM kernels for the MoE expert pathcascadia-ov-genai-shim/: C++ FFI shim wrapping openvino-genaicascadia-transport/: TCP activation relay (length-prefixed tensor wire format)cascadia-topology/: Per-link latency + bandwidth measurementscascadia-discovery/: mDNS peer discovery on _cascadia._tcp.local.--rank / --total / --listen / --next host:port on each worker. cascadia discover browses the LAN, but workers still need explicit ranks, full auto-ring formation is not yet wired into cascadia worker (tracked in #89).cascadia profile-devices --model <dir> benchmarks each OV device (iGPU / NPU / CPU) on a host and writes device_profile.json, step 1 toward automatic placement. See docs/perf/DEVICE_PROFILE.md.--target npu) shards, cascadia worker --device CPU --prefill-device NPU runs the compute-bound chunked prefill on the NPU and the bandwidth-bound decode on the CPU, sharing one host KV ring. See docs/perf/HYBRID_NPU_CPU.md.Cascadia does not daemonize itself, so it needs to be run under systemd / NSSM / launchd. See docs/deploy/ for a systemd unit template and Windows / macOS recipes. Cascadia handles SIGTERM cleanly.
Security: the HTTP API and inter-stage TCP relay are plaintext and unauthenticated. Bind only to trusted networks (LAN, loopback) or terminate TLS + auth at a reverse proxy in front of --api. See SECURITY.md for the threat model and built-in hardening.
Monitoring: stages started with --api serve Prometheus metrics at GET /metrics — request rate/latency, TTFT, inter-token latency, token throughput, cancellations, model load times, and inter-stage transport bytes. Metric inventory and example queries in docs/METRICS.md.
config.json not in <model dir>: ov-runtime reads the HF model config.json from the shard's tokenizer dir to derive rotary parameters. Older shard exports may not bundle config.json; copy it from the source model's HF cache (~/.cache/huggingface/hub/models--<repo>/snapshots/<sha>/config.json) into the shards root. Shards produced by cascadia shard bundle it automatically.
could not connect to downstream peer within timeout (engines wait 60 s): start the downstream worker first; check --listen on the downstream matches --next on the upstream and that the host's firewall allows the port.
Worker dies silently when SSH session closes: on Windows OpenSSH the child process is tied to the SSH parent. Run workers under systemd / NSSM / Task Scheduler in production.
cascadia doctor diagnoses most other environment/hardware issues.
| Doc | What's in it |
|---|---|
| QUICKSTART.md | 5-minute stub run → real inference |
| INSTALL.md | Full setup: OpenVINO SDK, GPU runtime, Docker |
| docs/CLI.md | Every command and flag |
| docs/SHARDING.md | Sharding flow + per-model-family support table |
| docs/ARCHITECTURE.md | Design decisions + crate responsibilities |
| docs/METRICS.md | Prometheus /metrics inventory + example queries |
| docs/engines/ | Per-engine deep dives |
| docs/architectures/ | Per-model-family export/support notes |
| docs/perf/ | Performance investigations and tuning |
| SECURITY.md | Threat model + vulnerability reporting |
See CONTRIBUTING.md for the build/test gate, crate layout, and commit conventions.
By participating you agree to the Code of Conduct.
Apache-2.0 (see LICENSE). Third-party attributions are in THIRD_PARTY_NOTICES.md.
Rust
78.4%
Python
18.0%
C++
2.0%
Run any model on Intel silicon
See the code
Run any model on Intel hardware.
Cascadia distributes LLM inference across Intel laptops, desktops, and AI PCs. Shard a model across the machines you already have and serve it through an OpenAI-compatible API. No cloud or NVIDIA GPUs required.
Frontier models don't fit on a single laptop. Cloud APIs are expensive, opaque, and require sending your data offsite. Cascadia lets you point a few Intel machines at each other and run models that none of them could handle alone.
/v1/chat/completions with SSE streaming; point existing clients at it unchangedcascadia shard cuts a HuggingFace model into INT4 per-stage shards; no external toolingmock, ov-genai, ov-runtime, ov-dist-spec (distributed speculative decoding), gemma4, a CPU-targeted sparse-moe engine for large mixture-of-experts models (Kimi K2.6, MiniMax-M2, GLM-5, DeepSeek-V4, Inkling), and qwen35 for the Qwen3.5 hybrid family (Qwen3.6 MoE, dense Qwen3.8)cascadia discover finds LAN peers over mDNScascadia doctor: diagnoses the one failure everyone hits: OpenVINO silently not seeing your GPU[!NOTE] Cascadia is in alpha status. It works on Intel AI PCs (Lunar Lake / Arrow Lake / Panther Lake) and Arc B-series (Battlemage) discrete GPUs. Intel Arc A-series discrete GPUs and Xeon CPU-only servers are on the roadmap.
On an Intel machine? Grab a self-contained bundle from Releases: OpenVINO runtime included, no build, no SDK. Unpack it and run the binary from inside — it is not installed on your PATH:
# Linux
tar -xzf cascadia-<ver>-linux-x86_64.tar.gz && cd cascadia-<ver>-linux-x86_64
./cascadia doctor
# Windows
Expand-Archive cascadia-<ver>-windows-x86_64.zip -DestinationPath .
cd cascadia-<ver>-windows-x86_64
.\cascadia.exe doctor
That's the whole install: no Rust, no OpenVINO SDK, no INTEL_OPENVINO_DIR. Commands below are written as cascadia …; from a bundle use ./cascadia (.\cascadia.exe on Windows), or add the directory to your PATH.
Or build from source. You must have Rust installed; OpenVINO isn't required, and Cascadia will mock responses:
cargo build --release -p cascadia
./target/release/cascadia doctor
./target/release/cascadia run mock-model --engine mock
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "mock-model",
"messages": [{"role": "user", "content": "Capital of France?"}]
}'
If you receive a JSON chat-completion back, it means the full path (API → engine → streaming) works.
There are follow-up steps for performing real inference in QUICKSTART.md.
Prebuilt bundles for Linux and Windows x86_64 are on the Releases page: binary + OpenVINO runtime libraries, ready to run. Building from source instead has two modes:
# Stub mode. Rust only. Good for dev / CI on macOS / Linux / Windows.
# Engines that need OpenVINO return a clean runtime error.
cargo build --release -p cascadia
# Real OpenVINO mode. Links against openvino-genai 2026.2.0+. Required
# for inference on real Intel hardware.
INTEL_OPENVINO_DIR=/path/to/openvino_genai_<platform>_2026.2.0.0 \
cargo build --release -p cascadia --features openvino
Prerequisites:
--features openvinog++ ≥ 12 on Linux) and the OpenVINO GenAI SDKINSTALL.md has download links, the Linux GPU-runtime steps (scripts/setup-openvino.sh automates them from a source checkout), and the Docker image.
[!IMPORTANT] After building, run
cascadia doctor. On Intel AI PCs, the GPU can be invisible to OpenVINO even with a working driver. That failure is otherwise silent (you would just get slow CPU inference).doctordetects the problem and tells you how to fix it.
Every command and flag is catalogued in docs/CLI.md.
Cascadia serves models from a local directory — it does not download or convert at run time. Only cascadia shard fetches from HuggingFace (caching under ~/.cache/cascadia/models/). Export once, then serve:
# Export to a 1-stage INT4 shard (export deps: see Installation above).
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \
--output-dir ~/cascadia/llama-8b-1stage \
--num-stages 1 --quantization int4
# `run` is single-machine sugar: one stage, OpenAI API on :8000.
cascadia run ~/cascadia/llama-8b-1stage --engine ov-runtime --device GPU
Or serve a whole-model OpenVINO IR through the ov-genai engine, which adds FastDraft speculative decode and prompt-lookup. That layout comes from Intel's exporter, not cascadia shard — download a pre-exported INT4 IR (Intel publishes many under the OpenVINO org) or build one with optimum-cli, see docs/engines/ov-genai.md:
cascadia run ~/models/llama-3.1-8b-int4-ov # defaults to --engine ov-genai --device GPU
For full control over engine, device, ports, and the speculative / sparsity knobs, use cascadia worker (cascadia worker --help):
cascadia worker --rank 0 --total 1 --engine ov-genai --device GPU \
--model ~/models/llama-3.1-8b-int4-ov \
--api :8000
Shard once, on whichever machine has the RAM and a Python install:
# Export-time deps (~3 GB; not needed at runtime). From a source checkout:
pip install -r tools/requirements.txt
# From a release bundle there is no tools/ — `cascadia doctor` prints the pinned line.
# Shard a HuggingFace model into 2 stages with INT4 weights:
cascadia shard --model unsloth/Meta-Llama-3.1-8B-Instruct \
--output-dir ~/cascadia/llama-8b-2stage \
--num-stages 2 --quantization int4
Copy the output directory to each node (scp -r / rsync), or re-shard separately on each node, whichever is faster on your network. Then run one worker per node: start the last stage first so the first stage finds it (if it isn't up yet, the first stage prints a clear "waiting for downstream peer" line and retries):
# Node B (last stage, listens for activations):
cascadia worker --rank 1 --total 2 --engine ov-runtime --device GPU \
--model ~/cascadia/llama-8b-2stage \
--listen :9100
# Node A (first stage, serves the API):
cascadia worker --rank 0 --total 2 --engine ov-runtime --device GPU \
--model ~/cascadia/llama-8b-2stage \
--next 10.0.0.2:9100 --api :8000
Not sure of a node's address? cascadia discover lists Cascadia peers on the LAN and the host:port to pass to --next.
For distributed speculative decoding, run every rank with --engine ov-dist-spec (they share a wire protocol) and give rank 0 --draft-model ~/models/llama-3.2-1b-int4-ov --spec-k 4 — the draft is a local OpenVINO IR directory, not an HF id. See (docs/engines/ov-dist-spec.md).
Workers started with --api also serve a browser dashboard at / — cluster topology with per-link latency/bandwidth, live request/token counters, and a chat surface. Release bundles newer than v0.1.8 include it. Source builds don't by default: the UI is a Vite SPA embedded into the binary behind the dashboard-embed cargo feature, and cargo can't run npm for you, so a plain cargo build serves the API plus a pointer page at / instead. To embed it (needs Node 20+):
cd crates/cascadia-dashboard/web
npm ci && npm run build
cd ../../..
cargo build --release -p cascadia --features dashboard-embed # add openvino for real inference
Use --release if you want the single-static-binary property: rust-embed
bakes the assets in for release builds, but reads web/dist from disk at
request time in debug ones, so a debug binary stops serving the UI if that
directory moves or is rebuilt.
Either way the JSON endpoints the UI reads (/api/topology, /api/stats) are always served, so during UI development npm run dev in crates/cascadia-dashboard/web hosts the SPA on :5173 and proxies API calls to a running worker (VITE_API_PROXY points it at a non-default host).
$ cascadia engines
mock deterministic word-echo engine for tests
ov-genai single-stage openvino_genai.LLMPipeline; FastDraft + Prompt Lookup
ov-runtime multi-stage stateful KV cache; pre-exported per-stage v3+ shards
ov-dist-spec multi-stage spec decode (mask-based KV rewind); v5 shards
gemma4 Gemma 4 multi-stage (per-layer-type attn, KV-sharing, PLI, sliding window); gemma4_cached_v1.x shards
sparse-moe sparse mixture-of-experts on CPU: Kimi K2.6 (AVX-512 int4 GEMM), MiniMax-M2 (OV-IR shells), GLM-5 / DeepSeek-V4 / Inkling (Rust shells + int4 mmap experts, N-rank pipeline)
qwen35 Qwen3.5-family staged chain (GatedDeltaNet; 3.5/3.6 MoE or 3.8 dense); qwen3_5* IR-surgery shards (alias: qwen36-moe)
sparse-moe consumes a manifest.json + per-expert artefact tree, not cascadia shard output, see docs/architectures/moe.md. In-repo exporters: tools/export_minimax_m2.py, export_glm5.py, export_deepseek_v4.py, export_inkling.py (per-family pages under docs/architectures/); the Kimi K2.6 artefacts come from an external pipeline that is not part of this repo. Tuning: docs/perf/A3_TOPK_REDUCTION.md, docs/perf/CHESS_PER_CHANNEL.md.
cascadia shard works today with Llama (1–3.3), Mistral (7B, NeMo, Small 3.x text), Qwen2 / Qwen2.5, Qwen3 dense, DeepSeek R1 Distills (Qwen and Llama variants), Phi-3, Phi-4 / Phi-4-mini (partial rotary), and Gemma 1 / Gemma 2 (logit softcapping + the 4-norm structure; sliding-window attention is treated as full-causal, so output is exact within the window). Gemma 4 (E2B / E4B / 31B) exports through a dedicated path (tools/export_gemma4.py, auto-dispatched by cascadia shard).
Qwen3.5/3.6 hybrid MoE (model_type: qwen3_5_moe) is special-cased: cascadia shard dispatches it to a dedicated IR-surgery exporter and it serves through the qwen35 engine (alias qwen36-moe) (docs/architectures/qwen36-moe-support.md). Other mixture-of-experts and architecturally-incompatible families like Llama 4, Qwen3-MoE, Mixtral, gpt-oss, full DeepSeek-V2/V3, Gemma 3, the Gemma 4 26B-A4B MoE variant, and Mamba hybrids are detected and rejected up front with a clear error. See docs/SHARDING.md and docs/architectures/ for the full per-family status table and deep-dives.
Cascadia is a Cargo workspace; one concern per crate. The Engine + Builder traits in cascadia-engine are the plugin seam. See docs/ARCHITECTURE.md for design rationale and per-crate responsibilities. Key crates:
cascadia-api/: OpenAI-compatible HTTP (axum)cascadia-metrics/: Prometheus metric registry shared by the API, runner, and transportcascadia-runner/: Per-stage runner; concurrent-safe chunk streamingcascadia-engine/: Engine + Builder traits (the plugin seam)cascadia-engine-openvino/: Five OV engines (ov-genai, ov-runtime, ov-dist-spec, gemma4, qwen35)cascadia-engine-sparse-moe/: Sparse-MoE engine; routes only the top-k experts per tokencascadia-int4-gemm/: hand-rolled AVX-512 INT4 GEMM kernels for the MoE expert pathcascadia-ov-genai-shim/: C++ FFI shim wrapping openvino-genaicascadia-transport/: TCP activation relay (length-prefixed tensor wire format)cascadia-topology/: Per-link latency + bandwidth measurementscascadia-discovery/: mDNS peer discovery on _cascadia._tcp.local.--rank / --total / --listen / --next host:port on each worker. cascadia discover browses the LAN, but workers still need explicit ranks, full auto-ring formation is not yet wired into cascadia worker (tracked in #89).cascadia profile-devices --model <dir> benchmarks each OV device (iGPU / NPU / CPU) on a host and writes device_profile.json, step 1 toward automatic placement. See docs/perf/DEVICE_PROFILE.md.--target npu) shards, cascadia worker --device CPU --prefill-device NPU runs the compute-bound chunked prefill on the NPU and the bandwidth-bound decode on the CPU, sharing one host KV ring. See docs/perf/HYBRID_NPU_CPU.md.Cascadia does not daemonize itself, so it needs to be run under systemd / NSSM / launchd. See docs/deploy/ for a systemd unit template and Windows / macOS recipes. Cascadia handles SIGTERM cleanly.
Security: the HTTP API and inter-stage TCP relay are plaintext and unauthenticated. Bind only to trusted networks (LAN, loopback) or terminate TLS + auth at a reverse proxy in front of --api. See SECURITY.md for the threat model and built-in hardening.
Monitoring: stages started with --api serve Prometheus metrics at GET /metrics — request rate/latency, TTFT, inter-token latency, token throughput, cancellations, model load times, and inter-stage transport bytes. Metric inventory and example queries in docs/METRICS.md.
config.json not in <model dir>: ov-runtime reads the HF model config.json from the shard's tokenizer dir to derive rotary parameters. Older shard exports may not bundle config.json; copy it from the source model's HF cache (~/.cache/huggingface/hub/models--<repo>/snapshots/<sha>/config.json) into the shards root. Shards produced by cascadia shard bundle it automatically.
could not connect to downstream peer within timeout (engines wait 60 s): start the downstream worker first; check --listen on the downstream matches --next on the upstream and that the host's firewall allows the port.
Worker dies silently when SSH session closes: on Windows OpenSSH the child process is tied to the SSH parent. Run workers under systemd / NSSM / Task Scheduler in production.
cascadia doctor diagnoses most other environment/hardware issues.
| Doc | What's in it |
|---|---|
| QUICKSTART.md | 5-minute stub run → real inference |
| INSTALL.md | Full setup: OpenVINO SDK, GPU runtime, Docker |
| docs/CLI.md | Every command and flag |
| docs/SHARDING.md | Sharding flow + per-model-family support table |
| docs/ARCHITECTURE.md | Design decisions + crate responsibilities |
| docs/METRICS.md | Prometheus /metrics inventory + example queries |
| docs/engines/ | Per-engine deep dives |
| docs/architectures/ | Per-model-family export/support notes |
| docs/perf/ | Performance investigations and tuning |
| SECURITY.md | Threat model + vulnerability reporting |
See CONTRIBUTING.md for the build/test gate, crate layout, and commit conventions.
By participating you agree to the Code of Conduct.
Apache-2.0 (see LICENSE). Third-party attributions are in THIRD_PARTY_NOTICES.md.
Rust
78.4%
Python
18.0%
C++
2.0%