Local inference stack (DeepSeek-V4-flash, DeepSeek-V4.1-Flash:, qwen3.8-27b, GLM 5.3, qwen3.8-next, )
See the codeA compact Rust router that puts each OpenAI-compatible inference request on
the healthy GPU replica where it can do the least repeated work.
ramjet sits between your clients and replicated model servers. It keeps conversations and shared system prompts near warm cache state, then lets live load override affinity before one replica becomes a hotspot. Clients keep the same OpenAI API; engines need no ramjet-specific integration.
How ramjet and the engines behind it were tuned for each model, written up on the Helix blog. GLM-5.3 ran on an 8× H200 server; the others on our 8× RTX PRO 6000 server:
| GLM-5.3 (8× H200) | GLM-5.3-Flash | Qwen3.8-Flash-Next | Qwen3.8-27B | DeepSeek V4 |
|---|---|---|---|---|
| 48 coding agents from one server (29 Sep) | Part 1: getting day-zero serving to work (27 Aug) | On eight GPUs: what actually helped (27 Aug) | Chasing a 454 tok/s tweet (22 Aug) | SGLang vs DwarfStar vs vLLM+DSpark (14 Aug) |
| Ramjet vs NVIDIA Dynamo (30 Sep) | Running on 2, 4 or 8 GPUs (14 Sep) | One model, two speeds: smart routing (28 Aug) | Doubling throughput by reading a log line (23 Aug) | V4.1 Flash: encoder, Engram and KV cache (10 Sep) |
| Ran out of cache snapshots, not cache tokens (26 Sep) | Swift 1.5 cut thinking tokens (27 Sep) | A better lm_head, tested and shipped (25 Aug) |
| Reuse more | Queue less | Fail cleanly |
|---|---|---|
| Bounded prefix fingerprints find the replica most likely to reuse prior work. | Size-weighted reservations spread cold prefills and concurrent decodes. | Active health probes, retryable failover, and immediate disconnect cancellation keep capacity honest. |
The ordinary router is stateless, privacy-bounded, and deliberately useful without raw KV-cache events. Optional DSpark enforcement persists only opaque quarantine commitments so an LB restart cannot forget a bad EngineCore. The production path remains the proxy plus your existing OpenAI-compatible engines.
System overview to see how your node is doing:
And specific serving tab:
You can also just plug it into prometheus, /metrics API is available.
| Model · server | Result | Measured outcome |
|---|---|---|
| DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000 | Shared-app concurrency, load-blind → ramjet | 298 → 469 output tok/s · 1.57× |
| DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000 | Fresh 3-app × 4-session locality run | 82.5% cached prompt tokens |
| DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000 | Whole-box deterministic code, c24/max256 | 1,820–1,844 output tok/s |
| Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000 | Related request queued behind a long one, phase-aware load release | TTFT 2,496 → 287 ms · long request keeps 99.1% throughput |
| Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000 | Direct vLLM → same engine through ramjet | −0.03% at c1 · −0.24% at c16 |
| GLM-5.3-Flash · 2× TP4, 8× H200 | Coding-agent swarm, prefix routing → marginal affinity basis | 197/194 → 243/249 turns/min · TTFT p90 5.3–5.5 → 3.2–3.3s |
| GLM-5.3 · DP8 attention, 8× H200 | 32 / 48 coding agents, NVIDIA Dynamo 1.5.0 → ramjet, same engine | 73.0 → 88.9 / 77.3 → 99.5 turns/min · 85% → 92% cached |
These are workload results, not theoretical peaks. Reproduce the DeepSeek rows from RESULTS.md; inspect every accepted and rejected experiment in EXPERIMENTS.md.
All measured on node06 — 8× RTX PRO 6000 Blackwell; the engine topology is listed per row. The full-box column reports the best qualified saturation point recorded for that stack, not a shared concurrency level.
| Model | Served as | Decode @ c1 | Best qualified full-box throughput | Measured shape | Compose |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash (sparse MoE) | deepseek-v4-flash | 245.1 tok/s | 1,891.2 tok/s | 2× TP4, c24/max256 | deploy/dspark_0731 |
| Qwen3.8-27B FP8 (dense, vLLM) | qwen3.8-27b | 77 tok/s · 121 with MTP | 7,890.9 tok/s | 2× TP4, c256/max256, MTP off | deploy/qwen38_27b |
| Qwen3.8-27B NVFP4/BF16 head (dense, SGLang + DFlash2) | qwen3.8-27b | 153.3 tok/s greedy median · +7.5% matched A/B | Not yet requalified (former Inferact target: 7,882.6 tok/s) | 8× TP1, 208 slots, bf16 SSM, DFlash2 on | deploy/qwen38_27b |
| Qwen3.8-Flash-Next FP8 (sparse MoE, vLLM) | qwen3.8-flash-next | 113 tok/s · 202 with MTP3 | 3,340.5 tok/s | 2× TP4+EP, c64, MTP3 on both | deploy/qwen38_flash_next |
| GLM-5.3-Flash W4A16, FP8 experts (sparse MoE, SGLang) | glm-5.3-flash | 164.8 tok/s with EAGLE | 388.2 tok/s per 2-GPU replica | TP2, c4/max256; whole box not yet saturated | deploy/glm53_flash_sm120 |
No model — and neither Qwen3.8-27B stack — is simply better. Single-stream
decode is what an interactive user feels; the full-box figure is a capacity
landmark for a saturated agent fleet. These maxima come from separate
model-specific workloads, so they are not a matched head-to-head benchmark.
The vLLM row's saturation result has MTP off because speculation improves
low-concurrency latency but wastes rejected drafts once the batch saturates
the GPU. The SGLang row uses RadixArk's immutable
BF16-lm_head checkpoint. Its matched one-engine canary measured 153.3 tok/s
against 142.6 for the former Inferact target (+7.5%), with the same 7/8
objective answers and 20/25 deterministic agent-protocol cases. The smaller
target exposes 26 running slots and 582,246 KV tokens per engine: 208 slots
across the fleet. Full-box saturation has not yet been requalified on these
weights; the former Inferact target reached 7,882.6 tok/s, within 0.1% of the
vLLM reference. On that earlier SGLang stack, bf16 SSM state reduced c128 TTFT
p95 from 3.99s to 0.221s. The same
3-app × 4-session × 2-turn locality run measured 87.3% cached prompt
tokens, and 12 concurrent same-app requests spread across 7 of 8 engines
at 714 tok/s. Its cost is cold long-context prefill: a 196K-token first
turn pays ~57s of TTFT on one GPU, with prefix-cached follow-ups at 2–4s.
Qwen3.8-Flash-Next shows the same speculation trade-off: on 256-token outputs MTP3 adds 79% at c1 but only 7.5% at c32. The qualified pair therefore runs MTP3 on one engine and standard decoding on the other, and ramjet uses the requested output length to pick between them only once cache and load tie. Its full-box figure predates that split, with MTP3 on both engines. GLM-5.3-Flash runs on two GPUs per replica; its prefix cache is bounded by saved linear-attention states rather than KV tokens. Keeping two states per path instead of four, plus a 4 GB host tier, took a probe of 12 cyclic 20k-token sessions from 0/12 to 12/12 cached. The Kev stack adds a 0.8B decision model to the same server, sharing one Qwen GPU behind a second API profile: 77 ms p50 per short three-question request at c1 and 20.3 requests/s at c4, measured beside live traffic rather than saturated. Model profiles covers the sizing, sharding, and speculative-decoding trade-offs behind these numbers.
deploy/glm53_h200 runs the 753B FP8
checkpoint as one SGLang engine with eight data-parallel attention ranks, and
ramjet lists each rank as its own upstream (RJ_UPSTREAM_DP_RANKS). On a
simulated team of continuously working coding agents:
| Result | Measured outcome |
|---|---|
| 16 agents, SGLang rank placement → ramjet per-rank routing | 43.0 → 72.7 turns/min · 65.6% → 92.8% cached |
| 64 agents, plus a 32 GB host KV tier per rank | 45.9 → 95.1 turns/min |
| Capacity per server | 32–48 agents · 98–109 turns/min · TTFT p50 1.3–2.8s |
The blog post walks through each step, and Ramjet vs NVIDIA Dynamo compares the router against Dynamo 1.5.0's KV router on the same engine; the raw cells are in EXPERIMENTS.md (2026-09-29 and 2026-09-30).
For existing engines, the upstream list is normally the only setting you need:
services:
ramjet:
image: ghcr.io/helixml/ramjet:v0.7.0@sha256:dca028638314ca3171120532a075faaa70483e1494dd3d04bddc4db4eb88c01d
restart: unless-stopped
ports:
- "8000:8000" # OpenAI API + /health
- "9090:9090" # Prometheus
environment:
RJ_UPSTREAM: http://model-server-1:8000,http://model-server-2:8000
# RJ_UPSTREAM_TOKEN: ${MODEL_SERVER_API_KEY} # if required
docker compose up -d
curl --fail http://localhost:8000/health
The example pins a released image by immutable digest; see
CHANGELOG.md for what each version contains.
Safe defaults enable locality/load routing and keep tokenizer, raw KV-event,
exact-placement, and snapshot paths off. See the complete
configuration table, or start from the
eight-replica Compose stack currently
running in production. The
two-replica DeepSeek-V4-Flash stack
is the previous deployment, kept as a reviewed alternative and rollback
record.
Backend compatibility:
model-server-1andmodel-server-2are example Docker DNS names—replace them with your backends. The default router is not tied to vLLM: it forwards OpenAI-compatible APIs and health-checks each server withGET /v1/models. The opt-in/tokenize, KV-event, exact-routing, and snapshot research paths are currently designed for vLLM/DSpark.
score(replica) = min(prefix overlap, affinity cap) − α × live load
ramjet fingerprints only a bounded prefix, scores every healthy replica, and reserves load before forwarding. Warm state wins when it is valuable; idle capacity wins when reuse no longer pays for the queue. Score ties prefer the deeper raw overlap.
ok, degraded, and unhealthy readiness at GET /health.ramjet_* Prometheus metrics on port 9090.X-Ramjet-Upstream route correlation without leaking hosts.Exact tokenization, fenced KV indexes, authenticated snapshot companions, exact-placement canaries, and session-affinity shadow telemetry remain opt-in research surfaces. The session path cannot change placement. These paths fail closed and are not dependencies of ordinary serving.
Naming: the project was renamed from ramjet to ramjet. Settings now use the
RJ_*prefix and responses carryX-Ramjet-*headers; the retiredMD_*prefix is refused at startup rather than silently ignored, so a stale overlay fails loudly instead of running a differently tuned proxy. Theramjet_*metric names are deliberately unchanged so existing Grafana history keeps resolving.
| Task | Start here |
|---|---|
| Deploy or roll back | Docker Compose operator guide |
| Configure the router | Environment reference |
| Serve a different model | Model profiles |
| Understand the design | Architecture and routing model |
| Inspect current work | Roadmap |
Codex-compatible repo skills are included for repeatable node operations:
$deploy-ramjet,
$optimize-ramjet-node,
$load-test-ramjet-node,
and
$troubleshoot-ramjet-node.
cargo fmt --check
cargo test --locked
cargo clippy --locked --all-targets --all-features -- -D warnings
For privacy-safe production-shape validation, bench/agent_trace.py accepts
only numeric/enumerated trace shapes and synthesizes all request content. A
bounded /tokenize preflight adjusts for the active chat-template overhead;
authoritative response usage still enforces the token-density gate. See the
sovereign trace replay contract.
See AGENTS.md for the GPU-free inner loop, full release gate, and node06 benchmark contract.
Everything we have written about serving on the Helix blog, grouped by topic and newest first. The per-model table above picks from the same posts.
Routing with ramjet
GLM-5.3-Flash
Qwen3.8
DeepSeek
Hardware
Local inference stack (DeepSeek-V4-flash, DeepSeek-V4.1-Flash:, qwen3.8-27b, GLM 5.3, qwen3.8-next, )
See the codeA compact Rust router that puts each OpenAI-compatible inference request on
the healthy GPU replica where it can do the least repeated work.
ramjet sits between your clients and replicated model servers. It keeps conversations and shared system prompts near warm cache state, then lets live load override affinity before one replica becomes a hotspot. Clients keep the same OpenAI API; engines need no ramjet-specific integration.
How ramjet and the engines behind it were tuned for each model, written up on the Helix blog. GLM-5.3 ran on an 8× H200 server; the others on our 8× RTX PRO 6000 server:
| GLM-5.3 (8× H200) | GLM-5.3-Flash | Qwen3.8-Flash-Next | Qwen3.8-27B | DeepSeek V4 |
|---|---|---|---|---|
| 48 coding agents from one server (29 Sep) | Part 1: getting day-zero serving to work (27 Aug) | On eight GPUs: what actually helped (27 Aug) | Chasing a 454 tok/s tweet (22 Aug) | SGLang vs DwarfStar vs vLLM+DSpark (14 Aug) |
| Ramjet vs NVIDIA Dynamo (30 Sep) | Running on 2, 4 or 8 GPUs (14 Sep) | One model, two speeds: smart routing (28 Aug) | Doubling throughput by reading a log line (23 Aug) | V4.1 Flash: encoder, Engram and KV cache (10 Sep) |
| Ran out of cache snapshots, not cache tokens (26 Sep) | Swift 1.5 cut thinking tokens (27 Sep) | A better lm_head, tested and shipped (25 Aug) |
| Reuse more | Queue less | Fail cleanly |
|---|---|---|
| Bounded prefix fingerprints find the replica most likely to reuse prior work. | Size-weighted reservations spread cold prefills and concurrent decodes. | Active health probes, retryable failover, and immediate disconnect cancellation keep capacity honest. |
The ordinary router is stateless, privacy-bounded, and deliberately useful without raw KV-cache events. Optional DSpark enforcement persists only opaque quarantine commitments so an LB restart cannot forget a bad EngineCore. The production path remains the proxy plus your existing OpenAI-compatible engines.
System overview to see how your node is doing:
And specific serving tab:
You can also just plug it into prometheus, /metrics API is available.
| Model · server | Result | Measured outcome |
|---|---|---|
| DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000 | Shared-app concurrency, load-blind → ramjet | 298 → 469 output tok/s · 1.57× |
| DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000 | Fresh 3-app × 4-session locality run | 82.5% cached prompt tokens |
| DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000 | Whole-box deterministic code, c24/max256 | 1,820–1,844 output tok/s |
| Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000 | Related request queued behind a long one, phase-aware load release | TTFT 2,496 → 287 ms · long request keeps 99.1% throughput |
| Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000 | Direct vLLM → same engine through ramjet | −0.03% at c1 · −0.24% at c16 |
| GLM-5.3-Flash · 2× TP4, 8× H200 | Coding-agent swarm, prefix routing → marginal affinity basis | 197/194 → 243/249 turns/min · TTFT p90 5.3–5.5 → 3.2–3.3s |
| GLM-5.3 · DP8 attention, 8× H200 | 32 / 48 coding agents, NVIDIA Dynamo 1.5.0 → ramjet, same engine | 73.0 → 88.9 / 77.3 → 99.5 turns/min · 85% → 92% cached |
These are workload results, not theoretical peaks. Reproduce the DeepSeek rows from RESULTS.md; inspect every accepted and rejected experiment in EXPERIMENTS.md.
All measured on node06 — 8× RTX PRO 6000 Blackwell; the engine topology is listed per row. The full-box column reports the best qualified saturation point recorded for that stack, not a shared concurrency level.
| Model | Served as | Decode @ c1 | Best qualified full-box throughput | Measured shape | Compose |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash (sparse MoE) | deepseek-v4-flash | 245.1 tok/s | 1,891.2 tok/s | 2× TP4, c24/max256 | deploy/dspark_0731 |
| Qwen3.8-27B FP8 (dense, vLLM) | qwen3.8-27b | 77 tok/s · 121 with MTP | 7,890.9 tok/s | 2× TP4, c256/max256, MTP off | deploy/qwen38_27b |
| Qwen3.8-27B NVFP4/BF16 head (dense, SGLang + DFlash2) | qwen3.8-27b | 153.3 tok/s greedy median · +7.5% matched A/B | Not yet requalified (former Inferact target: 7,882.6 tok/s) | 8× TP1, 208 slots, bf16 SSM, DFlash2 on | deploy/qwen38_27b |
| Qwen3.8-Flash-Next FP8 (sparse MoE, vLLM) | qwen3.8-flash-next | 113 tok/s · 202 with MTP3 | 3,340.5 tok/s | 2× TP4+EP, c64, MTP3 on both | deploy/qwen38_flash_next |
| GLM-5.3-Flash W4A16, FP8 experts (sparse MoE, SGLang) | glm-5.3-flash | 164.8 tok/s with EAGLE | 388.2 tok/s per 2-GPU replica | TP2, c4/max256; whole box not yet saturated | deploy/glm53_flash_sm120 |
No model — and neither Qwen3.8-27B stack — is simply better. Single-stream
decode is what an interactive user feels; the full-box figure is a capacity
landmark for a saturated agent fleet. These maxima come from separate
model-specific workloads, so they are not a matched head-to-head benchmark.
The vLLM row's saturation result has MTP off because speculation improves
low-concurrency latency but wastes rejected drafts once the batch saturates
the GPU. The SGLang row uses RadixArk's immutable
BF16-lm_head checkpoint. Its matched one-engine canary measured 153.3 tok/s
against 142.6 for the former Inferact target (+7.5%), with the same 7/8
objective answers and 20/25 deterministic agent-protocol cases. The smaller
target exposes 26 running slots and 582,246 KV tokens per engine: 208 slots
across the fleet. Full-box saturation has not yet been requalified on these
weights; the former Inferact target reached 7,882.6 tok/s, within 0.1% of the
vLLM reference. On that earlier SGLang stack, bf16 SSM state reduced c128 TTFT
p95 from 3.99s to 0.221s. The same
3-app × 4-session × 2-turn locality run measured 87.3% cached prompt
tokens, and 12 concurrent same-app requests spread across 7 of 8 engines
at 714 tok/s. Its cost is cold long-context prefill: a 196K-token first
turn pays ~57s of TTFT on one GPU, with prefix-cached follow-ups at 2–4s.
Qwen3.8-Flash-Next shows the same speculation trade-off: on 256-token outputs MTP3 adds 79% at c1 but only 7.5% at c32. The qualified pair therefore runs MTP3 on one engine and standard decoding on the other, and ramjet uses the requested output length to pick between them only once cache and load tie. Its full-box figure predates that split, with MTP3 on both engines. GLM-5.3-Flash runs on two GPUs per replica; its prefix cache is bounded by saved linear-attention states rather than KV tokens. Keeping two states per path instead of four, plus a 4 GB host tier, took a probe of 12 cyclic 20k-token sessions from 0/12 to 12/12 cached. The Kev stack adds a 0.8B decision model to the same server, sharing one Qwen GPU behind a second API profile: 77 ms p50 per short three-question request at c1 and 20.3 requests/s at c4, measured beside live traffic rather than saturated. Model profiles covers the sizing, sharding, and speculative-decoding trade-offs behind these numbers.
deploy/glm53_h200 runs the 753B FP8
checkpoint as one SGLang engine with eight data-parallel attention ranks, and
ramjet lists each rank as its own upstream (RJ_UPSTREAM_DP_RANKS). On a
simulated team of continuously working coding agents:
| Result | Measured outcome |
|---|---|
| 16 agents, SGLang rank placement → ramjet per-rank routing | 43.0 → 72.7 turns/min · 65.6% → 92.8% cached |
| 64 agents, plus a 32 GB host KV tier per rank | 45.9 → 95.1 turns/min |
| Capacity per server | 32–48 agents · 98–109 turns/min · TTFT p50 1.3–2.8s |
The blog post walks through each step, and Ramjet vs NVIDIA Dynamo compares the router against Dynamo 1.5.0's KV router on the same engine; the raw cells are in EXPERIMENTS.md (2026-09-29 and 2026-09-30).
For existing engines, the upstream list is normally the only setting you need:
services:
ramjet:
image: ghcr.io/helixml/ramjet:v0.7.0@sha256:dca028638314ca3171120532a075faaa70483e1494dd3d04bddc4db4eb88c01d
restart: unless-stopped
ports:
- "8000:8000" # OpenAI API + /health
- "9090:9090" # Prometheus
environment:
RJ_UPSTREAM: http://model-server-1:8000,http://model-server-2:8000
# RJ_UPSTREAM_TOKEN: ${MODEL_SERVER_API_KEY} # if required
docker compose up -d
curl --fail http://localhost:8000/health
The example pins a released image by immutable digest; see
CHANGELOG.md for what each version contains.
Safe defaults enable locality/load routing and keep tokenizer, raw KV-event,
exact-placement, and snapshot paths off. See the complete
configuration table, or start from the
eight-replica Compose stack currently
running in production. The
two-replica DeepSeek-V4-Flash stack
is the previous deployment, kept as a reviewed alternative and rollback
record.
Backend compatibility:
model-server-1andmodel-server-2are example Docker DNS names—replace them with your backends. The default router is not tied to vLLM: it forwards OpenAI-compatible APIs and health-checks each server withGET /v1/models. The opt-in/tokenize, KV-event, exact-routing, and snapshot research paths are currently designed for vLLM/DSpark.
score(replica) = min(prefix overlap, affinity cap) − α × live load
ramjet fingerprints only a bounded prefix, scores every healthy replica, and reserves load before forwarding. Warm state wins when it is valuable; idle capacity wins when reuse no longer pays for the queue. Score ties prefer the deeper raw overlap.
ok, degraded, and unhealthy readiness at GET /health.ramjet_* Prometheus metrics on port 9090.X-Ramjet-Upstream route correlation without leaking hosts.Exact tokenization, fenced KV indexes, authenticated snapshot companions, exact-placement canaries, and session-affinity shadow telemetry remain opt-in research surfaces. The session path cannot change placement. These paths fail closed and are not dependencies of ordinary serving.
Naming: the project was renamed from ramjet to ramjet. Settings now use the
RJ_*prefix and responses carryX-Ramjet-*headers; the retiredMD_*prefix is refused at startup rather than silently ignored, so a stale overlay fails loudly instead of running a differently tuned proxy. Theramjet_*metric names are deliberately unchanged so existing Grafana history keeps resolving.
| Task | Start here |
|---|---|
| Deploy or roll back | Docker Compose operator guide |
| Configure the router | Environment reference |
| Serve a different model | Model profiles |
| Understand the design | Architecture and routing model |
| Inspect current work | Roadmap |
Codex-compatible repo skills are included for repeatable node operations:
$deploy-ramjet,
$optimize-ramjet-node,
$load-test-ramjet-node,
and
$troubleshoot-ramjet-node.
cargo fmt --check
cargo test --locked
cargo clippy --locked --all-targets --all-features -- -D warnings
For privacy-safe production-shape validation, bench/agent_trace.py accepts
only numeric/enumerated trace shapes and synthesizes all request content. A
bounded /tokenize preflight adjusts for the active chat-template overhead;
authoritative response usage still enforces the token-density gate. See the
sovereign trace replay contract.
See AGENTS.md for the GPU-free inner loop, full release gate, and node06 benchmark contract.
Everything we have written about serving on the Helix blog, grouped by topic and newest first. The per-model table above picks from the same posts.
Routing with ramjet
GLM-5.3-Flash
Qwen3.8
DeepSeek
Hardware