helixml/ramjet

Local inference stack (DeepSeek-V4-flash, DeepSeek-V4.1-Flash:, qwen3.8-27b, GLM 5.3, qwen3.8-next, )

Rust

5

277 commits

updated Oct 7, 2026

See the code

See what people are saying

SourceMessageScoreDate

Ramjet - mini altermative to nvidia dynamo (r/LocalLLM)

Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well. If you have dgx spark, multi mac…

0

Oct 7, 2026

Ramjet - mini altermative to nvidia dynamo (r/LocalLLaMA)

Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well. If you have dgx spark, multi mac…

1

Oct 7, 2026

README

ramjet

Warm intake. Balanced burn.

A compact Rust router that puts each OpenAI-compatible inference request on
the healthy GPU replica where it can do the least repeated work.

Rust 1.95 or newer

An incoming prompt is scored by ramjet and routed to the GPU replica with the best combination of reusable prefix and available capacity

ramjet sits between your clients and replicated model servers. It keeps conversations and shared system prompts near warm cache state, then lets live load override affinity before one replica becomes a hotspot. Clients keep the same OpenAI API; engines need no ramjet-specific integration.

What we wrote about each model

How ramjet and the engines behind it were tuned for each model, written up on the Helix blog. GLM-5.3 ran on an 8× H200 server; the others on our 8× RTX PRO 6000 server:

Why it exists

Reuse moreQueue lessFail cleanly
Bounded prefix fingerprints find the replica most likely to reuse prior work.Size-weighted reservations spread cold prefills and concurrent decodes.Active health probes, retryable failover, and immediate disconnect cancellation keep capacity honest.

The ordinary router is stateless, privacy-bounded, and deliberately useful without raw KV-cache events. Optional DSpark enforcement persists only opaque quarantine commitments so an LB restart cannot forget a bad EngineCore. The production path remains the proxy plus your existing OpenAI-compatible engines.

Built-in dashboard

System overview to see how your node is doing:

image

And specific serving tab:

image

You can also just plug it into prometheus, /metrics API is available.

Measured on real hardware

Model · serverResultMeasured outcome
DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000Shared-app concurrency, load-blind → ramjet298 → 469 output tok/s · 1.57×
DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000Fresh 3-app × 4-session locality run82.5% cached prompt tokens
DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000Whole-box deterministic code, c24/max2561,820–1,844 output tok/s
Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000Related request queued behind a long one, phase-aware load releaseTTFT 2,496 → 287 ms · long request keeps 99.1% throughput
Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000Direct vLLM → same engine through ramjet−0.03% at c1 · −0.24% at c16
GLM-5.3-Flash · 2× TP4, 8× H200Coding-agent swarm, prefix routing → marginal affinity basis197/194 → 243/249 turns/min · TTFT p90 5.3–5.5 → 3.2–3.3s
GLM-5.3 · DP8 attention, 8× H20032 / 48 coding agents, NVIDIA Dynamo 1.5.0 → ramjet, same engine73.0 → 88.9 / 77.3 → 99.5 turns/min · 85% → 92% cached

These are workload results, not theoretical peaks. Reproduce the DeepSeek rows from RESULTS.md; inspect every accepted and rejected experiment in EXPERIMENTS.md.

Models with a validated stack

All measured on node06 — 8× RTX PRO 6000 Blackwell; the engine topology is listed per row. The full-box column reports the best qualified saturation point recorded for that stack, not a shared concurrency level.

ModelServed asDecode @ c1Best qualified full-box throughputMeasured shapeCompose
DeepSeek-V4-Flash (sparse MoE)deepseek-v4-flash245.1 tok/s1,891.2 tok/s2× TP4, c24/max256deploy/dspark_0731
Qwen3.8-27B FP8 (dense, vLLM)qwen3.8-27b77 tok/s · 121 with MTP7,890.9 tok/s2× TP4, c256/max256, MTP offdeploy/qwen38_27b
Qwen3.8-27B NVFP4/BF16 head (dense, SGLang + DFlash2)qwen3.8-27b153.3 tok/s greedy median · +7.5% matched A/BNot yet requalified (former Inferact target: 7,882.6 tok/s)8× TP1, 208 slots, bf16 SSM, DFlash2 ondeploy/qwen38_27b
Qwen3.8-Flash-Next FP8 (sparse MoE, vLLM)qwen3.8-flash-next113 tok/s · 202 with MTP33,340.5 tok/s2× TP4+EP, c64, MTP3 on bothdeploy/qwen38_flash_next
GLM-5.3-Flash W4A16, FP8 experts (sparse MoE, SGLang)glm-5.3-flash164.8 tok/s with EAGLE388.2 tok/s per 2-GPU replicaTP2, c4/max256; whole box not yet saturateddeploy/glm53_flash_sm120

No model — and neither Qwen3.8-27B stack — is simply better. Single-stream decode is what an interactive user feels; the full-box figure is a capacity landmark for a saturated agent fleet. These maxima come from separate model-specific workloads, so they are not a matched head-to-head benchmark. The vLLM row's saturation result has MTP off because speculation improves low-concurrency latency but wastes rejected drafts once the batch saturates the GPU. The SGLang row uses RadixArk's immutable BF16-lm_head checkpoint. Its matched one-engine canary measured 153.3 tok/s against 142.6 for the former Inferact target (+7.5%), with the same 7/8 objective answers and 20/25 deterministic agent-protocol cases. The smaller target exposes 26 running slots and 582,246 KV tokens per engine: 208 slots across the fleet. Full-box saturation has not yet been requalified on these weights; the former Inferact target reached 7,882.6 tok/s, within 0.1% of the vLLM reference. On that earlier SGLang stack, bf16 SSM state reduced c128 TTFT p95 from 3.99s to 0.221s. The same 3-app × 4-session × 2-turn locality run measured 87.3% cached prompt tokens, and 12 concurrent same-app requests spread across 7 of 8 engines at 714 tok/s. Its cost is cold long-context prefill: a 196K-token first turn pays ~57s of TTFT on one GPU, with prefix-cached follow-ups at 2–4s.

Qwen3.8-Flash-Next shows the same speculation trade-off: on 256-token outputs MTP3 adds 79% at c1 but only 7.5% at c32. The qualified pair therefore runs MTP3 on one engine and standard decoding on the other, and ramjet uses the requested output length to pick between them only once cache and load tie. Its full-box figure predates that split, with MTP3 on both engines. GLM-5.3-Flash runs on two GPUs per replica; its prefix cache is bounded by saved linear-attention states rather than KV tokens. Keeping two states per path instead of four, plus a 4 GB host tier, took a probe of 12 cyclic 20k-token sessions from 0/12 to 12/12 cached. The Kev stack adds a 0.8B decision model to the same server, sharing one Qwen GPU behind a second API profile: 77 ms p50 per short three-question request at c1 and 20.3 requests/s at c4, measured beside live traffic rather than saturated. Model profiles covers the sizing, sharding, and speculative-decoding trade-offs behind these numbers.

Full GLM-5.3 on one 8× H200 server

deploy/glm53_h200 runs the 753B FP8 checkpoint as one SGLang engine with eight data-parallel attention ranks, and ramjet lists each rank as its own upstream (RJ_UPSTREAM_DP_RANKS). On a simulated team of continuously working coding agents:

ResultMeasured outcome
16 agents, SGLang rank placement → ramjet per-rank routing43.0 → 72.7 turns/min · 65.6% → 92.8% cached
64 agents, plus a 32 GB host KV tier per rank45.9 → 95.1 turns/min
Capacity per server32–48 agents · 98–109 turns/min · TTFT p50 1.3–2.8s

The blog post walks through each step, and Ramjet vs NVIDIA Dynamo compares the router against Dynamo 1.5.0's KV router on the same engine; the raw cells are in EXPERIMENTS.md (2026-09-29 and 2026-09-30).

Start in one minute

For existing engines, the upstream list is normally the only setting you need:

services:
  ramjet:
    image: ghcr.io/helixml/ramjet:v0.7.0@sha256:dca028638314ca3171120532a075faaa70483e1494dd3d04bddc4db4eb88c01d
    restart: unless-stopped
    ports:
      - "8000:8000" # OpenAI API + /health
      - "9090:9090" # Prometheus
    environment:
      RJ_UPSTREAM: http://model-server-1:8000,http://model-server-2:8000
      # RJ_UPSTREAM_TOKEN: ${MODEL_SERVER_API_KEY} # if required
docker compose up -d
curl --fail http://localhost:8000/health

The example pins a released image by immutable digest; see CHANGELOG.md for what each version contains. Safe defaults enable locality/load routing and keep tokenizer, raw KV-event, exact-placement, and snapshot paths off. See the complete configuration table, or start from the eight-replica Compose stack currently running in production. The two-replica DeepSeek-V4-Flash stack is the previous deployment, kept as a reviewed alternative and rollback record.

Backend compatibility: model-server-1 and model-server-2 are example Docker DNS names—replace them with your backends. The default router is not tied to vLLM: it forwards OpenAI-compatible APIs and health-checks each server with GET /v1/models. The opt-in /tokenize, KV-event, exact-routing, and snapshot research paths are currently designed for vLLM/DSpark.

The routing rule

score(replica) = min(prefix overlap, affinity cap) − α × live load

ramjet fingerprints only a bounded prefix, scores every healthy replica, and reserves load before forwarding. Warm state wins when it is valuable; idle capacity wins when reuse no longer pays for the queue. Score ties prefer the deeper raw overlap.

Production surface

  • OpenAI-compatible chat/completions, streaming, reasoning, and tool calls.
  • ok, degraded, and unhealthy readiness at GET /health.
  • Optional SHA-pinned model/template compatibility admission for engines that expose the atomic identity contract, with fail-closed per-replica recovery; the node06 guide includes an opt-in, no-extra-hop vLLM middleware candidate.
  • Optional DSpark reliability observation and sticky per-replica quarantine when active K5 acceptance collapses to zero across multiple complete metric windows; enforcement fsyncs an opaque EngineCore commitment and only a different compatibility-attested EngineCore can durably rearm it. A precommitted dirty marker keeps unresolved replicas fenced after an unclean LB exit or failed state mutation.
  • Stable ramjet_* Prometheus metrics on port 9090.
  • Opaque X-Ramjet-Upstream route correlation without leaking hosts.
  • Bounded memory, request sanitization, model metadata rewriting, and upstream cancellation when the client disappears.

Exact tokenization, fenced KV indexes, authenticated snapshot companions, exact-placement canaries, and session-affinity shadow telemetry remain opt-in research surfaces. The session path cannot change placement. These paths fail closed and are not dependencies of ordinary serving.

Naming: the project was renamed from ramjet to ramjet. Settings now use the RJ_* prefix and responses carry X-Ramjet-* headers; the retired MD_* prefix is refused at startup rather than silently ignored, so a stale overlay fails loudly instead of running a differently tuned proxy. The ramjet_* metric names are deliberately unchanged so existing Grafana history keeps resolving.

Operate it

TaskStart here
Deploy or roll backDocker Compose operator guide
Configure the routerEnvironment reference
Serve a different modelModel profiles
Understand the designArchitecture and routing model
Inspect current workRoadmap

Codex-compatible repo skills are included for repeatable node operations: $deploy-ramjet, $optimize-ramjet-node, $load-test-ramjet-node, and $troubleshoot-ramjet-node.

Develop

cargo fmt --check
cargo test --locked
cargo clippy --locked --all-targets --all-features -- -D warnings
Privacy-safe production-shape replay

For privacy-safe production-shape validation, bench/agent_trace.py accepts only numeric/enumerated trace shapes and synthesizes all request content. A bounded /tokenize preflight adjusts for the active chat-template overhead; authoritative response usage still enforces the token-density gate. See the sovereign trace replay contract.

See AGENTS.md for the GPU-free inner loop, full release gate, and node06 benchmark contract.

Resources

Everything we have written about serving on the Helix blog, grouped by topic and newest first. The per-model table above picks from the same posts.

Routing with ramjet

GLM-5.3-Flash

Qwen3.8

DeepSeek

Hardware

License

Apache-2.0.

helixml/ramjet

Local inference stack (DeepSeek-V4-flash, DeepSeek-V4.1-Flash:, qwen3.8-27b, GLM 5.3, qwen3.8-next, )

Rust

5

277 commits

updated Oct 7, 2026

See the code

See what people are saying

SourceMessageScoreDate

Ramjet - mini altermative to nvidia dynamo (r/LocalLLM)

Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well. If you have dgx spark, multi mac…

0

Oct 7, 2026

Ramjet - mini altermative to nvidia dynamo (r/LocalLLaMA)

Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well. If you have dgx spark, multi mac…

1

Oct 7, 2026

README

ramjet

Warm intake. Balanced burn.

A compact Rust router that puts each OpenAI-compatible inference request on
the healthy GPU replica where it can do the least repeated work.

Rust 1.95 or newer

An incoming prompt is scored by ramjet and routed to the GPU replica with the best combination of reusable prefix and available capacity

ramjet sits between your clients and replicated model servers. It keeps conversations and shared system prompts near warm cache state, then lets live load override affinity before one replica becomes a hotspot. Clients keep the same OpenAI API; engines need no ramjet-specific integration.

What we wrote about each model

How ramjet and the engines behind it were tuned for each model, written up on the Helix blog. GLM-5.3 ran on an 8× H200 server; the others on our 8× RTX PRO 6000 server:

Why it exists

Reuse moreQueue lessFail cleanly
Bounded prefix fingerprints find the replica most likely to reuse prior work.Size-weighted reservations spread cold prefills and concurrent decodes.Active health probes, retryable failover, and immediate disconnect cancellation keep capacity honest.

The ordinary router is stateless, privacy-bounded, and deliberately useful without raw KV-cache events. Optional DSpark enforcement persists only opaque quarantine commitments so an LB restart cannot forget a bad EngineCore. The production path remains the proxy plus your existing OpenAI-compatible engines.

Built-in dashboard

System overview to see how your node is doing:

image

And specific serving tab:

image

You can also just plug it into prometheus, /metrics API is available.

Measured on real hardware

Model · serverResultMeasured outcome
DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000Shared-app concurrency, load-blind → ramjet298 → 469 output tok/s · 1.57×
DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000Fresh 3-app × 4-session locality run82.5% cached prompt tokens
DeepSeek-V4-Flash · 2× TP4, 8× RTX PRO 6000Whole-box deterministic code, c24/max2561,820–1,844 output tok/s
Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000Related request queued behind a long one, phase-aware load releaseTTFT 2,496 → 287 ms · long request keeps 99.1% throughput
Qwen3.8-Flash-Next · 2× TP4, 8× RTX PRO 6000Direct vLLM → same engine through ramjet−0.03% at c1 · −0.24% at c16
GLM-5.3-Flash · 2× TP4, 8× H200Coding-agent swarm, prefix routing → marginal affinity basis197/194 → 243/249 turns/min · TTFT p90 5.3–5.5 → 3.2–3.3s
GLM-5.3 · DP8 attention, 8× H20032 / 48 coding agents, NVIDIA Dynamo 1.5.0 → ramjet, same engine73.0 → 88.9 / 77.3 → 99.5 turns/min · 85% → 92% cached

These are workload results, not theoretical peaks. Reproduce the DeepSeek rows from RESULTS.md; inspect every accepted and rejected experiment in EXPERIMENTS.md.

Models with a validated stack

All measured on node06 — 8× RTX PRO 6000 Blackwell; the engine topology is listed per row. The full-box column reports the best qualified saturation point recorded for that stack, not a shared concurrency level.

ModelServed asDecode @ c1Best qualified full-box throughputMeasured shapeCompose
DeepSeek-V4-Flash (sparse MoE)deepseek-v4-flash245.1 tok/s1,891.2 tok/s2× TP4, c24/max256deploy/dspark_0731
Qwen3.8-27B FP8 (dense, vLLM)qwen3.8-27b77 tok/s · 121 with MTP7,890.9 tok/s2× TP4, c256/max256, MTP offdeploy/qwen38_27b
Qwen3.8-27B NVFP4/BF16 head (dense, SGLang + DFlash2)qwen3.8-27b153.3 tok/s greedy median · +7.5% matched A/BNot yet requalified (former Inferact target: 7,882.6 tok/s)8× TP1, 208 slots, bf16 SSM, DFlash2 ondeploy/qwen38_27b
Qwen3.8-Flash-Next FP8 (sparse MoE, vLLM)qwen3.8-flash-next113 tok/s · 202 with MTP33,340.5 tok/s2× TP4+EP, c64, MTP3 on bothdeploy/qwen38_flash_next
GLM-5.3-Flash W4A16, FP8 experts (sparse MoE, SGLang)glm-5.3-flash164.8 tok/s with EAGLE388.2 tok/s per 2-GPU replicaTP2, c4/max256; whole box not yet saturateddeploy/glm53_flash_sm120

No model — and neither Qwen3.8-27B stack — is simply better. Single-stream decode is what an interactive user feels; the full-box figure is a capacity landmark for a saturated agent fleet. These maxima come from separate model-specific workloads, so they are not a matched head-to-head benchmark. The vLLM row's saturation result has MTP off because speculation improves low-concurrency latency but wastes rejected drafts once the batch saturates the GPU. The SGLang row uses RadixArk's immutable BF16-lm_head checkpoint. Its matched one-engine canary measured 153.3 tok/s against 142.6 for the former Inferact target (+7.5%), with the same 7/8 objective answers and 20/25 deterministic agent-protocol cases. The smaller target exposes 26 running slots and 582,246 KV tokens per engine: 208 slots across the fleet. Full-box saturation has not yet been requalified on these weights; the former Inferact target reached 7,882.6 tok/s, within 0.1% of the vLLM reference. On that earlier SGLang stack, bf16 SSM state reduced c128 TTFT p95 from 3.99s to 0.221s. The same 3-app × 4-session × 2-turn locality run measured 87.3% cached prompt tokens, and 12 concurrent same-app requests spread across 7 of 8 engines at 714 tok/s. Its cost is cold long-context prefill: a 196K-token first turn pays ~57s of TTFT on one GPU, with prefix-cached follow-ups at 2–4s.

Qwen3.8-Flash-Next shows the same speculation trade-off: on 256-token outputs MTP3 adds 79% at c1 but only 7.5% at c32. The qualified pair therefore runs MTP3 on one engine and standard decoding on the other, and ramjet uses the requested output length to pick between them only once cache and load tie. Its full-box figure predates that split, with MTP3 on both engines. GLM-5.3-Flash runs on two GPUs per replica; its prefix cache is bounded by saved linear-attention states rather than KV tokens. Keeping two states per path instead of four, plus a 4 GB host tier, took a probe of 12 cyclic 20k-token sessions from 0/12 to 12/12 cached. The Kev stack adds a 0.8B decision model to the same server, sharing one Qwen GPU behind a second API profile: 77 ms p50 per short three-question request at c1 and 20.3 requests/s at c4, measured beside live traffic rather than saturated. Model profiles covers the sizing, sharding, and speculative-decoding trade-offs behind these numbers.

Full GLM-5.3 on one 8× H200 server

deploy/glm53_h200 runs the 753B FP8 checkpoint as one SGLang engine with eight data-parallel attention ranks, and ramjet lists each rank as its own upstream (RJ_UPSTREAM_DP_RANKS). On a simulated team of continuously working coding agents:

ResultMeasured outcome
16 agents, SGLang rank placement → ramjet per-rank routing43.0 → 72.7 turns/min · 65.6% → 92.8% cached
64 agents, plus a 32 GB host KV tier per rank45.9 → 95.1 turns/min
Capacity per server32–48 agents · 98–109 turns/min · TTFT p50 1.3–2.8s

The blog post walks through each step, and Ramjet vs NVIDIA Dynamo compares the router against Dynamo 1.5.0's KV router on the same engine; the raw cells are in EXPERIMENTS.md (2026-09-29 and 2026-09-30).

Start in one minute

For existing engines, the upstream list is normally the only setting you need:

services:
  ramjet:
    image: ghcr.io/helixml/ramjet:v0.7.0@sha256:dca028638314ca3171120532a075faaa70483e1494dd3d04bddc4db4eb88c01d
    restart: unless-stopped
    ports:
      - "8000:8000" # OpenAI API + /health
      - "9090:9090" # Prometheus
    environment:
      RJ_UPSTREAM: http://model-server-1:8000,http://model-server-2:8000
      # RJ_UPSTREAM_TOKEN: ${MODEL_SERVER_API_KEY} # if required
docker compose up -d
curl --fail http://localhost:8000/health

The example pins a released image by immutable digest; see CHANGELOG.md for what each version contains. Safe defaults enable locality/load routing and keep tokenizer, raw KV-event, exact-placement, and snapshot paths off. See the complete configuration table, or start from the eight-replica Compose stack currently running in production. The two-replica DeepSeek-V4-Flash stack is the previous deployment, kept as a reviewed alternative and rollback record.

Backend compatibility: model-server-1 and model-server-2 are example Docker DNS names—replace them with your backends. The default router is not tied to vLLM: it forwards OpenAI-compatible APIs and health-checks each server with GET /v1/models. The opt-in /tokenize, KV-event, exact-routing, and snapshot research paths are currently designed for vLLM/DSpark.

The routing rule

score(replica) = min(prefix overlap, affinity cap) − α × live load

ramjet fingerprints only a bounded prefix, scores every healthy replica, and reserves load before forwarding. Warm state wins when it is valuable; idle capacity wins when reuse no longer pays for the queue. Score ties prefer the deeper raw overlap.

Production surface

  • OpenAI-compatible chat/completions, streaming, reasoning, and tool calls.
  • ok, degraded, and unhealthy readiness at GET /health.
  • Optional SHA-pinned model/template compatibility admission for engines that expose the atomic identity contract, with fail-closed per-replica recovery; the node06 guide includes an opt-in, no-extra-hop vLLM middleware candidate.
  • Optional DSpark reliability observation and sticky per-replica quarantine when active K5 acceptance collapses to zero across multiple complete metric windows; enforcement fsyncs an opaque EngineCore commitment and only a different compatibility-attested EngineCore can durably rearm it. A precommitted dirty marker keeps unresolved replicas fenced after an unclean LB exit or failed state mutation.
  • Stable ramjet_* Prometheus metrics on port 9090.
  • Opaque X-Ramjet-Upstream route correlation without leaking hosts.
  • Bounded memory, request sanitization, model metadata rewriting, and upstream cancellation when the client disappears.

Exact tokenization, fenced KV indexes, authenticated snapshot companions, exact-placement canaries, and session-affinity shadow telemetry remain opt-in research surfaces. The session path cannot change placement. These paths fail closed and are not dependencies of ordinary serving.

Naming: the project was renamed from ramjet to ramjet. Settings now use the RJ_* prefix and responses carry X-Ramjet-* headers; the retired MD_* prefix is refused at startup rather than silently ignored, so a stale overlay fails loudly instead of running a differently tuned proxy. The ramjet_* metric names are deliberately unchanged so existing Grafana history keeps resolving.

Operate it

TaskStart here
Deploy or roll backDocker Compose operator guide
Configure the routerEnvironment reference
Serve a different modelModel profiles
Understand the designArchitecture and routing model
Inspect current workRoadmap

Codex-compatible repo skills are included for repeatable node operations: $deploy-ramjet, $optimize-ramjet-node, $load-test-ramjet-node, and $troubleshoot-ramjet-node.

Develop

cargo fmt --check
cargo test --locked
cargo clippy --locked --all-targets --all-features -- -D warnings
Privacy-safe production-shape replay

For privacy-safe production-shape validation, bench/agent_trace.py accepts only numeric/enumerated trace shapes and synthesizes all request content. A bounded /tokenize preflight adjusts for the active chat-template overhead; authoritative response usage still enforces the token-density gate. See the sovereign trace replay contract.

See AGENTS.md for the GPU-free inner loop, full release gate, and node06 benchmark contract.

Resources

Everything we have written about serving on the Helix blog, grouped by topic and newest first. The per-model table above picks from the same posts.

Routing with ramjet

GLM-5.3-Flash

Qwen3.8

DeepSeek

Hardware

License

Apache-2.0.