Ignis is a deliberately specialized inference engine: one model family, one GPU class.
Rust
6
922 commits
updated Oct 4, 2026
A single-GPU inference engine for Qwen3.8-27B, built for the load an agent makes.
OpenAI-compatible HTTP API · Rust core · C++/CUDA kernel leaf · one Blackwell card
| 367 tok/s one lane, coding prompts | 1,065 tok/s eight lanes, aggregate | 7.11x KV capacity of BF16 | ~36 ms a decision, zero tokens decoded | 512K context, YaRN x2 |
Measured on one RTX 5090, 64 GB DDR4-3200 MT/s, PCIe 3.0. How, and against what: How fast.
Ignis is a deliberately specialized engine: one model family, one class of
card. It gives up generality and takes back speed. It loads an NVFP4
Qwen3.8-27B onto a single Blackwell card — an RTX 5090 or an RTX PRO
6000, both SM120a (why not DGX Spark?) — serves streaming and non-streaming
completions over the OpenAI v1 API, and is shaped around the workload a
developer actually produces — one main agent plus a handful of subagents
hitting the same card at once. Eight resident decode lanes run as one
batch-wide round; the paged KV cache is budgeted in bytes rather than in
sequences; conversation state survives the request that built it; images are
evidence, not an afterthought; and a decision can be answered without
generating a single token. A Playground, a live Monitor and Prometheus
metrics ship in the binary.
Rust owns everything above the step — scheduling, admission, KV accounting, serving. The forward pass and all GPU compute live in a C++/CUDA static library behind a flat, device-resident step-level C ABI. The engine is its own dogfood target, and partly there already: it serves some of the coding agents that build it, and the rest of that is what the remaining work is for.
Take a build from Releases — a
Windows .zip, a Linux .tar.gz, or the linux/amd64 container image on
ghcr.io/gpillon/ignis. Run it on a Blackwell card:
ignis-server --max-context 524288 --rope-scaling yarn:2 \
--spec dflash2 --draft-tokens 7 --vision --metrics
No model yet? It offers to fetch one. Those flags turn on what the bare
ignis-server leaves off: the 512K context (it starts at 40,960), DFlash2
speculative decoding, image input, and the Prometheus metrics the Playground's
Monitor reads.
Then the OpenAI API is on http://127.0.0.1:8000/v1, the Playground and its Monitor on http://127.0.0.1:8000/ui/, and Prometheus on http://127.0.0.1:9464/metrics. The weights are on Hugging Face in two flavours, the standard one and an uncensored one: The weights. Everything else — the container, building from a checkout, every flag — is Quick start and docs/user.
Ignis on one RTX 5090 — Qwen3.8-27B NVFP4, hq-e8-2b KV, DFlash2 speculative
decoding with 7 draft tokens, greedy, short prompts. The served context is
512K tokens: the checkpoint's native 262,144 stretched by YaRN x2, as the
TL;DR flags and make set it. These runs used a 262,144-token maximum:
| load | throughput | source |
|---|---|---|
| one lane, eight coding prompts (write, edit, explain), geometric mean | 367 tok/s | upstream quick wins |
| one lane, predictable code / free prose | 491 / 167 tok/s | retained slots on the host |
| eight lanes at once, aggregate | 1,065 tok/s | upstream quick wins |
Other engines, same card, same model — single-stream decode as their authors published it:
| engine | weights and speculation | single-stream decode | source |
|---|---|---|---|
vLLM (main g41f179b57) | NVFP4, FP8 KV, DFlash2 drafter with 6 tokens | ~304 tok/s code at 8K, 221 at 60K | HF discussion #132 |
| llama.cpp | NVFP4-MTP GGUF, q8_0 KV, MTP with 4 tokens | 122–142 tok/s | DataCamp, 2026-08-25 |
| llama.cpp | MTP | 70.5 tok/s | HF discussion #132 |
Those rows are other people's runs: their prompts, their drafters, their context lengths. They are not a same-session benchmark, so read them as an order of magnitude, not a ratio. Prefill is where ignis is not ahead: vLLM's published ~12,500 tok/s at 60K is above ignis's measured 6,600 tok/s on a cold 46K prompt and 4,600 tok/s on a cold 105K one.
The one same-session comparison is against NInfer, the engine whose kernels ignis started from: same artifact, same card, one session, two process launches per side, before ignis had speculative decoding. ignis is 1.06x the reference at one lane and 3.07x at four, and 3% behind it on inter-token latency p95 (hq vs BF16 live/live).
Every number below is a measurement on a 5090.
--spec dflash2 --draft-tokens 7 runs draft-and-verify rounds against it.interactive or agent — through the
class extension field or an @<lane> suffix on the model name. The tag
drives protection, backfill priority and eviction order, so a burst of
subagents fills the lanes a foreground conversation is not using instead of
evicting it./v1/decide)POST /v1/decide prefills the
evidence once and reads the answer out of the logits at a single position.noul (yes/no), choice, score, and — one
constrained digit per step — number, point, box.point reads a click target off a 4096px screenshot to within 84 px worst
case, inside the button on every scene tested./v1/systemone, so a Jev client reaches
this engine by changing the URL and nothing else.--vision loads the tower and reserves its workspace; images arrive inline
as data URIs or as URLs the server fetches, prepares and caches.--rope-scaling yarn:F
rescales that envelope, yarn:4 putting the ceiling at 1,048,576
(factor up to 64). What a long-context probe must ask is a question about
relative position — a literal needle at 320K is recalled with the flag and
without it./ui/ — chat with tools and subagents, parallel
sessions, image input, a Decide tab and a live Monitor. Built into the
server, not a separate service.--metrics), with the VRAM plan, retained
state, KV pool occupancy, decision counters and request lifecycle.--api-key and --expose — a keyed public URL through a quick Cloudflare
tunnel, key always required.linux/amd64 container image on ghcr.io, and
release archives cut from the same build.Agents running in parallel, each on its own lane, with the engine's own timings beside every reply:

A decision over prose — the winner, the whole distribution, and a score read as levels — answered out of one prefill with nothing generated:

The same endpoint over an image: a point, a box, and a yes/no, each a digit at a time, with the model's self-reported uncertainty on every axis:

The Monitor scrapes the engine's own metrics: throughput, lane occupancy, latency quantiles:

...and what the load actually reserved, against the constants that bound it:

Two images of the same model, both on Hugging Face, both 18.07 GiB, both carrying the DFlash2 drafter:
| Hugging Face | file | |
|---|---|---|
| Standard | gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer | qwen3_8_27b_nvfp4full-v2.ninfer |
| Uncensored | gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer | qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer |
The server fetches the standard one by itself the first time it starts without a model, and checks its size and SHA-256 before using it. To download either one by hand:
hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer --local-dir models
hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer --local-dir models
The uncensored image is the standard one with the
huihui-ai abliteration
applied: 1,255 of its 1,325 objects are byte-identical, and only the 70 matrices
the abliteration changed are re-encoded. The server never downloads it on its
own; start it with --artifact ./models/qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer,
or make dev UNCENSORED=1 from a checkout. It does not refuse, so put the
guardrails in the application or tool layer and do not expose it to untrusted
users.
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models \
ghcr.io/gpillon/ignis:0.1.2 --model-download-path /models
(docker: --gpus all in place of --device.) With no model in that
directory the server fetches one and verifies it. The Playground is then at
http://127.0.0.1:8000/ui/.
make doctor # toolchain, rust target, artifact, web deps
make dev # build web + kernel + GPU server, then run it
make dev VISION=1 # the same, with the vision tower loaded for image input
make dev UNCENSORED=1 # the same, on the uncensored (abliterated) weights
make mock # the same with no GPU, no kernel, no artifact
make alone lists every target; make config prints the exact server command a
run will use.
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Hello"}],"max_tokens":256}'
Full flag table, API surface, model handling, container and release detail: docs/user.
Every field of the OpenAI request bodies is honoured, accepted as inert (a
value asking for what the server already does, like n: 1 or store: false),
or refused with a 400 naming it in param — never dropped in silence. The
table is in the API reference at /v1/docs/.
Choices, not gaps — so nobody goes looking for them:
/v1/decide, read
off the loaded model's own readout./v1/completions. A raw prompt skips the chat template the model was
tuned on, and the turn boundaries reuse is keyed on; chat completions and
responses cover it./v1/batches. An offline queue is a storage service, and this server
holds no request past its connection: send the requests at once, and the
scheduler batches them.Two parts, in this order: the C++/CUDA kernel leaf, then the Rust workspace that links it.
nvcc, target SM120a); CMake + Ninja; Node.js for the Playground.kernel/build.ps1 on Windows, kernel/build.sh on Linux.
Both configure and build into kernel/build/, which crates/*/build.rs link.cargo build --release -p ignis-server --features cuda for
the real engine; without --features cuda (or without an artifact) the server
runs a deterministic CPU-only mock, which is how protocol and scheduler work
gets done with no card.| Path | What lives there |
|---|---|
crates/core/ | Scheduler, admission, paged KV accounting, request state, host tier. |
crates/artifact/ | The .ninfer reader: reader, binder, materializer. |
crates/runtime/ | Safe wrapper over the step ABI (CudaLeaf, decode graphs). |
crates/server/ | HTTP, OpenAI schemas, decisions, metrics, telemetry. |
crates/logging/ | Structured logging: tracing layers, hotpath lint, trace context. |
crates/bench/ | Trace-replay harness, gate and canary runner. |
crates/vendor/ | The vendoring tool: manifest, hashes, patch records. |
kernel/ | The C++/CUDA leaf: the program and the vendored ops (CMake + nvcc). |
web/ | The Playground: React + Vite, embedded into the server at build time. |
make wraps all of it. Tests: cargo test is workspace-wide, CPU-only and
never touches the GPU; GPU work is checked by a separate explicit profile that
fails rather than skips when the card is busy. The 5090 fits one run at a
time — check make gpu-status before starting anything on it.
docs/user/ | Running it: flags, API, models, container, releases, telemetry. |
docs/adr/ | Every architectural decision, and why it was taken. |
docs/findings/ | Durable, evidence-backed measurements. |
docs/agents/ | Conventions for agents working in this repo. |
CONTEXT.md | The glossary. One vocabulary, one meaning per term. |
The OpenAPI description of the HTTP surface is served at /v1/openapi.json,
with a browsable reference at /v1/docs/.
Ignis descends from NInfer, a lineage
of Windows-oriented local-inference forks, and does not hide it. The pinned
reference is the fork gpillon/ninfer,
itself downstream of cometkim/ninfer
(kernel-perf, hyperquant KV, NVFP4-full, the original DFlash2 port) →
natpate/ninfer-windows (the
Windows port) → Neroued/ninfer (the
original engine). The .ninfer model artifact is NInfer's format,
and the kernel leaf vendors NInfer's CUDA ops verbatim under a manifest that
pins the reference commit and every file's content hash — a port claim you can
diff. What is ours is the layer above them: the forward pass, the sequence
state, the step ABI, the scheduler, and by now the kernels that measurement
asked us to rewrite.
Ignis is not a port of NInfer. It is a different architecture that starts from proven kernel work.
With thanks to the NInfer project and its contributors:
Apache License 2.0.
Rust
65.1%
TypeScript
11.6%
Cuda
8.8%
Python
6.7%
C++
4.4%
C
1.2%
Ignis is a deliberately specialized inference engine: one model family, one GPU class.
Rust
6
922 commits
updated Oct 4, 2026
A single-GPU inference engine for Qwen3.8-27B, built for the load an agent makes.
OpenAI-compatible HTTP API · Rust core · C++/CUDA kernel leaf · one Blackwell card
| 367 tok/s one lane, coding prompts | 1,065 tok/s eight lanes, aggregate | 7.11x KV capacity of BF16 | ~36 ms a decision, zero tokens decoded | 512K context, YaRN x2 |
Measured on one RTX 5090, 64 GB DDR4-3200 MT/s, PCIe 3.0. How, and against what: How fast.
Ignis is a deliberately specialized engine: one model family, one class of
card. It gives up generality and takes back speed. It loads an NVFP4
Qwen3.8-27B onto a single Blackwell card — an RTX 5090 or an RTX PRO
6000, both SM120a (why not DGX Spark?) — serves streaming and non-streaming
completions over the OpenAI v1 API, and is shaped around the workload a
developer actually produces — one main agent plus a handful of subagents
hitting the same card at once. Eight resident decode lanes run as one
batch-wide round; the paged KV cache is budgeted in bytes rather than in
sequences; conversation state survives the request that built it; images are
evidence, not an afterthought; and a decision can be answered without
generating a single token. A Playground, a live Monitor and Prometheus
metrics ship in the binary.
Rust owns everything above the step — scheduling, admission, KV accounting, serving. The forward pass and all GPU compute live in a C++/CUDA static library behind a flat, device-resident step-level C ABI. The engine is its own dogfood target, and partly there already: it serves some of the coding agents that build it, and the rest of that is what the remaining work is for.
Take a build from Releases — a
Windows .zip, a Linux .tar.gz, or the linux/amd64 container image on
ghcr.io/gpillon/ignis. Run it on a Blackwell card:
ignis-server --max-context 524288 --rope-scaling yarn:2 \
--spec dflash2 --draft-tokens 7 --vision --metrics
No model yet? It offers to fetch one. Those flags turn on what the bare
ignis-server leaves off: the 512K context (it starts at 40,960), DFlash2
speculative decoding, image input, and the Prometheus metrics the Playground's
Monitor reads.
Then the OpenAI API is on http://127.0.0.1:8000/v1, the Playground and its Monitor on http://127.0.0.1:8000/ui/, and Prometheus on http://127.0.0.1:9464/metrics. The weights are on Hugging Face in two flavours, the standard one and an uncensored one: The weights. Everything else — the container, building from a checkout, every flag — is Quick start and docs/user.
Ignis on one RTX 5090 — Qwen3.8-27B NVFP4, hq-e8-2b KV, DFlash2 speculative
decoding with 7 draft tokens, greedy, short prompts. The served context is
512K tokens: the checkpoint's native 262,144 stretched by YaRN x2, as the
TL;DR flags and make set it. These runs used a 262,144-token maximum:
| load | throughput | source |
|---|---|---|
| one lane, eight coding prompts (write, edit, explain), geometric mean | 367 tok/s | upstream quick wins |
| one lane, predictable code / free prose | 491 / 167 tok/s | retained slots on the host |
| eight lanes at once, aggregate | 1,065 tok/s | upstream quick wins |
Other engines, same card, same model — single-stream decode as their authors published it:
| engine | weights and speculation | single-stream decode | source |
|---|---|---|---|
vLLM (main g41f179b57) | NVFP4, FP8 KV, DFlash2 drafter with 6 tokens | ~304 tok/s code at 8K, 221 at 60K | HF discussion #132 |
| llama.cpp | NVFP4-MTP GGUF, q8_0 KV, MTP with 4 tokens | 122–142 tok/s | DataCamp, 2026-08-25 |
| llama.cpp | MTP | 70.5 tok/s | HF discussion #132 |
Those rows are other people's runs: their prompts, their drafters, their context lengths. They are not a same-session benchmark, so read them as an order of magnitude, not a ratio. Prefill is where ignis is not ahead: vLLM's published ~12,500 tok/s at 60K is above ignis's measured 6,600 tok/s on a cold 46K prompt and 4,600 tok/s on a cold 105K one.
The one same-session comparison is against NInfer, the engine whose kernels ignis started from: same artifact, same card, one session, two process launches per side, before ignis had speculative decoding. ignis is 1.06x the reference at one lane and 3.07x at four, and 3% behind it on inter-token latency p95 (hq vs BF16 live/live).
Every number below is a measurement on a 5090.
--spec dflash2 --draft-tokens 7 runs draft-and-verify rounds against it.interactive or agent — through the
class extension field or an @<lane> suffix on the model name. The tag
drives protection, backfill priority and eviction order, so a burst of
subagents fills the lanes a foreground conversation is not using instead of
evicting it./v1/decide)POST /v1/decide prefills the
evidence once and reads the answer out of the logits at a single position.noul (yes/no), choice, score, and — one
constrained digit per step — number, point, box.point reads a click target off a 4096px screenshot to within 84 px worst
case, inside the button on every scene tested./v1/systemone, so a Jev client reaches
this engine by changing the URL and nothing else.--vision loads the tower and reserves its workspace; images arrive inline
as data URIs or as URLs the server fetches, prepares and caches.--rope-scaling yarn:F
rescales that envelope, yarn:4 putting the ceiling at 1,048,576
(factor up to 64). What a long-context probe must ask is a question about
relative position — a literal needle at 320K is recalled with the flag and
without it./ui/ — chat with tools and subagents, parallel
sessions, image input, a Decide tab and a live Monitor. Built into the
server, not a separate service.--metrics), with the VRAM plan, retained
state, KV pool occupancy, decision counters and request lifecycle.--api-key and --expose — a keyed public URL through a quick Cloudflare
tunnel, key always required.linux/amd64 container image on ghcr.io, and
release archives cut from the same build.Agents running in parallel, each on its own lane, with the engine's own timings beside every reply:

A decision over prose — the winner, the whole distribution, and a score read as levels — answered out of one prefill with nothing generated:

The same endpoint over an image: a point, a box, and a yes/no, each a digit at a time, with the model's self-reported uncertainty on every axis:

The Monitor scrapes the engine's own metrics: throughput, lane occupancy, latency quantiles:

...and what the load actually reserved, against the constants that bound it:

Two images of the same model, both on Hugging Face, both 18.07 GiB, both carrying the DFlash2 drafter:
| Hugging Face | file | |
|---|---|---|
| Standard | gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer | qwen3_8_27b_nvfp4full-v2.ninfer |
| Uncensored | gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer | qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer |
The server fetches the standard one by itself the first time it starts without a model, and checks its size and SHA-256 before using it. To download either one by hand:
hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer --local-dir models
hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer --local-dir models
The uncensored image is the standard one with the
huihui-ai abliteration
applied: 1,255 of its 1,325 objects are byte-identical, and only the 70 matrices
the abliteration changed are re-encoded. The server never downloads it on its
own; start it with --artifact ./models/qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer,
or make dev UNCENSORED=1 from a checkout. It does not refuse, so put the
guardrails in the application or tool layer and do not expose it to untrusted
users.
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models \
ghcr.io/gpillon/ignis:0.1.2 --model-download-path /models
(docker: --gpus all in place of --device.) With no model in that
directory the server fetches one and verifies it. The Playground is then at
http://127.0.0.1:8000/ui/.
make doctor # toolchain, rust target, artifact, web deps
make dev # build web + kernel + GPU server, then run it
make dev VISION=1 # the same, with the vision tower loaded for image input
make dev UNCENSORED=1 # the same, on the uncensored (abliterated) weights
make mock # the same with no GPU, no kernel, no artifact
make alone lists every target; make config prints the exact server command a
run will use.
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Hello"}],"max_tokens":256}'
Full flag table, API surface, model handling, container and release detail: docs/user.
Every field of the OpenAI request bodies is honoured, accepted as inert (a
value asking for what the server already does, like n: 1 or store: false),
or refused with a 400 naming it in param — never dropped in silence. The
table is in the API reference at /v1/docs/.
Choices, not gaps — so nobody goes looking for them:
/v1/decide, read
off the loaded model's own readout./v1/completions. A raw prompt skips the chat template the model was
tuned on, and the turn boundaries reuse is keyed on; chat completions and
responses cover it./v1/batches. An offline queue is a storage service, and this server
holds no request past its connection: send the requests at once, and the
scheduler batches them.Two parts, in this order: the C++/CUDA kernel leaf, then the Rust workspace that links it.
nvcc, target SM120a); CMake + Ninja; Node.js for the Playground.kernel/build.ps1 on Windows, kernel/build.sh on Linux.
Both configure and build into kernel/build/, which crates/*/build.rs link.cargo build --release -p ignis-server --features cuda for
the real engine; without --features cuda (or without an artifact) the server
runs a deterministic CPU-only mock, which is how protocol and scheduler work
gets done with no card.| Path | What lives there |
|---|---|
crates/core/ | Scheduler, admission, paged KV accounting, request state, host tier. |
crates/artifact/ | The .ninfer reader: reader, binder, materializer. |
crates/runtime/ | Safe wrapper over the step ABI (CudaLeaf, decode graphs). |
crates/server/ | HTTP, OpenAI schemas, decisions, metrics, telemetry. |
crates/logging/ | Structured logging: tracing layers, hotpath lint, trace context. |
crates/bench/ | Trace-replay harness, gate and canary runner. |
crates/vendor/ | The vendoring tool: manifest, hashes, patch records. |
kernel/ | The C++/CUDA leaf: the program and the vendored ops (CMake + nvcc). |
web/ | The Playground: React + Vite, embedded into the server at build time. |
make wraps all of it. Tests: cargo test is workspace-wide, CPU-only and
never touches the GPU; GPU work is checked by a separate explicit profile that
fails rather than skips when the card is busy. The 5090 fits one run at a
time — check make gpu-status before starting anything on it.
docs/user/ | Running it: flags, API, models, container, releases, telemetry. |
docs/adr/ | Every architectural decision, and why it was taken. |
docs/findings/ | Durable, evidence-backed measurements. |
docs/agents/ | Conventions for agents working in this repo. |
CONTEXT.md | The glossary. One vocabulary, one meaning per term. |
The OpenAPI description of the HTTP surface is served at /v1/openapi.json,
with a browsable reference at /v1/docs/.
Ignis descends from NInfer, a lineage
of Windows-oriented local-inference forks, and does not hide it. The pinned
reference is the fork gpillon/ninfer,
itself downstream of cometkim/ninfer
(kernel-perf, hyperquant KV, NVFP4-full, the original DFlash2 port) →
natpate/ninfer-windows (the
Windows port) → Neroued/ninfer (the
original engine). The .ninfer model artifact is NInfer's format,
and the kernel leaf vendors NInfer's CUDA ops verbatim under a manifest that
pins the reference commit and every file's content hash — a port claim you can
diff. What is ours is the layer above them: the forward pass, the sequence
state, the step ABI, the scheduler, and by now the kernels that measurement
asked us to rewrite.
Ignis is not a port of NInfer. It is a different architecture that starts from proven kernel work.
With thanks to the NInfer project and its contributors:
Apache License 2.0.
Rust
65.1%
TypeScript
11.6%
Cuda
8.8%
Python
6.7%
C++
4.4%
C
1.2%