gpillon/ignis

Ignis is a deliberately specialized inference engine: one model family, one GPU class.

Rust

6

922 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

I made Ignis: an open-source engine that only runs Qwen3.8-27B on one RTX 5090, and that's the point. 512K context, Very Fast and JEV-live Support for Images and long text/logs retrival (r/LocalLLM)

Hi r/LocalLLM, Everything started from Ninfer... But it was not enough. So for the last few weeks I've been building from scratch (except for some kernels) **Ignis**, an Apache-2.0 inference engine in rust that deliberately gives up generality: one model family (Qwen3.8-27B NVFP4), one class of…

8

Oct 4, 2026

README

Ignis

Ignis

A single-GPU inference engine for Qwen3.8-27B, built for the load an agent makes.

OpenAI-compatible HTTP API  ·  Rust core  ·  C++/CUDA kernel leaf  ·  one Blackwell card

license Apache-2.0 target SM120a model Qwen3.8-27B NVFP4 API OpenAI-compatible
weights on Hugging Face uncensored weights on Hugging Face

367 tok/s
one lane, coding prompts
1,065 tok/s
eight lanes, aggregate
7.11x
KV capacity of BF16
~36 ms
a decision, zero tokens decoded
512K
context, YaRN x2

Measured on one RTX 5090, 64 GB DDR4-3200 MT/s, PCIe 3.0. How, and against what: How fast.


Ignis is a deliberately specialized engine: one model family, one class of card. It gives up generality and takes back speed. It loads an NVFP4 Qwen3.8-27B onto a single Blackwell card — an RTX 5090 or an RTX PRO 6000, both SM120a (why not DGX Spark?) — serves streaming and non-streaming completions over the OpenAI v1 API, and is shaped around the workload a developer actually produces — one main agent plus a handful of subagents hitting the same card at once. Eight resident decode lanes run as one batch-wide round; the paged KV cache is budgeted in bytes rather than in sequences; conversation state survives the request that built it; images are evidence, not an afterthought; and a decision can be answered without generating a single token. A Playground, a live Monitor and Prometheus metrics ship in the binary.

Rust owns everything above the step — scheduling, admission, KV accounting, serving. The forward pass and all GPU compute live in a C++/CUDA static library behind a flat, device-resident step-level C ABI. The engine is its own dogfood target, and partly there already: it serves some of the coding agents that build it, and the rest of that is what the remaining work is for.

TL;DR

Take a build from Releases — a Windows .zip, a Linux .tar.gz, or the linux/amd64 container image on ghcr.io/gpillon/ignis. Run it on a Blackwell card:

ignis-server --max-context 524288 --rope-scaling yarn:2 \
  --spec dflash2 --draft-tokens 7 --vision --metrics

No model yet? It offers to fetch one. Those flags turn on what the bare ignis-server leaves off: the 512K context (it starts at 40,960), DFlash2 speculative decoding, image input, and the Prometheus metrics the Playground's Monitor reads.

Then the OpenAI API is on http://127.0.0.1:8000/v1, the Playground and its Monitor on http://127.0.0.1:8000/ui/, and Prometheus on http://127.0.0.1:9464/metrics. The weights are on Hugging Face in two flavours, the standard one and an uncensored one: The weights. Everything else — the container, building from a checkout, every flag — is Quick start and docs/user.

How fast

Ignis on one RTX 5090 — Qwen3.8-27B NVFP4, hq-e8-2b KV, DFlash2 speculative decoding with 7 draft tokens, greedy, short prompts. The served context is 512K tokens: the checkpoint's native 262,144 stretched by YaRN x2, as the TL;DR flags and make set it. These runs used a 262,144-token maximum:

loadthroughputsource
one lane, eight coding prompts (write, edit, explain), geometric mean367 tok/supstream quick wins
one lane, predictable code / free prose491 / 167 tok/sretained slots on the host
eight lanes at once, aggregate1,065 tok/supstream quick wins

Other engines, same card, same model — single-stream decode as their authors published it:

engineweights and speculationsingle-stream decodesource
vLLM (main g41f179b57)NVFP4, FP8 KV, DFlash2 drafter with 6 tokens~304 tok/s code at 8K, 221 at 60KHF discussion #132
llama.cppNVFP4-MTP GGUF, q8_0 KV, MTP with 4 tokens122–142 tok/sDataCamp, 2026-08-25
llama.cppMTP70.5 tok/sHF discussion #132

Those rows are other people's runs: their prompts, their drafters, their context lengths. They are not a same-session benchmark, so read them as an order of magnitude, not a ratio. Prefill is where ignis is not ahead: vLLM's published ~12,500 tok/s at 60K is above ignis's measured 6,600 tok/s on a cold 46K prompt and 4,600 tok/s on a cold 105K one.

The one same-session comparison is against NInfer, the engine whose kernels ignis started from: same artifact, same card, one session, two process launches per side, before ignis had speculative decoding. ignis is 1.06x the reference at one lane and 3.07x at four, and 3% behind it on inter-token latency p95 (hq vs BF16 live/live).

What sets it apart

Every number below is a measurement on a 5090.

The decode round

  • One batch-wide round for all eight lanes, replayed from per-width CUDA graphs (widths 1..8), not one launch sequence per sequence.
  • The round is 1,166 graph nodes and 15.81 ms of device time that is weight streaming at the card's bandwidth — the backbone GEMM at 100% of roofline, the output head at 86%. Graph submission costs 2.2% of it, and the whole remaining headroom is the 4.8% the device spends idle.
  • Prefill and decode interleave at chunk level, so a long prompt never stalls the lanes already generating.

Speculative decoding (DFlash2)

  • The served artifact carries a grafted DFlash2 drafter; --spec dflash2 --draft-tokens 7 runs draft-and-verify rounds against it.
  • The drafter's vendored top-k gave one warp to each of seven columns over a 248k vocabulary and spent 3,145 µs — 16.4% of decode kernel time — to move 2.0 µs of memory. Our row-split replacement does it in 44 µs at a bit-identical answer: 17.6% of a decode round, +21% decode throughput at one lane. A vendored kernel measured as the bottleneck may be replaced — that is the one exemption to verbatim vendoring.

KV that costs bytes, not sequences

  • hq-e8-2b is the serving KV format: a sequence-token costs 9,216 bytes against BF16's 65,536 — 7.11x the capacity, so eight lanes at a 40,960 context need ~3.02 GB instead of 20 GiB.
  • BF16 is retained, not deprecated: it is the format every correctness oracle runs against.
  • The GQA workspace zeroing is dead work under hq — every byte the next launch reads it writes itself — and removing it is worth 2.12% of GQA layer device time on a 70K prefill.

State that outlives its request

  • Cross-request reuse: a prompt checkpoint captured at the generation opener, a retained prefix shared between siblings, and a lazy spill into a pinned KV-RAM host tier with unified eviction.
  • A device-to-device prefix clone is 0.25 ms for 148 MiB — 96x cheaper than a PCIe round trip and ~8,000x cheaper than re-prefilling the head; a full-context snapshot/restore round trip is ~90 ms, about 100x cheaper than the re-prefill it replaces.

A VRAM plan, decided at load

  • Every byte the process will hold is reserved and laid out before the first request, and the plan is printed: weights, workspaces, lane state, decode and drafter graphs, retained slots, media embeddings, and whatever is left becomes the KV pool. The Monitor shows that plan beside what is in use — see it.
  • The load that used to grow the process 2,734 MiB in 14 minutes now grows it 6 MiB in 24.5, and the printed plan matches what the OS reports to within 16 MiB. No paging, no surprise at minute forty.

Lanes that know who is asking

  • A request states its own lane tag — interactive or agent — through the class extension field or an @<lane> suffix on the model name. The tag drives protection, backfill priority and eviction order, so a burst of subagents fills the lanes a foreground conversation is not using instead of evicting it.
  • Admission is a full state machine: protection, backfill class, temporal credit, frontier distance.

Jev-like decisions, answered without generating (/v1/decide)

  • A classification is not a completion. POST /v1/decide prefills the evidence once and reads the answer out of the logits at a single position.
  • Six primitives: noul (yes/no), choice, score, and — one constrained digit per step — number, point, box.
  • On 144 authored decisions the declared options hold a median 99.8% of the model's distribution, and the readout costs ~36.5 ms and zero decoded tokens at 0.934 balanced accuracy on the layout first measured; moving the evidence into the system block takes it to 0.963.
  • point reads a click target off a 4096px screenshot to within 84 px worst case, inside the button on every scene tested.
  • The same handler is served as /v1/systemone, so a Jev client reaches this engine by changing the URL and nothing else.

Images as evidence

  • --vision loads the tower and reserves its workspace; images arrive inline as data URIs or as URLs the server fetches, prepares and caches.
  • The encoder output is kept past its request, keyed by content digest and grid: four questions over one 4096x4096 screenshot go from 27.92 s to 13.24 s on one encode instead of four, and 9.49 s when the picture was already seen.
  • The tower's quadratic curve is the checkpoint's own scheme, not a bug: 88.2% of a 16,384-column encode is the attention kernel, holding 164.7 TFLOP/s with 0.1% device idle, so 3.67 s is the worst case per image — and the only lever is sending a smaller one.

A context you can rescale

  • The checkpoint is trained to 262,144 positions; --rope-scaling yarn:F rescales that envelope, yarn:4 putting the ceiling at 1,048,576 (factor up to 64). What a long-context probe must ask is a question about relative position — a literal needle at 320K is recalled with the flag and without it.

Batteries in the binary

  • The Playground at /ui/ — chat with tools and subagents, parallel sessions, image input, a Decide tab and a live Monitor. Built into the server, not a separate service.
  • Prometheus on its own listener (--metrics), with the VRAM plan, retained state, KV pool occupancy, decision counters and request lifecycle.
  • It fetches its own model. Started with no artifact, a GPU build asks, then downloads and verifies it.
  • --api-key and --expose — a keyed public URL through a quick Cloudflare tunnel, key always required.
  • Linux and Windows, a linux/amd64 container image on ghcr.io, and release archives cut from the same build.

The Playground

Agents running in parallel, each on its own lane, with the engine's own timings beside every reply:

The Playground running four subagents

A decision over prose — the winner, the whole distribution, and a score read as levels — answered out of one prefill with nothing generated:

A text decision and its distribution

The same endpoint over an image: a point, a box, and a yes/no, each a digit at a time, with the model's self-reported uncertainty on every axis:

A decision over an image, with point and box

The Monitor scrapes the engine's own metrics: throughput, lane occupancy, latency quantiles:

The Monitor, live

...and what the load actually reserved, against the constants that bound it:

The VRAM plan and retained state

Quick start

The weights

Two images of the same model, both on Hugging Face, both 18.07 GiB, both carrying the DFlash2 drafter:

Hugging Facefile
Standardgpillon/Qwen3.8-27B-nvfp4full-dflash2-NInferqwen3_8_27b_nvfp4full-v2.ninfer
Uncensoredgpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInferqwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer

The server fetches the standard one by itself the first time it starts without a model, and checks its size and SHA-256 before using it. To download either one by hand:

hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer --local-dir models
hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer --local-dir models

The uncensored image is the standard one with the huihui-ai abliteration applied: 1,255 of its 1,325 objects are byte-identical, and only the 70 matrices the abliteration changed are re-encoded. The server never downloads it on its own; start it with --artifact ./models/qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer, or make dev UNCENSORED=1 from a checkout. It does not refuse, so put the guardrails in the application or tool layer and do not expose it to untrusted users.

Container

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models \
  ghcr.io/gpillon/ignis:0.1.2 --model-download-path /models

(docker: --gpus all in place of --device.) With no model in that directory the server fetches one and verifies it. The Playground is then at http://127.0.0.1:8000/ui/.

From a checkout

make doctor          # toolchain, rust target, artifact, web deps
make dev             # build web + kernel + GPU server, then run it
make dev VISION=1    # the same, with the vision tower loaded for image input
make dev UNCENSORED=1  # the same, on the uncensored (abliterated) weights
make mock            # the same with no GPU, no kernel, no artifact

make alone lists every target; make config prints the exact server command a run will use.

A request

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Hello"}],"max_tokens":256}'

Full flag table, API surface, model handling, container and release detail: docs/user.

Every field of the OpenAI request bodies is honoured, accepted as inert (a value asking for what the server already does, like n: 1 or store: false), or refused with a 400 naming it in param — never dropped in silence. The table is in the API reference at /v1/docs/.

What it does not serve

Choices, not gaps — so nobody goes looking for them:

  • Embeddings, or a generic reranker. Either is a second model on the card; the VRAM plan gives the card to one. A classification is /v1/decide, read off the loaded model's own readout.
  • /v1/completions. A raw prompt skips the chat template the model was tuned on, and the turn boundaries reuse is keyed on; chat completions and responses cover it.
  • /v1/batches. An offline queue is a storage service, and this server holds no request past its connection: send the requests at once, and the scheduler batches them.
  • Audio, in or out. The model has no audio tower and no audio head.
  • LoRA adapters. They are trained against BF16 weights, and the checkpoint is NVFP4: merge the adapter and re-export the checkpoint instead.
  • More than one loaded model. One model family, one card: the VRAM plan is decided at load, for one.

Building it

Two parts, in this order: the C++/CUDA kernel leaf, then the Rust workspace that links it.

  • Toolchain — Rust; MSVC C++ build tools (Windows) or GCC (Linux); the CUDA Toolkit (nvcc, target SM120a); CMake + Ninja; Node.js for the Playground.
  • Kernel leaf — kernel/build.ps1 on Windows, kernel/build.sh on Linux. Both configure and build into kernel/build/, which crates/*/build.rs link.
  • Workspace — cargo build --release -p ignis-server --features cuda for the real engine; without --features cuda (or without an artifact) the server runs a deterministic CPU-only mock, which is how protocol and scheduler work gets done with no card.

The repo

PathWhat lives there
crates/core/Scheduler, admission, paged KV accounting, request state, host tier.
crates/artifact/The .ninfer reader: reader, binder, materializer.
crates/runtime/Safe wrapper over the step ABI (CudaLeaf, decode graphs).
crates/server/HTTP, OpenAI schemas, decisions, metrics, telemetry.
crates/logging/Structured logging: tracing layers, hotpath lint, trace context.
crates/bench/Trace-replay harness, gate and canary runner.
crates/vendor/The vendoring tool: manifest, hashes, patch records.
kernel/The C++/CUDA leaf: the program and the vendored ops (CMake + nvcc).
web/The Playground: React + Vite, embedded into the server at build time.

make wraps all of it. Tests: cargo test is workspace-wide, CPU-only and never touches the GPU; GPU work is checked by a separate explicit profile that fails rather than skips when the card is busy. The 5090 fits one run at a time — check make gpu-status before starting anything on it.

Documentation

docs/user/Running it: flags, API, models, container, releases, telemetry.
docs/adr/Every architectural decision, and why it was taken.
docs/findings/Durable, evidence-backed measurements.
docs/agents/Conventions for agents working in this repo.
CONTEXT.mdThe glossary. One vocabulary, one meaning per term.

The OpenAPI description of the HTTP surface is served at /v1/openapi.json, with a browsable reference at /v1/docs/.

Lineage

Ignis descends from NInfer, a lineage of Windows-oriented local-inference forks, and does not hide it. The pinned reference is the fork gpillon/ninfer, itself downstream of cometkim/ninfer (kernel-perf, hyperquant KV, NVFP4-full, the original DFlash2 port) → natpate/ninfer-windows (the Windows port) → Neroued/ninfer (the original engine). The .ninfer model artifact is NInfer's format, and the kernel leaf vendors NInfer's CUDA ops verbatim under a manifest that pins the reference commit and every file's content hash — a port claim you can diff. What is ours is the layer above them: the forward pass, the sequence state, the step ABI, the scheduler, and by now the kernels that measurement asked us to rewrite.

Ignis is not a port of NInfer. It is a different architecture that starts from proven kernel work.

Credits

With thanks to the NInfer project and its contributors:

  • cometkim — kernel-perf integration (PDL decode chain, split-K prefill, per-request error boundary), the DFlash2 base port, the hyperquant KV cache, the 1M-context envelope, the NVFP4-full target.
  • Mirko Covizzi — RTX 5090 Laptop compatibility, MTP adaptive verification-width tuning.
  • mr-september — warmup/readiness decoupling, frontend streaming fixes.
  • dylan (dylanbrodiefafard) — the RAM KV cache concept, LRU lane eviction, the decode CPU-spin fix.
  • Neroued — the original NInfer engine.

License

Apache License 2.0.

gpillon/ignis

Ignis is a deliberately specialized inference engine: one model family, one GPU class.

Rust

6

922 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

I made Ignis: an open-source engine that only runs Qwen3.8-27B on one RTX 5090, and that's the point. 512K context, Very Fast and JEV-live Support for Images and long text/logs retrival (r/LocalLLM)

Hi r/LocalLLM, Everything started from Ninfer... But it was not enough. So for the last few weeks I've been building from scratch (except for some kernels) **Ignis**, an Apache-2.0 inference engine in rust that deliberately gives up generality: one model family (Qwen3.8-27B NVFP4), one class of…

8

Oct 4, 2026

README

Ignis

Ignis

A single-GPU inference engine for Qwen3.8-27B, built for the load an agent makes.

OpenAI-compatible HTTP API  ·  Rust core  ·  C++/CUDA kernel leaf  ·  one Blackwell card

license Apache-2.0 target SM120a model Qwen3.8-27B NVFP4 API OpenAI-compatible
weights on Hugging Face uncensored weights on Hugging Face

367 tok/s
one lane, coding prompts
1,065 tok/s
eight lanes, aggregate
7.11x
KV capacity of BF16
~36 ms
a decision, zero tokens decoded
512K
context, YaRN x2

Measured on one RTX 5090, 64 GB DDR4-3200 MT/s, PCIe 3.0. How, and against what: How fast.


Ignis is a deliberately specialized engine: one model family, one class of card. It gives up generality and takes back speed. It loads an NVFP4 Qwen3.8-27B onto a single Blackwell card — an RTX 5090 or an RTX PRO 6000, both SM120a (why not DGX Spark?) — serves streaming and non-streaming completions over the OpenAI v1 API, and is shaped around the workload a developer actually produces — one main agent plus a handful of subagents hitting the same card at once. Eight resident decode lanes run as one batch-wide round; the paged KV cache is budgeted in bytes rather than in sequences; conversation state survives the request that built it; images are evidence, not an afterthought; and a decision can be answered without generating a single token. A Playground, a live Monitor and Prometheus metrics ship in the binary.

Rust owns everything above the step — scheduling, admission, KV accounting, serving. The forward pass and all GPU compute live in a C++/CUDA static library behind a flat, device-resident step-level C ABI. The engine is its own dogfood target, and partly there already: it serves some of the coding agents that build it, and the rest of that is what the remaining work is for.

TL;DR

Take a build from Releases — a Windows .zip, a Linux .tar.gz, or the linux/amd64 container image on ghcr.io/gpillon/ignis. Run it on a Blackwell card:

ignis-server --max-context 524288 --rope-scaling yarn:2 \
  --spec dflash2 --draft-tokens 7 --vision --metrics

No model yet? It offers to fetch one. Those flags turn on what the bare ignis-server leaves off: the 512K context (it starts at 40,960), DFlash2 speculative decoding, image input, and the Prometheus metrics the Playground's Monitor reads.

Then the OpenAI API is on http://127.0.0.1:8000/v1, the Playground and its Monitor on http://127.0.0.1:8000/ui/, and Prometheus on http://127.0.0.1:9464/metrics. The weights are on Hugging Face in two flavours, the standard one and an uncensored one: The weights. Everything else — the container, building from a checkout, every flag — is Quick start and docs/user.

How fast

Ignis on one RTX 5090 — Qwen3.8-27B NVFP4, hq-e8-2b KV, DFlash2 speculative decoding with 7 draft tokens, greedy, short prompts. The served context is 512K tokens: the checkpoint's native 262,144 stretched by YaRN x2, as the TL;DR flags and make set it. These runs used a 262,144-token maximum:

loadthroughputsource
one lane, eight coding prompts (write, edit, explain), geometric mean367 tok/supstream quick wins
one lane, predictable code / free prose491 / 167 tok/sretained slots on the host
eight lanes at once, aggregate1,065 tok/supstream quick wins

Other engines, same card, same model — single-stream decode as their authors published it:

engineweights and speculationsingle-stream decodesource
vLLM (main g41f179b57)NVFP4, FP8 KV, DFlash2 drafter with 6 tokens~304 tok/s code at 8K, 221 at 60KHF discussion #132
llama.cppNVFP4-MTP GGUF, q8_0 KV, MTP with 4 tokens122–142 tok/sDataCamp, 2026-08-25
llama.cppMTP70.5 tok/sHF discussion #132

Those rows are other people's runs: their prompts, their drafters, their context lengths. They are not a same-session benchmark, so read them as an order of magnitude, not a ratio. Prefill is where ignis is not ahead: vLLM's published ~12,500 tok/s at 60K is above ignis's measured 6,600 tok/s on a cold 46K prompt and 4,600 tok/s on a cold 105K one.

The one same-session comparison is against NInfer, the engine whose kernels ignis started from: same artifact, same card, one session, two process launches per side, before ignis had speculative decoding. ignis is 1.06x the reference at one lane and 3.07x at four, and 3% behind it on inter-token latency p95 (hq vs BF16 live/live).

What sets it apart

Every number below is a measurement on a 5090.

The decode round

  • One batch-wide round for all eight lanes, replayed from per-width CUDA graphs (widths 1..8), not one launch sequence per sequence.
  • The round is 1,166 graph nodes and 15.81 ms of device time that is weight streaming at the card's bandwidth — the backbone GEMM at 100% of roofline, the output head at 86%. Graph submission costs 2.2% of it, and the whole remaining headroom is the 4.8% the device spends idle.
  • Prefill and decode interleave at chunk level, so a long prompt never stalls the lanes already generating.

Speculative decoding (DFlash2)

  • The served artifact carries a grafted DFlash2 drafter; --spec dflash2 --draft-tokens 7 runs draft-and-verify rounds against it.
  • The drafter's vendored top-k gave one warp to each of seven columns over a 248k vocabulary and spent 3,145 µs — 16.4% of decode kernel time — to move 2.0 µs of memory. Our row-split replacement does it in 44 µs at a bit-identical answer: 17.6% of a decode round, +21% decode throughput at one lane. A vendored kernel measured as the bottleneck may be replaced — that is the one exemption to verbatim vendoring.

KV that costs bytes, not sequences

  • hq-e8-2b is the serving KV format: a sequence-token costs 9,216 bytes against BF16's 65,536 — 7.11x the capacity, so eight lanes at a 40,960 context need ~3.02 GB instead of 20 GiB.
  • BF16 is retained, not deprecated: it is the format every correctness oracle runs against.
  • The GQA workspace zeroing is dead work under hq — every byte the next launch reads it writes itself — and removing it is worth 2.12% of GQA layer device time on a 70K prefill.

State that outlives its request

  • Cross-request reuse: a prompt checkpoint captured at the generation opener, a retained prefix shared between siblings, and a lazy spill into a pinned KV-RAM host tier with unified eviction.
  • A device-to-device prefix clone is 0.25 ms for 148 MiB — 96x cheaper than a PCIe round trip and ~8,000x cheaper than re-prefilling the head; a full-context snapshot/restore round trip is ~90 ms, about 100x cheaper than the re-prefill it replaces.

A VRAM plan, decided at load

  • Every byte the process will hold is reserved and laid out before the first request, and the plan is printed: weights, workspaces, lane state, decode and drafter graphs, retained slots, media embeddings, and whatever is left becomes the KV pool. The Monitor shows that plan beside what is in use — see it.
  • The load that used to grow the process 2,734 MiB in 14 minutes now grows it 6 MiB in 24.5, and the printed plan matches what the OS reports to within 16 MiB. No paging, no surprise at minute forty.

Lanes that know who is asking

  • A request states its own lane tag — interactive or agent — through the class extension field or an @<lane> suffix on the model name. The tag drives protection, backfill priority and eviction order, so a burst of subagents fills the lanes a foreground conversation is not using instead of evicting it.
  • Admission is a full state machine: protection, backfill class, temporal credit, frontier distance.

Jev-like decisions, answered without generating (/v1/decide)

  • A classification is not a completion. POST /v1/decide prefills the evidence once and reads the answer out of the logits at a single position.
  • Six primitives: noul (yes/no), choice, score, and — one constrained digit per step — number, point, box.
  • On 144 authored decisions the declared options hold a median 99.8% of the model's distribution, and the readout costs ~36.5 ms and zero decoded tokens at 0.934 balanced accuracy on the layout first measured; moving the evidence into the system block takes it to 0.963.
  • point reads a click target off a 4096px screenshot to within 84 px worst case, inside the button on every scene tested.
  • The same handler is served as /v1/systemone, so a Jev client reaches this engine by changing the URL and nothing else.

Images as evidence

  • --vision loads the tower and reserves its workspace; images arrive inline as data URIs or as URLs the server fetches, prepares and caches.
  • The encoder output is kept past its request, keyed by content digest and grid: four questions over one 4096x4096 screenshot go from 27.92 s to 13.24 s on one encode instead of four, and 9.49 s when the picture was already seen.
  • The tower's quadratic curve is the checkpoint's own scheme, not a bug: 88.2% of a 16,384-column encode is the attention kernel, holding 164.7 TFLOP/s with 0.1% device idle, so 3.67 s is the worst case per image — and the only lever is sending a smaller one.

A context you can rescale

  • The checkpoint is trained to 262,144 positions; --rope-scaling yarn:F rescales that envelope, yarn:4 putting the ceiling at 1,048,576 (factor up to 64). What a long-context probe must ask is a question about relative position — a literal needle at 320K is recalled with the flag and without it.

Batteries in the binary

  • The Playground at /ui/ — chat with tools and subagents, parallel sessions, image input, a Decide tab and a live Monitor. Built into the server, not a separate service.
  • Prometheus on its own listener (--metrics), with the VRAM plan, retained state, KV pool occupancy, decision counters and request lifecycle.
  • It fetches its own model. Started with no artifact, a GPU build asks, then downloads and verifies it.
  • --api-key and --expose — a keyed public URL through a quick Cloudflare tunnel, key always required.
  • Linux and Windows, a linux/amd64 container image on ghcr.io, and release archives cut from the same build.

The Playground

Agents running in parallel, each on its own lane, with the engine's own timings beside every reply:

The Playground running four subagents

A decision over prose — the winner, the whole distribution, and a score read as levels — answered out of one prefill with nothing generated:

A text decision and its distribution

The same endpoint over an image: a point, a box, and a yes/no, each a digit at a time, with the model's self-reported uncertainty on every axis:

A decision over an image, with point and box

The Monitor scrapes the engine's own metrics: throughput, lane occupancy, latency quantiles:

The Monitor, live

...and what the load actually reserved, against the constants that bound it:

The VRAM plan and retained state

Quick start

The weights

Two images of the same model, both on Hugging Face, both 18.07 GiB, both carrying the DFlash2 drafter:

Hugging Facefile
Standardgpillon/Qwen3.8-27B-nvfp4full-dflash2-NInferqwen3_8_27b_nvfp4full-v2.ninfer
Uncensoredgpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInferqwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer

The server fetches the standard one by itself the first time it starts without a model, and checks its size and SHA-256 before using it. To download either one by hand:

hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer --local-dir models
hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer --local-dir models

The uncensored image is the standard one with the huihui-ai abliteration applied: 1,255 of its 1,325 objects are byte-identical, and only the 70 matrices the abliteration changed are re-encoded. The server never downloads it on its own; start it with --artifact ./models/qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer, or make dev UNCENSORED=1 from a checkout. It does not refuse, so put the guardrails in the application or tool layer and do not expose it to untrusted users.

Container

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models \
  ghcr.io/gpillon/ignis:0.1.2 --model-download-path /models

(docker: --gpus all in place of --device.) With no model in that directory the server fetches one and verifies it. The Playground is then at http://127.0.0.1:8000/ui/.

From a checkout

make doctor          # toolchain, rust target, artifact, web deps
make dev             # build web + kernel + GPU server, then run it
make dev VISION=1    # the same, with the vision tower loaded for image input
make dev UNCENSORED=1  # the same, on the uncensored (abliterated) weights
make mock            # the same with no GPU, no kernel, no artifact

make alone lists every target; make config prints the exact server command a run will use.

A request

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Hello"}],"max_tokens":256}'

Full flag table, API surface, model handling, container and release detail: docs/user.

Every field of the OpenAI request bodies is honoured, accepted as inert (a value asking for what the server already does, like n: 1 or store: false), or refused with a 400 naming it in param — never dropped in silence. The table is in the API reference at /v1/docs/.

What it does not serve

Choices, not gaps — so nobody goes looking for them:

  • Embeddings, or a generic reranker. Either is a second model on the card; the VRAM plan gives the card to one. A classification is /v1/decide, read off the loaded model's own readout.
  • /v1/completions. A raw prompt skips the chat template the model was tuned on, and the turn boundaries reuse is keyed on; chat completions and responses cover it.
  • /v1/batches. An offline queue is a storage service, and this server holds no request past its connection: send the requests at once, and the scheduler batches them.
  • Audio, in or out. The model has no audio tower and no audio head.
  • LoRA adapters. They are trained against BF16 weights, and the checkpoint is NVFP4: merge the adapter and re-export the checkpoint instead.
  • More than one loaded model. One model family, one card: the VRAM plan is decided at load, for one.

Building it

Two parts, in this order: the C++/CUDA kernel leaf, then the Rust workspace that links it.

  • Toolchain — Rust; MSVC C++ build tools (Windows) or GCC (Linux); the CUDA Toolkit (nvcc, target SM120a); CMake + Ninja; Node.js for the Playground.
  • Kernel leaf — kernel/build.ps1 on Windows, kernel/build.sh on Linux. Both configure and build into kernel/build/, which crates/*/build.rs link.
  • Workspace — cargo build --release -p ignis-server --features cuda for the real engine; without --features cuda (or without an artifact) the server runs a deterministic CPU-only mock, which is how protocol and scheduler work gets done with no card.

The repo

PathWhat lives there
crates/core/Scheduler, admission, paged KV accounting, request state, host tier.
crates/artifact/The .ninfer reader: reader, binder, materializer.
crates/runtime/Safe wrapper over the step ABI (CudaLeaf, decode graphs).
crates/server/HTTP, OpenAI schemas, decisions, metrics, telemetry.
crates/logging/Structured logging: tracing layers, hotpath lint, trace context.
crates/bench/Trace-replay harness, gate and canary runner.
crates/vendor/The vendoring tool: manifest, hashes, patch records.
kernel/The C++/CUDA leaf: the program and the vendored ops (CMake + nvcc).
web/The Playground: React + Vite, embedded into the server at build time.

make wraps all of it. Tests: cargo test is workspace-wide, CPU-only and never touches the GPU; GPU work is checked by a separate explicit profile that fails rather than skips when the card is busy. The 5090 fits one run at a time — check make gpu-status before starting anything on it.

Documentation

docs/user/Running it: flags, API, models, container, releases, telemetry.
docs/adr/Every architectural decision, and why it was taken.
docs/findings/Durable, evidence-backed measurements.
docs/agents/Conventions for agents working in this repo.
CONTEXT.mdThe glossary. One vocabulary, one meaning per term.

The OpenAPI description of the HTTP surface is served at /v1/openapi.json, with a browsable reference at /v1/docs/.

Lineage

Ignis descends from NInfer, a lineage of Windows-oriented local-inference forks, and does not hide it. The pinned reference is the fork gpillon/ninfer, itself downstream of cometkim/ninfer (kernel-perf, hyperquant KV, NVFP4-full, the original DFlash2 port) → natpate/ninfer-windows (the Windows port) → Neroued/ninfer (the original engine). The .ninfer model artifact is NInfer's format, and the kernel leaf vendors NInfer's CUDA ops verbatim under a manifest that pins the reference commit and every file's content hash — a port claim you can diff. What is ours is the layer above them: the forward pass, the sequence state, the step ABI, the scheduler, and by now the kernels that measurement asked us to rewrite.

Ignis is not a port of NInfer. It is a different architecture that starts from proven kernel work.

Credits

With thanks to the NInfer project and its contributors:

  • cometkim — kernel-perf integration (PDL decode chain, split-K prefill, per-request error boundary), the DFlash2 base port, the hyperquant KV cache, the 1M-context envelope, the NVFP4-full target.
  • Mirko Covizzi — RTX 5090 Laptop compatibility, MTP adaptive verification-width tuning.
  • mr-september — warmup/readiness decoupling, frontend streaming fixes.
  • dylan (dylanbrodiefafard) — the RAM KV cache concept, LRU lane eviction, the decode CPU-spin fix.
  • Neroued — the original NInfer engine.

License

Apache License 2.0.

Languages

Rust

65.1%

TypeScript

11.6%

Cuda

8.8%

Python

6.7%

C++

4.4%

C

1.2%