peonist-ai/halogen-server

The fastest way to run Qwen3.8 27B on Strix Halo (gfx1151)

Python

121

11 commits

updated Sep 14, 2026

See the code

README

halogen — the peon's inference engine

halogen

The fastest way to run Qwen3.8-27B on AMD Strix Halo — at higher precision than anything that comes close.

Every kernel is written for this one GPU and this one model family. No general-purpose runtime, no portability layer, no fallback path — which is why it can do things a general engine cannot, and why it runs on exactly one piece of silicon.

On a 32K prompt with a 256-token answer, against the fastest numbers anyone else has published for this model on this hardware:

prefilldecodetotal
halogen (6.32 bpw)57.9 s8.1 s66.0 s
KyaniteLabs (Q4_K_XL)84.0 s8.5 s92.5 s
q38rocm (4.26 bpw)133.7 s7.8 s141.5 s

2.1× faster than q38rocm and 1.4× faster than KyaniteLabs end-to-end, while carrying ~1.5× their weight precision. Prefill is where that is won, and on any prompt with real context prefill is most of the wall clock.

Output is also byte-identical to serial greedy decode — speculation here is a pure speed optimization, verified on every release, not a quality trade.

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -v /path/to/models:/models:ro -v /path/to/tokenizer:/tokenizer:ro \
  ghcr.io/peonist-ai/halogen:0.1.4

An OpenAI-compatible endpoint comes up on :8731.

On Docker rather than Podman, replace --group-add keep-groups with --group-add video --group-add render. keep-groups is a Podman keyword that Docker does not understand: Docker resolves --group-add names against the container's /etc/group and fails with unable to find group keep-groups.


Get the weights

The image contains no model weights — it is 3.5 GB of engine, and the checkpoint is 35.9 GB. Download it once and mount it:

pip install -U "huggingface_hub[cli]"
hf download peonist-ai/halogen-qwen3.8-27b \
  --local-dir ~/halogen-models

That repository carries both the .hgn checkpoint and a flat tokenizer directory, so there is nothing to assemble by hand:

~/halogen-models/
  qwen3.8-27b-p1w4d-d2.hgn      35.9 GB   the checkpoint
  tokenizer/                              tokenizer.json, chat template, ...

Then point the container at both:

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -v ~/halogen-models:/models:ro \
  -v ~/halogen-models/tokenizer:/tokenizer:ro \
  ghcr.io/peonist-ai/halogen:0.1.4

Or let it fetch them for you

If you would rather not download separately, set HALOGEN_DOWNLOAD and the container fetches the weights on first start:

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -e HALOGEN_DOWNLOAD=peonist-ai/halogen-qwen3.8-27b \
  -e HALOGEN_TOKENIZER=/models/tokenizer \
  -v ~/halogen-models:/models \
  ghcr.io/peonist-ai/halogen:0.1.4

Two differences from the manual route. The models volume is mounted read-write — it has to be, to download into. And there is only one mount: the download brings the tokenizer with it, so HALOGEN_TOKENIZER points inside /models rather than at a second volume. Mounting ~/halogen-models/tokenizer here would fail on a first run, because the container runtime would create it as an empty directory before the download had a chance to populate it. It fires only when the checkpoint is actually missing, so restarts do not re-download, and an interrupted transfer resumes rather than starting over.

With HALOGEN_DOWNLOAD unset, the container makes no outbound network connections at all — no telemetry, no license check, no model fetch. If the checkpoint is not on disk where HALOGEN_CHECKPOINT points, it says so and exits rather than reaching for the network. That default is deliberate: a 35.9 GB transfer should not begin because someone ran podman run to see what would happen.

Model weights are licensed separately from the engine by their original authors; see the model repository for those terms.


Performance

Measured on a Ryzen AI Max+ 395 (Radeon 8060S, 128 GB LPDDR5X), ROCm 7.14.0, checkpoint p1w4d-d2 (~6.3 bits/weight effective at decode), 262,144 context.

Prefill

All three measured over the HTTP endpoint with the bundled sweep, so they are one instrument rather than a mix.

testt/s
pp512620
pp2048710
pp32768566

Decode

Over the HTTP endpoint, ten prompt shapes, greedy, DFlash2 drafter:

mean t/srange
DFlash2 (default)31.7120.8 – 44.2
MTP26.8819.8 – 34.2
serial (no speculation)10.58—

Aggregate throughput at 8 concurrent requests: 48.6 t/s (4.87×).

Read the range, not just the mean

halogen's decode rate is a distribution, not a number. Speculative decoding accepts more drafted tokens when the text is predictable, so the same build on the same hardware does:

  • prose / chat: 20.8 – 23.5 t/s
  • procedures: 25.6 – 38.7 t/s
  • code / proofs: 32.7 – 44.2 t/s

A single headline figure hides a 2× spread. Any decode number quoted from this project — by us or anyone else — should name the prompt set that produced it, or it is not reproducible. bench prints the mean; sweep prints mean, standard deviation, and min–max, deliberately.

An engine without speculative decoding has a content-independent decode rate and can honestly quote one number. We can't.


How it compares

Published numbers from other projects running the same model on the same silicon. These are their figures on their configurations, not a head-to-head we ran — quantization, KV-cache settings and context differ, so read this as orientation, not as a controlled benchmark.

halogenq38rocmKyaniteLabs
backendcustom HIPROCm/RADVllama.cpp
weights~6.3 bpw4.26 bpwUD-Q4_K_XL
prefill @32K566 t/s245 t/s~390 t/s
decode, speculated30.98 mean (20.8–44.2)30.56 – 36.04prose 11–24, code 29–40
decode, unassisted (diagnostic)10.58 t/s14.02 t/s—

Where we win: prefill, by 1.4–2.3×, and end-to-end on any prompt with real context. That is what the engine was built for.

Where we lose: unassisted decode — and that row is a diagnostic, not a product configuration. Nobody ships serial decode; every project in this table runs speculation by default. The gap is also not kernel quality: decode is bandwidth-bound, q38rocm streams ~17 GB/token against our 23.5, and fewer bits is simply faster. We spend those bits deliberately (see QUANT.md) — the only 4-bit tensors in our trunk are ones somebody else calibrated, and the aggressive technique is fenced to prefill where it never touches token generation.

Our batch-1 decode is at the hardware wall. 10.58 t/s × 23.51 GB/token = 249 GB/s against a measured ceiling of 240 GB/s. There is no kernel win left there for anyone; the levers are fewer bits, better draft acceptance, and batching.

On the 148–163 t/s figure circulating for llama.cpp on this hardware: that is an ngram-repetition artifact on back-to-back identical runs, and KyaniteLabs — whose benchmark it is — says so plainly and warns against quoting it for chat. Their honest conversational numbers are in the table. We think that is the right way to publish, and we have tried to match it.


Benchmark it yourself

The image ships both benchmarks. No fixtures, no extra downloads, no cooperation from us required.

# ten real prompt shapes over the HTTP endpoint — the number of record
podman run --rm --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -v /path/to/models:/models:ro -v /path/to/tokenizer:/tokenizer:ro \
  ghcr.io/peonist-ai/halogen:0.1.4 bench dflash2 256 low 3

# llama-bench-shaped pp/tg sweep, for putting a number beside another engine
podman run --rm ... ghcr.io/peonist-ai/halogen:0.1.4 \
  sweep -p 512,2048,8192 -n 128,256 -d dflash2,mtp -r 3

sweep --json emits machine-readable output.

Publishing the results is expressly permitted — no approval, no notice, no prior review. We only ask that figures name the version and the prompt set, for the reason above. That request is not a licensing condition.


What it does

Byte-identical speculative decoding. Draft-then-verify commits only tokens the full model would have produced, so output is bit-for-bit identical to serial greedy decode. This is gated on every release across all three drafters — not asserted, measured. Speculation here is a pure speed optimization with no quality cost, and you can turn it off per request to check.

Native 262,144-token context, with decode that barely degrades at depth — Gated DeltaNet carries O(1) state, so 48 of 64 layers have no KV cache at all.

Prompt cache — a follow-up turn on a long conversation resumes instead of re-prefilling, worth roughly 20× on time-to-first-token at 32K. Warm answers are byte-identical to cold ones by construction.

Batched decode — 8 concurrent sequences, 4.87× aggregate, each byte-identical to running alone. Off by default, and it trades away speculation when enabled; see Configuration before turning it on.

OpenAI-compatible API — /v1/chat/completions, /v1/completions, streaming, tool calling, sampling with seeds, reasoning-effort control.

Three selectable drafters — dflash2 (default), mtp, serial. Choose per request; output is identical, only speed changes.


Configuration

Everything is set by environment variable — there is no config file. The complete list of levers, with defaults and whether each one can change output, is in docs/FLAGS.md. These are the ones most people touch:

variabledefaultwhat it does
HALOGEN_CHECKPOINT/models/qwen3.8-27b-p1w4d-d2.hgnwhich checkpoint to load
HALOGEN_TOKENIZER/tokenizerflat tokenizer directory
HALOGEN_API_PORT8731the published port
HALOGEN_DRAFTER2 (DFlash2)default drafter: 0 serial, 1 MTP, 2 DFlash2
HALOGEN_CACHE_MBautoprompt cache budget in MB — a byte budget, not a slot count; 0 disables
HALOGEN_MAX_TOKENS_CAP65536largest max_tokens a request may ask for — over it is a 400, never a silent truncation
HALOGEN_QUEUE_TIMEOUT7200seconds a queued request will wait — coupled to the cap, see below
HALOGEN_KV_SLOTS1concurrent resident sequences — see below
HALOGEN_SLOT_CTX262144context each slot holds — see below

Per-request settings — drafter, temperature, top_p, seed, reasoning effort, tools — go in the JSON body and override the server defaults.

The token budget covers thinking, not just the answer

This model reasons before it replies and those tokens count against the budget, so a budget that runs out mid-thought does not shorten the answer, it removes it: the reply comes back with finish_reason: "length", an empty content, and the partial reasoning in reasoning_content, which most OpenAI clients do not display. The per-request default is 8192, which finished every ordinary prompt we measured with room to spare; the ceiling is HALOGEN_MAX_TOKENS_CAP.

Any of three field names works, and they mean the same thing here: max_completion_tokens (current OpenAI Chat Completions), max_output_tokens (OpenAI Responses), or max_tokens (deprecated upstream, still widely sent). Send one, or send several as long as they agree; two different values is a 400 rather than a guess about which you meant. /health lists all three under token_budget_aliases and reports the default as max_tokens_default.

If a reply looks empty or cut off, read finish_reason first: "stop" means you have the whole answer, "length" means you ran out of budget. Pass a larger budget, or "reasoning_effort": "low" to make the model think less.

The thinking controls — reasoning_effort, enable_thinking, preserve_thinking — are accepted at the top level of the request and also nested under chat_template_kwargs, the vLLM/SGLang spelling. Both reach the same template; if you send both they must agree, and an unsupported key under chat_template_kwargs is a 400 rather than a silent drop. Note that the template's own effort default is xhigh, so a request that sends nothing thinks at the most expensive setting. "reasoning_effort": "none" is thinking off, the same as "enable_thinking": false. The OpenAI developer role is accepted and rendered as system.

response_format with json_object or json_schema is a 400, not a silent drop: this build has no constrained decoding, so a schema could not be enforced and you would get prose with a 200. /health lists it under not_implemented. Prompt for JSON and validate the reply.

Every response carries a timings object in llama-server's shape (prompt_n, predicted_n, prompt_ms, predicted_ms, the two rates, cache_n, draft_n, draft_n_accepted), copied from the engine's own accounting, so llama-swap and similar routers show real prefill and decode rates. prompt_n is the count the engine processed — the prompt minus what the cache covered — because that is what prompt_ms measures.

Streaming responses send a : keepalive SSE comment after HALOGEN_SSE_KEEPALIVE_S (10) seconds of silence, so a long prefill does not trip a client's body timeout; and idle keep-alive connections live HALOGEN_KEEPALIVE_TIMEOUT (300) seconds rather than uvicorn's 5.

Concurrency, and the one trap

By default halogen serves one request at a time with speculative decoding on. That is the right setting for a single user: you get ~31 t/s.

Raising HALOGEN_KV_SLOTS lets several sequences be resident at once and raises aggregate throughput to about 49 t/s at 8 concurrent requests — but speculation and batching are currently mutually exclusive. With more than one slot the drafter is off, so each individual stream runs at serial speed (~6 t/s at 8 slots). One user is much better off with the default; a shared server with steady concurrent load is better off with slots.

The trap: the KV pool costs slots x slot_ctx x 64 KiB, so raising slots without lowering the per-slot context multiplies the allocation. Eight slots at the native 262,144 context asks for 137 GB and will not fit. Keep the product at or below the native context:

KV_SLOTSSLOT_CTXpool
126214417.2 GB (default)
213107217.2 GB
46553617.2 GB
83276817.2 GB
8262144137 GB — will not fit

A prompt longer than SLOT_CTX is a hard error naming the limit. It is never silently truncated.

Raising the output cap

HALOGEN_MAX_TOKENS_CAP and HALOGEN_QUEUE_TIMEOUT are coupled and should not be moved independently. The cap bounds how long one request can hold the GPU; the timeout bounds how long the next client waits for it. If a full-length request can outlast the timeout, everyone queued behind it gets a 503.

At high reasoning effort decode runs around 10 t/s, so:

capworst-case requestneeds a timeout above
4,0966.8 min410 s
16,38427.3 min1,640 s
32,76854.6 min3,280 s
65,536 (default)109.2 min6,550 s, and the default timeout is 7,200 s

The cap is a ceiling on what a client may ask for, not a promise about throughput. Almost nothing reaches it: the model stops on its own when the answer is done. It is set high so that a long reasoning problem is not cut off by server policy, and the timeout is set above it so that a client who does ask for a full-length reply does not 503 the next one in the queue. Lower both together if you would rather bound how long one request can hold the GPU.

Asking for more than the cap returns a 400 naming the limit. It is never silently truncated — a truncated response and a model that stopped on its own both end with finish_reason: "length", so a client cannot tell them apart.

Prompt cache

HALOGEN_CACHE_MB is empty by default, meaning auto: the engine sizes the cache from available memory at startup. That suits a machine dedicated to serving. Set an explicit value in MB to pin it, or 0 to disable.

One caveat if you pin it: a single full-context entry is about 18.4 GB at 262K, so a small explicit budget produces a cache that reports itself enabled and never actually hits. The engine warns at startup when this happens.

Warm answers are byte-identical to cold ones by construction.

Requirements

  • AMD Strix Halo (gfx1151) — Ryzen AI Max+ 395 or equivalent. The build hard-rejects every other architecture; this will not run on your discrete GPU, and that is deliberate.
  • 128 GB unified memory recommended. The checkpoint is 35.9 GB and is mapped, not copied.
  • ROCm-capable kernel with /dev/kfd and /dev/dri accessible.
  • A checkpoint and a tokenizer, mounted at /models and /tokenizer — see Get the weights. The tokenizer directory must be flat; the published model repository is already flat, so this only bites if you point at a HuggingFace cache snapshot, whose entries are symlinks into a sibling blobs/ and dangle inside a container.

Modes

commandwhat it does
(default)engine + API in one container, one published port
engine / apisplit roles for a two-container deployment
benchten real prompt shapes over HTTP
sweeppp/tg size sweep

The engine's token protocol has no authentication. In the default mode it binds loopback inside the container and only the API port is published. If you split the roles, keeping the engine port unpublished is your responsibility.


Honest limits

  • One GPU target. gfx1151 only, by construction.
  • Text only. The model has a vision encoder; halogen does not use it.
  • Unassisted decode is not our strong suit — see the comparison above.
  • Cold time-to-first-token at very long context is slow. A genuinely cold 262K prompt is a multi-minute prefill. The prompt cache makes the second turn fast; it cannot make the first one fast.
  • One default is not byte-identical to the engine's built-in one. The image ships full W4A4 promotion, worth +9% prefill, against about −0.45 pt top-1 aggregate (better at deep context, worse in the first ~12%). It does not affect the guarantees above — speculation is still exact against serial greedy, warm cache still matches cold, batched still matches solo. Roll it back with one environment variable; see docs/FLAGS.md.
  • The comparison table is cross-published, not head-to-head. We have not run the other engines ourselves on our box under matched settings. When we do, we will publish whatever it says.

License

Free for any use, including commercial. Unmodified redistribution permitted. Benchmark publication expressly permitted. See LICENSE and THIRD-PARTY-NOTICES, both also at /licenses inside the image.

Model weights are not included and are not covered by that license. They are obtained separately and licensed by their original authors.

amd
gfx1151
gpu-inference
hip
inference-engine
llm
llm-serving
local-llm
openai-api
qwen
qwen3
radeon
rocm
ryzen-ai
speculative-decoding
strix-halo

Significant stargazers

Lucas Charles

85 followers · starred Sep 2026

Falco

9 followers · starred Sep 2026

peonist-ai/halogen-server

The fastest way to run Qwen3.8 27B on Strix Halo (gfx1151)

Python

121

11 commits

updated Sep 14, 2026

See the code

README

halogen — the peon's inference engine

halogen

The fastest way to run Qwen3.8-27B on AMD Strix Halo — at higher precision than anything that comes close.

Every kernel is written for this one GPU and this one model family. No general-purpose runtime, no portability layer, no fallback path — which is why it can do things a general engine cannot, and why it runs on exactly one piece of silicon.

On a 32K prompt with a 256-token answer, against the fastest numbers anyone else has published for this model on this hardware:

prefilldecodetotal
halogen (6.32 bpw)57.9 s8.1 s66.0 s
KyaniteLabs (Q4_K_XL)84.0 s8.5 s92.5 s
q38rocm (4.26 bpw)133.7 s7.8 s141.5 s

2.1× faster than q38rocm and 1.4× faster than KyaniteLabs end-to-end, while carrying ~1.5× their weight precision. Prefill is where that is won, and on any prompt with real context prefill is most of the wall clock.

Output is also byte-identical to serial greedy decode — speculation here is a pure speed optimization, verified on every release, not a quality trade.

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -v /path/to/models:/models:ro -v /path/to/tokenizer:/tokenizer:ro \
  ghcr.io/peonist-ai/halogen:0.1.4

An OpenAI-compatible endpoint comes up on :8731.

On Docker rather than Podman, replace --group-add keep-groups with --group-add video --group-add render. keep-groups is a Podman keyword that Docker does not understand: Docker resolves --group-add names against the container's /etc/group and fails with unable to find group keep-groups.


Get the weights

The image contains no model weights — it is 3.5 GB of engine, and the checkpoint is 35.9 GB. Download it once and mount it:

pip install -U "huggingface_hub[cli]"
hf download peonist-ai/halogen-qwen3.8-27b \
  --local-dir ~/halogen-models

That repository carries both the .hgn checkpoint and a flat tokenizer directory, so there is nothing to assemble by hand:

~/halogen-models/
  qwen3.8-27b-p1w4d-d2.hgn      35.9 GB   the checkpoint
  tokenizer/                              tokenizer.json, chat template, ...

Then point the container at both:

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -v ~/halogen-models:/models:ro \
  -v ~/halogen-models/tokenizer:/tokenizer:ro \
  ghcr.io/peonist-ai/halogen:0.1.4

Or let it fetch them for you

If you would rather not download separately, set HALOGEN_DOWNLOAD and the container fetches the weights on first start:

podman run --rm -p 8731:8731 \
  --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -e HALOGEN_DOWNLOAD=peonist-ai/halogen-qwen3.8-27b \
  -e HALOGEN_TOKENIZER=/models/tokenizer \
  -v ~/halogen-models:/models \
  ghcr.io/peonist-ai/halogen:0.1.4

Two differences from the manual route. The models volume is mounted read-write — it has to be, to download into. And there is only one mount: the download brings the tokenizer with it, so HALOGEN_TOKENIZER points inside /models rather than at a second volume. Mounting ~/halogen-models/tokenizer here would fail on a first run, because the container runtime would create it as an empty directory before the download had a chance to populate it. It fires only when the checkpoint is actually missing, so restarts do not re-download, and an interrupted transfer resumes rather than starting over.

With HALOGEN_DOWNLOAD unset, the container makes no outbound network connections at all — no telemetry, no license check, no model fetch. If the checkpoint is not on disk where HALOGEN_CHECKPOINT points, it says so and exits rather than reaching for the network. That default is deliberate: a 35.9 GB transfer should not begin because someone ran podman run to see what would happen.

Model weights are licensed separately from the engine by their original authors; see the model repository for those terms.


Performance

Measured on a Ryzen AI Max+ 395 (Radeon 8060S, 128 GB LPDDR5X), ROCm 7.14.0, checkpoint p1w4d-d2 (~6.3 bits/weight effective at decode), 262,144 context.

Prefill

All three measured over the HTTP endpoint with the bundled sweep, so they are one instrument rather than a mix.

testt/s
pp512620
pp2048710
pp32768566

Decode

Over the HTTP endpoint, ten prompt shapes, greedy, DFlash2 drafter:

mean t/srange
DFlash2 (default)31.7120.8 – 44.2
MTP26.8819.8 – 34.2
serial (no speculation)10.58—

Aggregate throughput at 8 concurrent requests: 48.6 t/s (4.87×).

Read the range, not just the mean

halogen's decode rate is a distribution, not a number. Speculative decoding accepts more drafted tokens when the text is predictable, so the same build on the same hardware does:

  • prose / chat: 20.8 – 23.5 t/s
  • procedures: 25.6 – 38.7 t/s
  • code / proofs: 32.7 – 44.2 t/s

A single headline figure hides a 2× spread. Any decode number quoted from this project — by us or anyone else — should name the prompt set that produced it, or it is not reproducible. bench prints the mean; sweep prints mean, standard deviation, and min–max, deliberately.

An engine without speculative decoding has a content-independent decode rate and can honestly quote one number. We can't.


How it compares

Published numbers from other projects running the same model on the same silicon. These are their figures on their configurations, not a head-to-head we ran — quantization, KV-cache settings and context differ, so read this as orientation, not as a controlled benchmark.

halogenq38rocmKyaniteLabs
backendcustom HIPROCm/RADVllama.cpp
weights~6.3 bpw4.26 bpwUD-Q4_K_XL
prefill @32K566 t/s245 t/s~390 t/s
decode, speculated30.98 mean (20.8–44.2)30.56 – 36.04prose 11–24, code 29–40
decode, unassisted (diagnostic)10.58 t/s14.02 t/s—

Where we win: prefill, by 1.4–2.3×, and end-to-end on any prompt with real context. That is what the engine was built for.

Where we lose: unassisted decode — and that row is a diagnostic, not a product configuration. Nobody ships serial decode; every project in this table runs speculation by default. The gap is also not kernel quality: decode is bandwidth-bound, q38rocm streams ~17 GB/token against our 23.5, and fewer bits is simply faster. We spend those bits deliberately (see QUANT.md) — the only 4-bit tensors in our trunk are ones somebody else calibrated, and the aggressive technique is fenced to prefill where it never touches token generation.

Our batch-1 decode is at the hardware wall. 10.58 t/s × 23.51 GB/token = 249 GB/s against a measured ceiling of 240 GB/s. There is no kernel win left there for anyone; the levers are fewer bits, better draft acceptance, and batching.

On the 148–163 t/s figure circulating for llama.cpp on this hardware: that is an ngram-repetition artifact on back-to-back identical runs, and KyaniteLabs — whose benchmark it is — says so plainly and warns against quoting it for chat. Their honest conversational numbers are in the table. We think that is the right way to publish, and we have tried to match it.


Benchmark it yourself

The image ships both benchmarks. No fixtures, no extra downloads, no cooperation from us required.

# ten real prompt shapes over the HTTP endpoint — the number of record
podman run --rm --device /dev/kfd --device /dev/dri --group-add keep-groups \
  --security-opt seccomp=unconfined --ipc=host \
  -v /path/to/models:/models:ro -v /path/to/tokenizer:/tokenizer:ro \
  ghcr.io/peonist-ai/halogen:0.1.4 bench dflash2 256 low 3

# llama-bench-shaped pp/tg sweep, for putting a number beside another engine
podman run --rm ... ghcr.io/peonist-ai/halogen:0.1.4 \
  sweep -p 512,2048,8192 -n 128,256 -d dflash2,mtp -r 3

sweep --json emits machine-readable output.

Publishing the results is expressly permitted — no approval, no notice, no prior review. We only ask that figures name the version and the prompt set, for the reason above. That request is not a licensing condition.


What it does

Byte-identical speculative decoding. Draft-then-verify commits only tokens the full model would have produced, so output is bit-for-bit identical to serial greedy decode. This is gated on every release across all three drafters — not asserted, measured. Speculation here is a pure speed optimization with no quality cost, and you can turn it off per request to check.

Native 262,144-token context, with decode that barely degrades at depth — Gated DeltaNet carries O(1) state, so 48 of 64 layers have no KV cache at all.

Prompt cache — a follow-up turn on a long conversation resumes instead of re-prefilling, worth roughly 20× on time-to-first-token at 32K. Warm answers are byte-identical to cold ones by construction.

Batched decode — 8 concurrent sequences, 4.87× aggregate, each byte-identical to running alone. Off by default, and it trades away speculation when enabled; see Configuration before turning it on.

OpenAI-compatible API — /v1/chat/completions, /v1/completions, streaming, tool calling, sampling with seeds, reasoning-effort control.

Three selectable drafters — dflash2 (default), mtp, serial. Choose per request; output is identical, only speed changes.


Configuration

Everything is set by environment variable — there is no config file. The complete list of levers, with defaults and whether each one can change output, is in docs/FLAGS.md. These are the ones most people touch:

variabledefaultwhat it does
HALOGEN_CHECKPOINT/models/qwen3.8-27b-p1w4d-d2.hgnwhich checkpoint to load
HALOGEN_TOKENIZER/tokenizerflat tokenizer directory
HALOGEN_API_PORT8731the published port
HALOGEN_DRAFTER2 (DFlash2)default drafter: 0 serial, 1 MTP, 2 DFlash2
HALOGEN_CACHE_MBautoprompt cache budget in MB — a byte budget, not a slot count; 0 disables
HALOGEN_MAX_TOKENS_CAP65536largest max_tokens a request may ask for — over it is a 400, never a silent truncation
HALOGEN_QUEUE_TIMEOUT7200seconds a queued request will wait — coupled to the cap, see below
HALOGEN_KV_SLOTS1concurrent resident sequences — see below
HALOGEN_SLOT_CTX262144context each slot holds — see below

Per-request settings — drafter, temperature, top_p, seed, reasoning effort, tools — go in the JSON body and override the server defaults.

The token budget covers thinking, not just the answer

This model reasons before it replies and those tokens count against the budget, so a budget that runs out mid-thought does not shorten the answer, it removes it: the reply comes back with finish_reason: "length", an empty content, and the partial reasoning in reasoning_content, which most OpenAI clients do not display. The per-request default is 8192, which finished every ordinary prompt we measured with room to spare; the ceiling is HALOGEN_MAX_TOKENS_CAP.

Any of three field names works, and they mean the same thing here: max_completion_tokens (current OpenAI Chat Completions), max_output_tokens (OpenAI Responses), or max_tokens (deprecated upstream, still widely sent). Send one, or send several as long as they agree; two different values is a 400 rather than a guess about which you meant. /health lists all three under token_budget_aliases and reports the default as max_tokens_default.

If a reply looks empty or cut off, read finish_reason first: "stop" means you have the whole answer, "length" means you ran out of budget. Pass a larger budget, or "reasoning_effort": "low" to make the model think less.

The thinking controls — reasoning_effort, enable_thinking, preserve_thinking — are accepted at the top level of the request and also nested under chat_template_kwargs, the vLLM/SGLang spelling. Both reach the same template; if you send both they must agree, and an unsupported key under chat_template_kwargs is a 400 rather than a silent drop. Note that the template's own effort default is xhigh, so a request that sends nothing thinks at the most expensive setting. "reasoning_effort": "none" is thinking off, the same as "enable_thinking": false. The OpenAI developer role is accepted and rendered as system.

response_format with json_object or json_schema is a 400, not a silent drop: this build has no constrained decoding, so a schema could not be enforced and you would get prose with a 200. /health lists it under not_implemented. Prompt for JSON and validate the reply.

Every response carries a timings object in llama-server's shape (prompt_n, predicted_n, prompt_ms, predicted_ms, the two rates, cache_n, draft_n, draft_n_accepted), copied from the engine's own accounting, so llama-swap and similar routers show real prefill and decode rates. prompt_n is the count the engine processed — the prompt minus what the cache covered — because that is what prompt_ms measures.

Streaming responses send a : keepalive SSE comment after HALOGEN_SSE_KEEPALIVE_S (10) seconds of silence, so a long prefill does not trip a client's body timeout; and idle keep-alive connections live HALOGEN_KEEPALIVE_TIMEOUT (300) seconds rather than uvicorn's 5.

Concurrency, and the one trap

By default halogen serves one request at a time with speculative decoding on. That is the right setting for a single user: you get ~31 t/s.

Raising HALOGEN_KV_SLOTS lets several sequences be resident at once and raises aggregate throughput to about 49 t/s at 8 concurrent requests — but speculation and batching are currently mutually exclusive. With more than one slot the drafter is off, so each individual stream runs at serial speed (~6 t/s at 8 slots). One user is much better off with the default; a shared server with steady concurrent load is better off with slots.

The trap: the KV pool costs slots x slot_ctx x 64 KiB, so raising slots without lowering the per-slot context multiplies the allocation. Eight slots at the native 262,144 context asks for 137 GB and will not fit. Keep the product at or below the native context:

KV_SLOTSSLOT_CTXpool
126214417.2 GB (default)
213107217.2 GB
46553617.2 GB
83276817.2 GB
8262144137 GB — will not fit

A prompt longer than SLOT_CTX is a hard error naming the limit. It is never silently truncated.

Raising the output cap

HALOGEN_MAX_TOKENS_CAP and HALOGEN_QUEUE_TIMEOUT are coupled and should not be moved independently. The cap bounds how long one request can hold the GPU; the timeout bounds how long the next client waits for it. If a full-length request can outlast the timeout, everyone queued behind it gets a 503.

At high reasoning effort decode runs around 10 t/s, so:

capworst-case requestneeds a timeout above
4,0966.8 min410 s
16,38427.3 min1,640 s
32,76854.6 min3,280 s
65,536 (default)109.2 min6,550 s, and the default timeout is 7,200 s

The cap is a ceiling on what a client may ask for, not a promise about throughput. Almost nothing reaches it: the model stops on its own when the answer is done. It is set high so that a long reasoning problem is not cut off by server policy, and the timeout is set above it so that a client who does ask for a full-length reply does not 503 the next one in the queue. Lower both together if you would rather bound how long one request can hold the GPU.

Asking for more than the cap returns a 400 naming the limit. It is never silently truncated — a truncated response and a model that stopped on its own both end with finish_reason: "length", so a client cannot tell them apart.

Prompt cache

HALOGEN_CACHE_MB is empty by default, meaning auto: the engine sizes the cache from available memory at startup. That suits a machine dedicated to serving. Set an explicit value in MB to pin it, or 0 to disable.

One caveat if you pin it: a single full-context entry is about 18.4 GB at 262K, so a small explicit budget produces a cache that reports itself enabled and never actually hits. The engine warns at startup when this happens.

Warm answers are byte-identical to cold ones by construction.

Requirements

  • AMD Strix Halo (gfx1151) — Ryzen AI Max+ 395 or equivalent. The build hard-rejects every other architecture; this will not run on your discrete GPU, and that is deliberate.
  • 128 GB unified memory recommended. The checkpoint is 35.9 GB and is mapped, not copied.
  • ROCm-capable kernel with /dev/kfd and /dev/dri accessible.
  • A checkpoint and a tokenizer, mounted at /models and /tokenizer — see Get the weights. The tokenizer directory must be flat; the published model repository is already flat, so this only bites if you point at a HuggingFace cache snapshot, whose entries are symlinks into a sibling blobs/ and dangle inside a container.

Modes

commandwhat it does
(default)engine + API in one container, one published port
engine / apisplit roles for a two-container deployment
benchten real prompt shapes over HTTP
sweeppp/tg size sweep

The engine's token protocol has no authentication. In the default mode it binds loopback inside the container and only the API port is published. If you split the roles, keeping the engine port unpublished is your responsibility.


Honest limits

  • One GPU target. gfx1151 only, by construction.
  • Text only. The model has a vision encoder; halogen does not use it.
  • Unassisted decode is not our strong suit — see the comparison above.
  • Cold time-to-first-token at very long context is slow. A genuinely cold 262K prompt is a multi-minute prefill. The prompt cache makes the second turn fast; it cannot make the first one fast.
  • One default is not byte-identical to the engine's built-in one. The image ships full W4A4 promotion, worth +9% prefill, against about −0.45 pt top-1 aggregate (better at deep context, worse in the first ~12%). It does not affect the guarantees above — speculation is still exact against serial greedy, warm cache still matches cold, batched still matches solo. Roll it back with one environment variable; see docs/FLAGS.md.
  • The comparison table is cross-published, not head-to-head. We have not run the other engines ourselves on our box under matched settings. When we do, we will publish whatever it says.

License

Free for any use, including commercial. Unmodified redistribution permitted. Benchmark publication expressly permitted. See LICENSE and THIRD-PARTY-NOTICES, both also at /licenses inside the image.

Model weights are not included and are not covered by that license. They are obtained separately and licensed by their original authors.

amd
gfx1151
gpu-inference
hip
inference-engine
llm
llm-serving
local-llm
openai-api
qwen
qwen3
radeon
rocm
ryzen-ai
speculative-decoding
strix-halo

Significant stargazers

Lucas Charles

85 followers · starred Sep 2026

Falco

9 followers · starred Sep 2026