Qwen3.8-27B-GSQ3 on one RTX 4080: 100k context with rk4v4-e8, up to 2720 tok/s prefill and 262 tok/s generation
C++
0
1,265 commits
updated Oct 3, 2026
Qwen3.8-27B on a single 16 GB RTX 4080, at 100K context, with vision and speculative decoding.
NInfer-4080 runs the ISTA-DASLab Qwen3.8-27B 3-bit GSQ checkpoint on one NVIDIA GeForce RTX 4080 using more of the hardware than the general-purpose engines do: up to 2,720 tok/s prefill and 262 tok/s generation are measured on this card (sweep below), with the full 100,000-token context, vision, and MTP3/DFlash2 speculative decoding profiles resident at once. The artifact and a binaries-only container image are published, so there is nothing to convert before you run it.
It is an sm_89 port of NInfer-4090, which derives
from NInfer-3090 and
Neroued/ninfer, a specialized C++20/CUDA inference engine.
This fork registers the Q3G128_F16S 3-bit weight scheme and the gsq3 weights profile because
no registered Q4-or-wider allocation fits a 16 GB card at 100K context. The engine itself is
inherited: paged KV, compatible-prefix reuse, CUDA Graphs, speculative decoding, reasoning-effort
control, the OpenAI/Anthropic-compatible APIs, and the ReplaySSM state transactions all work as
documented in docs/.
roofkid/ninfer-4080:gsq3 (Docker Hub, tags gsq3 and 0.6.1-rtx4080).docker run, serve.Requirements: an RTX 4080 (16 GB, sm_89) with a CUDA 13.1-or-newer driver, and Docker with the
NVIDIA Container Toolkit. A source build needs CUDA 13.1 and only accepts
CMAKE_CUDA_ARCHITECTURES=89.
./scripts/download-qwen38-gsq3.sh # .bat on Windows
The script fetches qwen3_8_27b_gsq3.ninfer from the pinned Hugging Face revision, resumes an
interrupted download, and verifies the SHA-256. The artifact lands in models/.
docker run --rm --gpus all -p 8080:8080 \
-v "$PWD/models/qwen3_8_27b_gsq3.ninfer:/models/model.ninfer:ro" \
roofkid/ninfer-4080:gsq3 \
ninfer-serve /models/model.ninfer \
--host 0.0.0.0 --port 8080 \
--max-context 102400 --kv-capacity 102400 --kv-dtype rk4v4-e8 \
--max-concurrency 1 --max-pending-requests 16 --prefill-chunk 2688 \
--host-kv-mib 4096 \
--spec mtp --draft-tokens 3 --lm-head-draft \
--vision --preserve-thinking
scripts/run-ninfer-4080.sh (or .bat) does the same in one step and pulls the image for you.
scripts/run-ninfer-4080-dflash2.{sh,bat} serves the DFlash2 K=7 profile instead. Both pick the
artifact up from out/ or models/, accept NINFER_IMAGE, NINFER_PORT, NINFER_BIND,
NINFER_KV_DTYPE, NINFER_CONTEXT, and NINFER_API_KEY, and fall back to a local source build
when the image cannot be pulled (force one with NINFER_BUILD=1).
curl -s http://localhost:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Say hello in five words."}],"max_tokens":64}'
The server speaks OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages; see
HTTP serving and CLI usage. For the CLI without Docker, build
from source (cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DNINFER_BUILD_APPS=ON,
then cmake --build build --parallel) and run ./build/apps/ninfer models/qwen3_8_27b_gsq3.ninfer --prompt "..." --kv-dtype rk4v4-e8 --spec mtp --draft-tokens 3 --lm-head-draft.
Conditions: single request, rk4v4-e8 KV, --prefill-chunk 1024, the 131,072-token tiled corpus,
one warmup and one measured repetition per point. Acceptance is a tiled-corpus fixture property
(repeated text), not a model result.
| Depth | Prefill t/s | MTP3 decode t/s | DFlash2 K=7 decode t/s |
|---|---|---|---|
| 8K | 2,719.9 | 151.2 | 166.7 |
| 32K | 2,424.9 | 141.7 | 262.3 |
| 64K | 2,125.5 | 130.7 | 239.1 |
| 98K | 1,895.1 | 122.3 | 212.7 |
At the documented 100K prefill profile (--prefill-chunk 2688) the same build measures
1,971.4 tok/s, and 2,470.0 tok/s at 32,768 tokens. Against the llama.cpp IQ3_S reference on
this card (1,043 tok/s at 100K), the fork's 100K prefill now leads, and decode leads at every
depth. In real use the maintainer sees the 2K+ prefill rates on long prompts and roughly 150-200
tok/s decode on coding, about 100 tok/s on prose. Raw session logs and method: the port's working
record is docs/maintainer/rtx-4080-plan.md section 11 while
the port is active; the same numbers are repeated in the model card.
--max-context 102400, --host-kv-mib 4096, rk4v4-e8): about
11.2 GiB of device weights; the KV and runtime reservation validate before the server listens,
leaving roughly 0.9 GiB free.--spec dflash2 --draft-tokens 7 --lm-head-draft).rk4v4-e8 is the 100K accuracy profile. int8 and rk8v4 trade context for
accuracy and fit below 100K. The pinned host pools hold one deep checkpoint so rewrites reuse
the prefix instead of re-prefilling.groupwise-int artifact does not load on this card; the 3-bit GSQ3
artifact is the registered 4080 fit.The maintainer's own runs keep MBPP in the 90-92% range and HumanEval in the 95-96% range. The
weight-quality anchor on the fork's 1M-token corpus is 4.596095 quick perplexity (INT8 KV),
identical to the value recorded at conversion, and an apples-to-apples MBPP run put the artifact
at 90% against 90-92% for the beellama kvarn5/5/kvarn4/4 reference (within noise). Small
deviations within measurement noise are possible, mainly from KV quantization: rk4v4-e8 is the
validated 100K profile, and int8 or rk8v4 trade context for accuracy if you want the
conservative option.
After watching the community build NInfer for the 5090, 4090, and 3090, there was still nothing for a 16 GB RTX 4080. This fork is that missing piece. It was built as a pet project by the maintainer, who has 20+ years of software engineering and architecture experience but no GPU kernel background: the work was approached from a requirements and product-owner perspective, with a focus on engineering practices, measurable changes, and business decisions rather than hand-judging kernel code. The guiding principles were: fit the 16 GB card; start from the GSQ checkpoint the maintainer already used daily (its reasoning quality is discussed in the ByteShape GSQ article); use DFlash2 speculative decoding; reach 100K+ context; beat the general-purpose engines on prefill and decode; and re-measure quality after every change. Agent tooling drove the implementation (Pi as the harness with DeepSeek V4.1 Flash), at a total model spend of about $13.
rk4v4-e8 here is indistinguishable from kvarn5/5 for the
maintainer, with no "fast garbage" effect, and it is what makes 100K fit. At these speeds the
maintainer went back to xhigh thinking.| Field | Value |
|---|---|
| Filename | qwen3_8_27b_gsq3.ninfer |
| Size | 13,330,776,576 bytes (12.41 GiB) |
| SHA-256 | c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7 |
| NInfer identity | qwen3.8-27b / gsq3 / qwen3_8_27b |
| Source | ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ @ b5ce0b76, Apache-2.0 |
The 323 packed matrices from the publisher's checkpoint (the 320 Text-body matrices, the token
embedding, the full output head, and the draft head) are a lossless repack: every code and
group scale is copied unchanged, with only the publisher's unsigned code + 4 bit-plane layout
inverted into two's-complement fields. The only value deviation is 218 of 240,271,360 bf16 scales
that are not binary16-exact; all are subnormals rounded with an error bounded by 2**-25 (max
2.98e-8). The publisher's task evaluations describe the represented Text and vocabulary weights.
The MTP layer, Vision tower, and DFlash2 companion come from the official BF16 sources and are
quantized by the converter. The full model card, provenance table, and verification evidence are
in model-cards/Qwen3.8-27B-GSQ3-NInfer.
Verify a downloaded file with:
printf '%s %s\n' \
'c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7' \
'qwen3_8_27b_gsq3.ninfer' | sha256sum --check
Q3G128_F16S (symmetric 3-bit codes, one FP16 scale
per 128 values, 3.125 bpw) in the artifact codec, the converter/verifier, the C++ binder, and
the Qwen38Gsq3 profile, plus Q3 execution leaves for every Text site.pp100000 1,710 → 1,971 tok/s; decode +12% at K=7).Q4G64_F16S and bound in the gsq3 identity, which
raised the K=7 caps from 57K to 100,000 tokens text-only.scripts/run-ninfer-4080.sh, .bat, the artifact download
scripts, and the published image.weights_id = gsq3 is not registered upstream).sm_89 only.Qwen3.8-27B has three trained reasoning depths plus an off switch. OpenAI Chat Completions
accepts a top-level reasoning_effort field (low, medium, xhigh) and a top-level
enable_thinking boolean; hidden reasoning returns separately as message.reasoning_content.
Only those three levels are accepted, plus none to turn thinking off. high, minimal, and
max are rejected as reasoning_effort_not_supported, so a client that offers a high setting
must map it to xhigh. A token budget for reasoning is separate from the effort level. It is
set only through the Anthropic Messages path (thinking.budget_tokens) or server-wide with
--default-thinking-budget N; the OpenAI paths have no field for it. Without a budget,
reasoning is bounded only by the request's max_tokens, which is what the model card
recommends, and the model_thinking_tokens field of the request JSONL reads zero, because that
counter runs only under a budget. The chat_template_kwargs request field of llama.cpp is not
supported and is rejected. For the CLI, pass --reasoning-effort or --no-thinking. Sampling
defaults come from the model card and switch with the thinking mode: temperature=1.0,
top_p=0.95, top_k=20 in thinking mode; temperature=0.7, top_p=0.80, top_k=20,
presence_penalty=1.5 in non-thinking mode.
OpenAI Chat Completions, OpenAI Responses with streaming and local continuation state, Anthropic Messages, prompt-rendered function tools with parsed tool calls, compatible-prefix reuse, and JSONL request logs. See HTTP serving and CLI usage.
If you have another 16 GB RTX 4xxx card (4070 Ti Super, 4060 Ti 16 GB, 4080 Super, ...) I would like to know whether this works there and what speeds you see, because it is hard to judge how tied to the 4080 the tuning is. Issues and measurement reports are welcome on the fork's tracker. If you have a 4080, enjoy.
Thanks to everyone who worked on NInfer before this fork; you provided a stable base. A special shout-out to sergiuszm for the NInfer-4090 SM_89 heavy lifting, and to ISTA-DASLab for GSQ - cheers to Austria from Germany.
sm_120a), Apache-2.0.rtx4080-port.rk8v4, rk4v4, rk4v4-e8, rk2v4-e8) and the E8 codecs, merged
with authorship preserved.If this fork is useful to you and you would like to support its continued development, you can buy me a coffee. The upstream engine project also accepts support at ko-fi.com/neroued.
Support is entirely voluntary. It is not a purchase or investment and does not come with financial returns, promised services or features, or a role in project decisions. The project's direction, priorities, technical choices, and release schedule remain independently determined by the maintainer.
Apache License 2.0. See LICENSE.
C++
61.2%
Cuda
27.8%
Python
9.9%
Qwen3.8-27B-GSQ3 on one RTX 4080: 100k context with rk4v4-e8, up to 2720 tok/s prefill and 262 tok/s generation
C++
0
1,265 commits
updated Oct 3, 2026
Qwen3.8-27B on a single 16 GB RTX 4080, at 100K context, with vision and speculative decoding.
NInfer-4080 runs the ISTA-DASLab Qwen3.8-27B 3-bit GSQ checkpoint on one NVIDIA GeForce RTX 4080 using more of the hardware than the general-purpose engines do: up to 2,720 tok/s prefill and 262 tok/s generation are measured on this card (sweep below), with the full 100,000-token context, vision, and MTP3/DFlash2 speculative decoding profiles resident at once. The artifact and a binaries-only container image are published, so there is nothing to convert before you run it.
It is an sm_89 port of NInfer-4090, which derives
from NInfer-3090 and
Neroued/ninfer, a specialized C++20/CUDA inference engine.
This fork registers the Q3G128_F16S 3-bit weight scheme and the gsq3 weights profile because
no registered Q4-or-wider allocation fits a 16 GB card at 100K context. The engine itself is
inherited: paged KV, compatible-prefix reuse, CUDA Graphs, speculative decoding, reasoning-effort
control, the OpenAI/Anthropic-compatible APIs, and the ReplaySSM state transactions all work as
documented in docs/.
roofkid/ninfer-4080:gsq3 (Docker Hub, tags gsq3 and 0.6.1-rtx4080).docker run, serve.Requirements: an RTX 4080 (16 GB, sm_89) with a CUDA 13.1-or-newer driver, and Docker with the
NVIDIA Container Toolkit. A source build needs CUDA 13.1 and only accepts
CMAKE_CUDA_ARCHITECTURES=89.
./scripts/download-qwen38-gsq3.sh # .bat on Windows
The script fetches qwen3_8_27b_gsq3.ninfer from the pinned Hugging Face revision, resumes an
interrupted download, and verifies the SHA-256. The artifact lands in models/.
docker run --rm --gpus all -p 8080:8080 \
-v "$PWD/models/qwen3_8_27b_gsq3.ninfer:/models/model.ninfer:ro" \
roofkid/ninfer-4080:gsq3 \
ninfer-serve /models/model.ninfer \
--host 0.0.0.0 --port 8080 \
--max-context 102400 --kv-capacity 102400 --kv-dtype rk4v4-e8 \
--max-concurrency 1 --max-pending-requests 16 --prefill-chunk 2688 \
--host-kv-mib 4096 \
--spec mtp --draft-tokens 3 --lm-head-draft \
--vision --preserve-thinking
scripts/run-ninfer-4080.sh (or .bat) does the same in one step and pulls the image for you.
scripts/run-ninfer-4080-dflash2.{sh,bat} serves the DFlash2 K=7 profile instead. Both pick the
artifact up from out/ or models/, accept NINFER_IMAGE, NINFER_PORT, NINFER_BIND,
NINFER_KV_DTYPE, NINFER_CONTEXT, and NINFER_API_KEY, and fall back to a local source build
when the image cannot be pulled (force one with NINFER_BUILD=1).
curl -s http://localhost:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Say hello in five words."}],"max_tokens":64}'
The server speaks OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages; see
HTTP serving and CLI usage. For the CLI without Docker, build
from source (cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release -DNINFER_BUILD_APPS=ON,
then cmake --build build --parallel) and run ./build/apps/ninfer models/qwen3_8_27b_gsq3.ninfer --prompt "..." --kv-dtype rk4v4-e8 --spec mtp --draft-tokens 3 --lm-head-draft.
Conditions: single request, rk4v4-e8 KV, --prefill-chunk 1024, the 131,072-token tiled corpus,
one warmup and one measured repetition per point. Acceptance is a tiled-corpus fixture property
(repeated text), not a model result.
| Depth | Prefill t/s | MTP3 decode t/s | DFlash2 K=7 decode t/s |
|---|---|---|---|
| 8K | 2,719.9 | 151.2 | 166.7 |
| 32K | 2,424.9 | 141.7 | 262.3 |
| 64K | 2,125.5 | 130.7 | 239.1 |
| 98K | 1,895.1 | 122.3 | 212.7 |
At the documented 100K prefill profile (--prefill-chunk 2688) the same build measures
1,971.4 tok/s, and 2,470.0 tok/s at 32,768 tokens. Against the llama.cpp IQ3_S reference on
this card (1,043 tok/s at 100K), the fork's 100K prefill now leads, and decode leads at every
depth. In real use the maintainer sees the 2K+ prefill rates on long prompts and roughly 150-200
tok/s decode on coding, about 100 tok/s on prose. Raw session logs and method: the port's working
record is docs/maintainer/rtx-4080-plan.md section 11 while
the port is active; the same numbers are repeated in the model card.
--max-context 102400, --host-kv-mib 4096, rk4v4-e8): about
11.2 GiB of device weights; the KV and runtime reservation validate before the server listens,
leaving roughly 0.9 GiB free.--spec dflash2 --draft-tokens 7 --lm-head-draft).rk4v4-e8 is the 100K accuracy profile. int8 and rk8v4 trade context for
accuracy and fit below 100K. The pinned host pools hold one deep checkpoint so rewrites reuse
the prefix instead of re-prefilling.groupwise-int artifact does not load on this card; the 3-bit GSQ3
artifact is the registered 4080 fit.The maintainer's own runs keep MBPP in the 90-92% range and HumanEval in the 95-96% range. The
weight-quality anchor on the fork's 1M-token corpus is 4.596095 quick perplexity (INT8 KV),
identical to the value recorded at conversion, and an apples-to-apples MBPP run put the artifact
at 90% against 90-92% for the beellama kvarn5/5/kvarn4/4 reference (within noise). Small
deviations within measurement noise are possible, mainly from KV quantization: rk4v4-e8 is the
validated 100K profile, and int8 or rk8v4 trade context for accuracy if you want the
conservative option.
After watching the community build NInfer for the 5090, 4090, and 3090, there was still nothing for a 16 GB RTX 4080. This fork is that missing piece. It was built as a pet project by the maintainer, who has 20+ years of software engineering and architecture experience but no GPU kernel background: the work was approached from a requirements and product-owner perspective, with a focus on engineering practices, measurable changes, and business decisions rather than hand-judging kernel code. The guiding principles were: fit the 16 GB card; start from the GSQ checkpoint the maintainer already used daily (its reasoning quality is discussed in the ByteShape GSQ article); use DFlash2 speculative decoding; reach 100K+ context; beat the general-purpose engines on prefill and decode; and re-measure quality after every change. Agent tooling drove the implementation (Pi as the harness with DeepSeek V4.1 Flash), at a total model spend of about $13.
rk4v4-e8 here is indistinguishable from kvarn5/5 for the
maintainer, with no "fast garbage" effect, and it is what makes 100K fit. At these speeds the
maintainer went back to xhigh thinking.| Field | Value |
|---|---|
| Filename | qwen3_8_27b_gsq3.ninfer |
| Size | 13,330,776,576 bytes (12.41 GiB) |
| SHA-256 | c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7 |
| NInfer identity | qwen3.8-27b / gsq3 / qwen3_8_27b |
| Source | ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ @ b5ce0b76, Apache-2.0 |
The 323 packed matrices from the publisher's checkpoint (the 320 Text-body matrices, the token
embedding, the full output head, and the draft head) are a lossless repack: every code and
group scale is copied unchanged, with only the publisher's unsigned code + 4 bit-plane layout
inverted into two's-complement fields. The only value deviation is 218 of 240,271,360 bf16 scales
that are not binary16-exact; all are subnormals rounded with an error bounded by 2**-25 (max
2.98e-8). The publisher's task evaluations describe the represented Text and vocabulary weights.
The MTP layer, Vision tower, and DFlash2 companion come from the official BF16 sources and are
quantized by the converter. The full model card, provenance table, and verification evidence are
in model-cards/Qwen3.8-27B-GSQ3-NInfer.
Verify a downloaded file with:
printf '%s %s\n' \
'c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7' \
'qwen3_8_27b_gsq3.ninfer' | sha256sum --check
Q3G128_F16S (symmetric 3-bit codes, one FP16 scale
per 128 values, 3.125 bpw) in the artifact codec, the converter/verifier, the C++ binder, and
the Qwen38Gsq3 profile, plus Q3 execution leaves for every Text site.pp100000 1,710 → 1,971 tok/s; decode +12% at K=7).Q4G64_F16S and bound in the gsq3 identity, which
raised the K=7 caps from 57K to 100,000 tokens text-only.scripts/run-ninfer-4080.sh, .bat, the artifact download
scripts, and the published image.weights_id = gsq3 is not registered upstream).sm_89 only.Qwen3.8-27B has three trained reasoning depths plus an off switch. OpenAI Chat Completions
accepts a top-level reasoning_effort field (low, medium, xhigh) and a top-level
enable_thinking boolean; hidden reasoning returns separately as message.reasoning_content.
Only those three levels are accepted, plus none to turn thinking off. high, minimal, and
max are rejected as reasoning_effort_not_supported, so a client that offers a high setting
must map it to xhigh. A token budget for reasoning is separate from the effort level. It is
set only through the Anthropic Messages path (thinking.budget_tokens) or server-wide with
--default-thinking-budget N; the OpenAI paths have no field for it. Without a budget,
reasoning is bounded only by the request's max_tokens, which is what the model card
recommends, and the model_thinking_tokens field of the request JSONL reads zero, because that
counter runs only under a budget. The chat_template_kwargs request field of llama.cpp is not
supported and is rejected. For the CLI, pass --reasoning-effort or --no-thinking. Sampling
defaults come from the model card and switch with the thinking mode: temperature=1.0,
top_p=0.95, top_k=20 in thinking mode; temperature=0.7, top_p=0.80, top_k=20,
presence_penalty=1.5 in non-thinking mode.
OpenAI Chat Completions, OpenAI Responses with streaming and local continuation state, Anthropic Messages, prompt-rendered function tools with parsed tool calls, compatible-prefix reuse, and JSONL request logs. See HTTP serving and CLI usage.
If you have another 16 GB RTX 4xxx card (4070 Ti Super, 4060 Ti 16 GB, 4080 Super, ...) I would like to know whether this works there and what speeds you see, because it is hard to judge how tied to the 4080 the tuning is. Issues and measurement reports are welcome on the fork's tracker. If you have a 4080, enjoy.
Thanks to everyone who worked on NInfer before this fork; you provided a stable base. A special shout-out to sergiuszm for the NInfer-4090 SM_89 heavy lifting, and to ISTA-DASLab for GSQ - cheers to Austria from Germany.
sm_120a), Apache-2.0.rtx4080-port.rk8v4, rk4v4, rk4v4-e8, rk2v4-e8) and the E8 codecs, merged
with authorship preserved.If this fork is useful to you and you would like to support its continued development, you can buy me a coffee. The upstream engine project also accepts support at ko-fi.com/neroued.
Support is entirely voluntary. It is not a purchase or investment and does not come with financial returns, promised services or features, or a role in project decisions. The project's direction, priorities, technical choices, and release schedule remain independently determined by the maintainer.
Apache License 2.0. See LICENSE.
C++
61.2%
Cuda
27.8%
Python
9.9%