This model card is the version-controlled source for roofkid/Qwen3.8-27B-GSQ3-NInfer.
The repository contains a verbatim repack of the 3-bit
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
checkpoint into the native NInfer .ninfer artifact
format, sized to run the full 100,000-token, vision, and speculative-decoding profile on a single
16 GB RTX 4080. It is intended only for the RTX 4080 NInfer fork; it is not a Transformers
checkpoint, Safetensors distribution, or GGUF file.
| Field | Value |
|---|---|
| Filename | qwen3_8_27b_gsq3.ninfer |
| Size | 13,330,776,576 bytes (12.41 GiB) |
| SHA-256 | c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7 |
| Container version | 2 |
| NInfer model ID | qwen3.8-27b |
| NInfer weights ID | gsq3 |
| NInfer target key | qwen3_8_27b |
| Stored objects | 1,190 (1,184 tensors and 6 resources) |
The Text body is 320 matrices in the registered Q3G128_F16S scheme: symmetric 3-bit codes in
[-4, 3], one FP16 multiplier per 128-value group, 3.125 bits per weight. The token embedding,
full output head, and optimized draft head use Q4G64_F16S; the MTP layer and Vision tower follow
NInfer's registered groupwise-int recipes; the DFlash2 companion was requantized to
Q4G64_F16S. The file also carries the
tokenizer, chat-template, generation, and media-processor objects required by NInfer.
Verify a downloaded file with:
printf '%s %s\n' \
'c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7' \
'qwen3_8_27b_gsq3.ninfer' | sha256sum --check
The 323 packed matrices that come from the publisher's checkpoint β the 320 Text-body matrices,
the token embedding, the full output head, and the draft head (a row gather of the output head) β
are a lossless repack: every code and every group scale was copied unchanged, with only the
publisher's shifted bit-plane transform inverted (compressed-tensors stores the unsigned
code + 2^(bits-1) form; the artifact stores two's-complement fields). No code was requantized.
The only value deviation from the source is 218 of 240,271,360 multipliers: bf16 subnormal
words whose FP16 rounding error is at most 2**-25 (max observed 2.98e-8), audited per group in
the conversion report. The publisher's task evaluations therefore describe the represented Text
and vocabulary weights.
The MTP layer (official BF16 checkpoint), the Vision tower (BF16 in the GSQ release), and the
DFlash2 companion (BF16 from z-lab) are quantized by NInfer's converter, so the verbatim claim
does not extend to them.
rtx4080-port,
built from source (sm_89; the fork's CMake accepts only CMAKE_CUDA_ARCHITECTURES=89), or
the published container image below. Upstream NInfer and the RTX 3090/4090 forks do not
register the gsq3 weights profile or the Q3G128_F16S scheme and reject this file.sm_89);rk4v4-e8 KV cache, which is what makes the 100K profile fit.hf download roofkid/Qwen3.8-27B-GSQ3-NInfer qwen3_8_27b_gsq3.ninfer --local-dir models
The turnkey profile is the rtx4080-port fork's scripts/run-ninfer-4080.sh (or .bat), which
serves 100,000 tokens with vision and MTP3. The equivalent container invocation is:
docker run --rm --gpus all -p 8080:8080 \
-v "$PWD/models/qwen3_8_27b_gsq3.ninfer:/models/model.ninfer:ro" \
roofkid/ninfer-4080:gsq3 \
ninfer-serve /models/model.ninfer \
--host 0.0.0.0 --port 8080 \
--max-context 102400 --kv-capacity 102400 --kv-dtype rk4v4-e8 \
--max-concurrency 1 --max-pending-requests 16 --prefill-chunk 2688 \
--host-kv-mib 4096 \
--spec mtp --draft-tokens 3 --lm-head-draft \
--vision --preserve-thinking
Without Docker, the CLI is:
./build/apps/ninfer models/qwen3_8_27b_gsq3.ninfer \
--prompt "Explain prefill and decode in three sentences." \
--max-context 32768 --max-new 8192 --kv-dtype rk4v4-e8 \
--spec mtp --draft-tokens 3 --lm-head-draft
Measured with ninfer_bench on the session-26 build (rtx4080-port tip), rk4v4-e8,
--prefill-chunk 1024, the fork's 131,072-token tiled corpus, one warmup and one measured
repetition per point. Acceptance is a tiled-corpus fixture property (repeated text), not a
model result.
| Depth | Prefill t/s | MTP3 decode t/s | DFlash2 K=7 decode t/s |
|---|---|---|---|
| 8K | 2,719.9 | 151.2 | 166.7 |
| 32K | 2,424.9 | 141.7 | 262.3 |
| 64K | 2,125.5 | 130.7 | 239.1 |
| 98K | 1,895.1 | 122.3 | 212.7 |
At the documented 100K prefill profile (--prefill-chunk 2688) the same build measures
1,971.4 tok/s, and 2,470.0 tok/s at 32,768 tokens. MTP3 and DFlash2 accept 3 and 7 draft
tokens respectively; the DFlash2 K=7 profile is greedy-lossless against the engine's own routes
in the fork's real-artifact tests, and the MTP3 route is covered by the same suite.
--max-context 102400, --host-kv-mib 4096, rk4v4-e8): about
11 GiB of device weights, KV plus runtime reservation validated before the server starts
listening, with roughly 0.9 GiB of headroom after startup.--host-kv-mib 4096, default host state slots) hold a deep 100K
checkpoint so rewrites reuse the prefix instead of re-prefilling.Accuracy profiles that need more KV bytes (int8, rk8v4) fit at shorter contexts; rk4v4-e8
is the accuracy profile validated at 100K.
rk4v4-e8, against 90β92% for the beellama kvarn5/5/kvarn4/4 reference (within noise).| Field | Value |
|---|---|
| Quantized source | ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ |
| Quantized source revision | b5ce0b76f60020a875dee4f6ec9d934cca4121e4 |
| Base model | Qwen/Qwen3.8-27B |
| Base revision (vocabulary, MTP) | 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 |
| DFlash2 source | z-lab/Qwen3.8-27B-DFlash2 |
| DFlash2 revision | 50307d4c4cde6860d4eee73e2547cd786fe8e8a4 |
| Conversion recipe | qwen3_8_27b_gsq3-v1 |
| Converter repository | https://github.com/roofkid/ninfer-4080 |
| Converter revision | ce85711b5f71296c4027f839284e41acd9a80669 |
| Minimum runtime revision | db4a7d9ee4c03adb26cf1958463b41c5a59db39b |
| Ranking input SHA-256 | c692dc76388132c910547589b4fb4a0503fbd6ad50aaac6a509bbcb192a8afa5 |
The conversion verifier rebuilt every packed plane from the source shards and compared them
word-for-word with the artifact: 323 packed objects, 4,527,104 rows, 240,271,360 groups,
base bytes equal 323/323, scales equal 323/323 (218 rounded subnormals). The artifact
identity, object inventory, and conversion provenance are published in
artifact-manifest.json.
This NInfer artifact is distributed under the Apache License 2.0. The Qwen3.8-27B base model and the 3-bit GSQ checkpoint are also licensed under Apache-2.0, and the DFlash2 companion comes from z-lab/Qwen3.8-27B-DFlash2 (Apache-2.0). Users remain responsible for complying with the licenses and applicable laws.
If this artifact is useful to you, you can support the maintainer at Buy Me a Coffee.
This model card is the version-controlled source for roofkid/Qwen3.8-27B-GSQ3-NInfer.
The repository contains a verbatim repack of the 3-bit
ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
checkpoint into the native NInfer .ninfer artifact
format, sized to run the full 100,000-token, vision, and speculative-decoding profile on a single
16 GB RTX 4080. It is intended only for the RTX 4080 NInfer fork; it is not a Transformers
checkpoint, Safetensors distribution, or GGUF file.
| Field | Value |
|---|---|
| Filename | qwen3_8_27b_gsq3.ninfer |
| Size | 13,330,776,576 bytes (12.41 GiB) |
| SHA-256 | c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7 |
| Container version | 2 |
| NInfer model ID | qwen3.8-27b |
| NInfer weights ID | gsq3 |
| NInfer target key | qwen3_8_27b |
| Stored objects | 1,190 (1,184 tensors and 6 resources) |
The Text body is 320 matrices in the registered Q3G128_F16S scheme: symmetric 3-bit codes in
[-4, 3], one FP16 multiplier per 128-value group, 3.125 bits per weight. The token embedding,
full output head, and optimized draft head use Q4G64_F16S; the MTP layer and Vision tower follow
NInfer's registered groupwise-int recipes; the DFlash2 companion was requantized to
Q4G64_F16S. The file also carries the
tokenizer, chat-template, generation, and media-processor objects required by NInfer.
Verify a downloaded file with:
printf '%s %s\n' \
'c6f27073393e5bcc629489420470d71f52a27553bfc5c360fef07a25b3b550d7' \
'qwen3_8_27b_gsq3.ninfer' | sha256sum --check
The 323 packed matrices that come from the publisher's checkpoint β the 320 Text-body matrices,
the token embedding, the full output head, and the draft head (a row gather of the output head) β
are a lossless repack: every code and every group scale was copied unchanged, with only the
publisher's shifted bit-plane transform inverted (compressed-tensors stores the unsigned
code + 2^(bits-1) form; the artifact stores two's-complement fields). No code was requantized.
The only value deviation from the source is 218 of 240,271,360 multipliers: bf16 subnormal
words whose FP16 rounding error is at most 2**-25 (max observed 2.98e-8), audited per group in
the conversion report. The publisher's task evaluations therefore describe the represented Text
and vocabulary weights.
The MTP layer (official BF16 checkpoint), the Vision tower (BF16 in the GSQ release), and the
DFlash2 companion (BF16 from z-lab) are quantized by NInfer's converter, so the verbatim claim
does not extend to them.
rtx4080-port,
built from source (sm_89; the fork's CMake accepts only CMAKE_CUDA_ARCHITECTURES=89), or
the published container image below. Upstream NInfer and the RTX 3090/4090 forks do not
register the gsq3 weights profile or the Q3G128_F16S scheme and reject this file.sm_89);rk4v4-e8 KV cache, which is what makes the 100K profile fit.hf download roofkid/Qwen3.8-27B-GSQ3-NInfer qwen3_8_27b_gsq3.ninfer --local-dir models
The turnkey profile is the rtx4080-port fork's scripts/run-ninfer-4080.sh (or .bat), which
serves 100,000 tokens with vision and MTP3. The equivalent container invocation is:
docker run --rm --gpus all -p 8080:8080 \
-v "$PWD/models/qwen3_8_27b_gsq3.ninfer:/models/model.ninfer:ro" \
roofkid/ninfer-4080:gsq3 \
ninfer-serve /models/model.ninfer \
--host 0.0.0.0 --port 8080 \
--max-context 102400 --kv-capacity 102400 --kv-dtype rk4v4-e8 \
--max-concurrency 1 --max-pending-requests 16 --prefill-chunk 2688 \
--host-kv-mib 4096 \
--spec mtp --draft-tokens 3 --lm-head-draft \
--vision --preserve-thinking
Without Docker, the CLI is:
./build/apps/ninfer models/qwen3_8_27b_gsq3.ninfer \
--prompt "Explain prefill and decode in three sentences." \
--max-context 32768 --max-new 8192 --kv-dtype rk4v4-e8 \
--spec mtp --draft-tokens 3 --lm-head-draft
Measured with ninfer_bench on the session-26 build (rtx4080-port tip), rk4v4-e8,
--prefill-chunk 1024, the fork's 131,072-token tiled corpus, one warmup and one measured
repetition per point. Acceptance is a tiled-corpus fixture property (repeated text), not a
model result.
| Depth | Prefill t/s | MTP3 decode t/s | DFlash2 K=7 decode t/s |
|---|---|---|---|
| 8K | 2,719.9 | 151.2 | 166.7 |
| 32K | 2,424.9 | 141.7 | 262.3 |
| 64K | 2,125.5 | 130.7 | 239.1 |
| 98K | 1,895.1 | 122.3 | 212.7 |
At the documented 100K prefill profile (--prefill-chunk 2688) the same build measures
1,971.4 tok/s, and 2,470.0 tok/s at 32,768 tokens. MTP3 and DFlash2 accept 3 and 7 draft
tokens respectively; the DFlash2 K=7 profile is greedy-lossless against the engine's own routes
in the fork's real-artifact tests, and the MTP3 route is covered by the same suite.
--max-context 102400, --host-kv-mib 4096, rk4v4-e8): about
11 GiB of device weights, KV plus runtime reservation validated before the server starts
listening, with roughly 0.9 GiB of headroom after startup.--host-kv-mib 4096, default host state slots) hold a deep 100K
checkpoint so rewrites reuse the prefix instead of re-prefilling.Accuracy profiles that need more KV bytes (int8, rk8v4) fit at shorter contexts; rk4v4-e8
is the accuracy profile validated at 100K.
rk4v4-e8, against 90β92% for the beellama kvarn5/5/kvarn4/4 reference (within noise).| Field | Value |
|---|---|
| Quantized source | ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ |
| Quantized source revision | b5ce0b76f60020a875dee4f6ec9d934cca4121e4 |
| Base model | Qwen/Qwen3.8-27B |
| Base revision (vocabulary, MTP) | 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 |
| DFlash2 source | z-lab/Qwen3.8-27B-DFlash2 |
| DFlash2 revision | 50307d4c4cde6860d4eee73e2547cd786fe8e8a4 |
| Conversion recipe | qwen3_8_27b_gsq3-v1 |
| Converter repository | https://github.com/roofkid/ninfer-4080 |
| Converter revision | ce85711b5f71296c4027f839284e41acd9a80669 |
| Minimum runtime revision | db4a7d9ee4c03adb26cf1958463b41c5a59db39b |
| Ranking input SHA-256 | c692dc76388132c910547589b4fb4a0503fbd6ad50aaac6a509bbcb192a8afa5 |
The conversion verifier rebuilt every packed plane from the source shards and compared them
word-for-word with the artifact: 323 packed objects, 4,527,104 rows, 240,271,360 groups,
base bytes equal 323/323, scales equal 323/323 (218 rounded subnormals). The artifact
identity, object inventory, and conversion provenance are published in
artifact-manifest.json.
This NInfer artifact is distributed under the Apache License 2.0. The Qwen3.8-27B base model and the 3-bit GSQ checkpoint are also licensed under Apache-2.0, and the DFlash2 companion comes from z-lab/Qwen3.8-27B-DFlash2 (Apache-2.0). Users remain responsible for complying with the licenses and applicable laws.
If this artifact is useful to you, you can support the maintainer at Buy Me a Coffee.