jabbatheduck/ninfer-ext-models

Model

NInfer models — Qwen3.8-27B EXL3

0

15 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

NInfer models — Qwen3.8-27B EXL3

Quantized artifacts for NInfer Ext, a from-scratch C++/CUDA inference engine for maximum single-GPU performance.

[!IMPORTANT] These artifacts only work with giveen/ninfer-ext. They are not Transformers checkpoints, not GGUF, not safetensors weights, and cannot be loaded by transformers, vLLM, llama.cpp, or exllamav3. .ninfer is NInfer's own artifact format, and the EXL3 trellis layout is decoded by kernels that live in that repository. Loading them anywhere else will fail.

Contents

FileSizeNotes
qwen3_8_27b_exl3_4bpw.ninfer15.68 GiBText + MTP + Vision, 4.0 bpw body
qwen3_8_27b_exl3_4bpw.ninfer.conversion.json489 KiBConversion provenance and per-tensor rates
qwen3_8_27b_exl3_3p5bpw.ninfer14.24 GiBText + MTP + Vision, 3.5 bpw body
qwen3_8_27b_exl3_3p5bpw.ninfer.conversion.json489 KiBConversion provenance and per-tensor rates

Both are the same model at two points on the size curve. 4.0 bpw is the one to lead with: it is the only artifact here that beats both of the engine's other Qwen3.8-27B builds on both perplexity and KL divergence, while being smaller than either (see Quality). 3.5 bpw is the size-optimised tier — best PPL per byte, but it does not carry that advantage into divergence.

More NInfer artifacts will be added to this repository over time.

Model

Qwen3.8-27B (Qwen3_5ForCausalLM): 64 layers (48 GDN linear-attention, 16 full attention), hidden 5120, intermediate 17408, vocab 248,320, plus a separate MTP layer and a Vision tower.

Quantization

NInfer's native EXL3 format (exl3_mul1 + trellis_t16_v1) — a three-instruction trellis codebook over 16x16 tiles, with the input and output Hadamard rotations folded into the kernels. Weights are produced by NInfer's own C++/CUDA quantizer (ninfer-quantize) from full-precision source tensors; no exllamav3 checkpoint is imported.

Per-tensor rates (half bits, i.e. X.5 bpw, are first-class trellis rates):

Scope4.0 bpw artifact3.5 bpw artifact
MLP and GDN projections4.0 bpw3.5 bpw
Attention projections (the -hq promotion)5.0 bpw4.5 bpw
Vocabulary head6.0 bpw6.0 bpw
MTP layer4.0 bpw (5.0 attention), calibrated3.5 bpw (4.5 attention), calibrated
Vision towergroupwise Q6/Q8/Q4/Q5 (not EXL3 — its MLP intermediate is not 128-aligned)same

The MTP layer is calibrated from the final hidden states and next-token embeddings.

Quality

Full-corpus perplexity over 261,167 tokens (context/stride 4096/2048, FP8 KV, greedy), and KL divergence against the full-precision model over the same 2,940 positions of a 23.5k-token wikitext slice:

ArtifactSizePPLKL(P_BF16 ‖ P)
EXL3 4.0 bpw15.68 GiB4.29390.0332
Groupwise INT416.96 GiB4.34390.0429
NVFP422.09 GiB4.31490.0510
EXL3 3.5 bpw14.24 GiB4.31010.0624

The two metrics rank these differently, and both are reported for that reason: 4.0 bpw wins on both, but 3.5 bpw's better perplexity than INT4 and NVFP4 does not survive as divergence. Only the KL ordering should be compared across runs — the absolute values move by roughly 2× with the text, and the ordering was reproduced in both halves of the reference.

Performance

Measured on one NVIDIA GeForce RTX 5090, CUDA 13.3, a single request, greedy, 64-256 output tokens, --prefill-chunk 1024:

RegimeEXL3 4.0 bpwEXL3 3.5 bpwQ4NVFP4
Decode, plain75 tok/s64 tok/s83 tok/s73 tok/s
Decode, MTP K=3137 tok/s130 tok/s134 tok/s142 tok/s
Decode, MTP K=5 + --lm-head-draft145 tok/s131 tok/s144 tok/s167 tok/s
Prefill, 0.54k / 7.6k-token prompt1.97k / 2.33k tok/s1.72k / 2.11k tok/s2.42k / 2.91k tok/s5.28k / 8.60k tok/s

Measuring MTP at K=3 for all four keeps the draft length equal across formats; the K=5 row is each artifact's own best setting, which Q4 and NVFP4 reach with --lm-head-draft. Both EXL3 tiers carry the indexed proposal head that flag needs, so it is available to them too.

4.0 bpw is close to the other native formats on every regime: prefill 1.19–1.23x behind Q4, plain decode within 10%, and at a comparable draft length its MTP matches Q4's exactly, on an artifact 1.3 GiB smaller (15.68 against 16.96 GiB). NVFP4 remains the prefill leader — as it is for this engine's other models — because its tensor-core contraction needs no per-weight decoding, which a 4-bit trellis does: the contraction issues exactly the same number of MMAs as Q4's, and the difference is the funnel, bit-field extracts and IMAD/DP4A per decoded window that the trellis costs.

3.5 bpw is slower than 4.0 bpw on every regime, on the same kernels. Its weights are 9% smaller, but its odd half-rates take the heavier exl3_windows_half window decode — two funnel shifts for the eight windows against one funnel and five bit-field extracts — and that costs more than the bytes it saves. Its case is size and perplexity per byte, not speed.

These are single-request spot measurements, not the engine's methodology-conforming performance tables; docs/performance.md records the published coverage and the difference.

Usage

Build the engine from source, then serve an artifact:

git clone https://github.com/giveen/ninfer-ext
cd ninfer-ext && cmake -B build -DCMAKE_BUILD_TYPE=Release \
  -DPython3_EXECUTABLE=$PWD/.venv/bin/python && cmake --build build -j

build/apps/ninfer-serve qwen3_8_27b_exl3_4bpw.ninfer \
  --port 8099 --spec mtp --draft-tokens 3 --fixed-draft

--vision enables image input; the Vision tower loads lazily. See the repository's README.md, docs/cli.md, and docs/serving.md for the full command surface.

License

The quantized weights follow the base model's license. The quantization format, kernels, and tooling are part of giveen/ninfer-ext.

code
cuda
exl3
ninfer
ninfer-ext
quantization
speculative-decoding
text-generation
trellis
vision

jabbatheduck/ninfer-ext-models

Model

NInfer models — Qwen3.8-27B EXL3

0

15 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

NInfer models — Qwen3.8-27B EXL3

Quantized artifacts for NInfer Ext, a from-scratch C++/CUDA inference engine for maximum single-GPU performance.

[!IMPORTANT] These artifacts only work with giveen/ninfer-ext. They are not Transformers checkpoints, not GGUF, not safetensors weights, and cannot be loaded by transformers, vLLM, llama.cpp, or exllamav3. .ninfer is NInfer's own artifact format, and the EXL3 trellis layout is decoded by kernels that live in that repository. Loading them anywhere else will fail.

Contents

FileSizeNotes
qwen3_8_27b_exl3_4bpw.ninfer15.68 GiBText + MTP + Vision, 4.0 bpw body
qwen3_8_27b_exl3_4bpw.ninfer.conversion.json489 KiBConversion provenance and per-tensor rates
qwen3_8_27b_exl3_3p5bpw.ninfer14.24 GiBText + MTP + Vision, 3.5 bpw body
qwen3_8_27b_exl3_3p5bpw.ninfer.conversion.json489 KiBConversion provenance and per-tensor rates

Both are the same model at two points on the size curve. 4.0 bpw is the one to lead with: it is the only artifact here that beats both of the engine's other Qwen3.8-27B builds on both perplexity and KL divergence, while being smaller than either (see Quality). 3.5 bpw is the size-optimised tier — best PPL per byte, but it does not carry that advantage into divergence.

More NInfer artifacts will be added to this repository over time.

Model

Qwen3.8-27B (Qwen3_5ForCausalLM): 64 layers (48 GDN linear-attention, 16 full attention), hidden 5120, intermediate 17408, vocab 248,320, plus a separate MTP layer and a Vision tower.

Quantization

NInfer's native EXL3 format (exl3_mul1 + trellis_t16_v1) — a three-instruction trellis codebook over 16x16 tiles, with the input and output Hadamard rotations folded into the kernels. Weights are produced by NInfer's own C++/CUDA quantizer (ninfer-quantize) from full-precision source tensors; no exllamav3 checkpoint is imported.

Per-tensor rates (half bits, i.e. X.5 bpw, are first-class trellis rates):

Scope4.0 bpw artifact3.5 bpw artifact
MLP and GDN projections4.0 bpw3.5 bpw
Attention projections (the -hq promotion)5.0 bpw4.5 bpw
Vocabulary head6.0 bpw6.0 bpw
MTP layer4.0 bpw (5.0 attention), calibrated3.5 bpw (4.5 attention), calibrated
Vision towergroupwise Q6/Q8/Q4/Q5 (not EXL3 — its MLP intermediate is not 128-aligned)same

The MTP layer is calibrated from the final hidden states and next-token embeddings.

Quality

Full-corpus perplexity over 261,167 tokens (context/stride 4096/2048, FP8 KV, greedy), and KL divergence against the full-precision model over the same 2,940 positions of a 23.5k-token wikitext slice:

ArtifactSizePPLKL(P_BF16 ‖ P)
EXL3 4.0 bpw15.68 GiB4.29390.0332
Groupwise INT416.96 GiB4.34390.0429
NVFP422.09 GiB4.31490.0510
EXL3 3.5 bpw14.24 GiB4.31010.0624

The two metrics rank these differently, and both are reported for that reason: 4.0 bpw wins on both, but 3.5 bpw's better perplexity than INT4 and NVFP4 does not survive as divergence. Only the KL ordering should be compared across runs — the absolute values move by roughly 2× with the text, and the ordering was reproduced in both halves of the reference.

Performance

Measured on one NVIDIA GeForce RTX 5090, CUDA 13.3, a single request, greedy, 64-256 output tokens, --prefill-chunk 1024:

RegimeEXL3 4.0 bpwEXL3 3.5 bpwQ4NVFP4
Decode, plain75 tok/s64 tok/s83 tok/s73 tok/s
Decode, MTP K=3137 tok/s130 tok/s134 tok/s142 tok/s
Decode, MTP K=5 + --lm-head-draft145 tok/s131 tok/s144 tok/s167 tok/s
Prefill, 0.54k / 7.6k-token prompt1.97k / 2.33k tok/s1.72k / 2.11k tok/s2.42k / 2.91k tok/s5.28k / 8.60k tok/s

Measuring MTP at K=3 for all four keeps the draft length equal across formats; the K=5 row is each artifact's own best setting, which Q4 and NVFP4 reach with --lm-head-draft. Both EXL3 tiers carry the indexed proposal head that flag needs, so it is available to them too.

4.0 bpw is close to the other native formats on every regime: prefill 1.19–1.23x behind Q4, plain decode within 10%, and at a comparable draft length its MTP matches Q4's exactly, on an artifact 1.3 GiB smaller (15.68 against 16.96 GiB). NVFP4 remains the prefill leader — as it is for this engine's other models — because its tensor-core contraction needs no per-weight decoding, which a 4-bit trellis does: the contraction issues exactly the same number of MMAs as Q4's, and the difference is the funnel, bit-field extracts and IMAD/DP4A per decoded window that the trellis costs.

3.5 bpw is slower than 4.0 bpw on every regime, on the same kernels. Its weights are 9% smaller, but its odd half-rates take the heavier exl3_windows_half window decode — two funnel shifts for the eight windows against one funnel and five bit-field extracts — and that costs more than the bytes it saves. Its case is size and perplexity per byte, not speed.

These are single-request spot measurements, not the engine's methodology-conforming performance tables; docs/performance.md records the published coverage and the difference.

Usage

Build the engine from source, then serve an artifact:

git clone https://github.com/giveen/ninfer-ext
cd ninfer-ext && cmake -B build -DCMAKE_BUILD_TYPE=Release \
  -DPython3_EXECUTABLE=$PWD/.venv/bin/python && cmake --build build -j

build/apps/ninfer-serve qwen3_8_27b_exl3_4bpw.ninfer \
  --port 8099 --spec mtp --draft-tokens 3 --fixed-draft

--vision enables image input; the Vision tower loads lazily. See the repository's README.md, docs/cli.md, and docs/serving.md for the full command surface.

License

The quantized weights follow the base model's license. The quantization format, kernels, and tooling are part of giveen/ninfer-ext.

code
cuda
exl3
ninfer
ninfer-ext
quantization
speculative-decoding
text-generation
trellis
vision