antirez/qwen3.8-flash-next-gguf

Model

Qwen3.8 Flash Next for DwarfStar

12

4 commits

2 linked in READMEs

updated Sep 14, 2026

See the code

README

Qwen3.8 Flash Next for DwarfStar

Self-contained GGUFs for DwarfStar, with original BF16 n-grams read directly from disk. Keep the file on a local SSD. No n-gram sidecar is needed. MTP weights are included; enable them with --mtp. Requires DwarfStar's native BF16 n-gram reader. Older versions expecting --ple cannot load these files.

FileFile sizeMain/MTP weights
Qwen3.8-Flash-Next-Q2.gguf137.10 GiB41.73 GiB
Qwen3.8-Flash-Next-Q4.gguf165.11 GiB69.74 GiB

Each file contains the same 95.37 GiB BF16 n-gram table. It is not made resident or quantized. Context and runtime buffers require additional RAM.

The main/MTP tensor bytes come unchanged from Ivan Fioravanti's Q2 and Q4 releases. Q2 has imatrix-calibrated IQ2_XXS gate/up and Q2_K down routed experts, with the down rows padded from 640 to 768 inputs. Q4 has calibrated Q4_K gate/up and MXFP4 down. Other tensor formats and MTP weights are unchanged.

The n-grams are copied byte-for-byte from Qwen's original checkpoint at de4b8e4d43b917e7706784d8bb445c9af86a3540, including its padding rows. The packer checks the original hash constants and verifies every copied payload. These are quantized language-model weights, not lossless copies of the whole original model.

Original model: Qwen. Quantized main/MTP releases: Ivan Fioravanti. Native n-gram packaging and disk-only integration: Salvatore Sanfilippo / DwarfStar. The original Qwen Community License is included as LICENSE.

conversational
endpoints_compatible
gguf
qwen4exp
text-generation

antirez/qwen3.8-flash-next-gguf

Model

Qwen3.8 Flash Next for DwarfStar

12

4 commits

2 linked in READMEs

updated Sep 14, 2026

See the code

README

Qwen3.8 Flash Next for DwarfStar

Self-contained GGUFs for DwarfStar, with original BF16 n-grams read directly from disk. Keep the file on a local SSD. No n-gram sidecar is needed. MTP weights are included; enable them with --mtp. Requires DwarfStar's native BF16 n-gram reader. Older versions expecting --ple cannot load these files.

FileFile sizeMain/MTP weights
Qwen3.8-Flash-Next-Q2.gguf137.10 GiB41.73 GiB
Qwen3.8-Flash-Next-Q4.gguf165.11 GiB69.74 GiB

Each file contains the same 95.37 GiB BF16 n-gram table. It is not made resident or quantized. Context and runtime buffers require additional RAM.

The main/MTP tensor bytes come unchanged from Ivan Fioravanti's Q2 and Q4 releases. Q2 has imatrix-calibrated IQ2_XXS gate/up and Q2_K down routed experts, with the down rows padded from 640 to 768 inputs. Q4 has calibrated Q4_K gate/up and MXFP4 down. Other tensor formats and MTP weights are unchanged.

The n-grams are copied byte-for-byte from Qwen's original checkpoint at de4b8e4d43b917e7706784d8bb445c9af86a3540, including its padding rows. The packer checks the original hash constants and verifies every copied payload. These are quantized language-model weights, not lossless copies of the whole original model.

Original model: Qwen. Quantized main/MTP releases: Ivan Fioravanti. Native n-gram packaging and disk-only integration: Salvatore Sanfilippo / DwarfStar. The original Qwen Community License is included as LICENSE.

conversational
endpoints_compatible
gguf
qwen4exp
text-generation