A GGUF of Qwen/Qwen3.6-35B-A3B with higher-precision dense weights and the model's native MTP (multi-token prediction) block, for fast speculative decoding.
| Tensors | Type |
|---|---|
| Trunk dense weights (attention, DeltaNet, shared expert, output head) | Q6_K |
| Routed experts | as Unsloth UD-Q4_K_XL (gate/up Q4_K, down Q5_K; a few layers Q5_K/Q6_K) |
| Token embeddings, MTP block dense weights | Q8_0 |
| Norms, router, SSM parameters | F32 (MTP router BF16) |
File: Qwen3.6-35B-A3B-Q6dense.gguf, 22,388,168,960 bytes, SHA-256
842ed2bf58034c2f4856239d04de57acec27183ef4012e469d1a423a8de23108.
On AMD Strix Halo (gfx1151) with gufo, the Q6_K dense weights gave the same task quality as UD-Q4_K_XL and were faster with MTP than other dense quantizations we tried. Greedy next-token agreement with llama.cpp on this file: 443/445 positions, mean KL 0.002.
# gufo (native MTP from the same file)
gufo serve llm --model Qwen3.6-35B-A3B-Q6dense.gguf --speculative mtp
# llama.cpp (b11069 or later)
llama-server -m Qwen3.6-35B-A3B-Q6dense.gguf -ngl 999 -fa on --spec-type draft-mtp
Source: the BF16 GGUF and importance matrix of
unsloth/Qwen3.6-35B-A3B-MTP-GGUF
at revision 5bc3e238d916f48a861bac2f8a1990a0e9b7e98d. Quantized with llama.cpp
b11069 (commit 68d9053a). q6dense.types in this repo sets the type of
every tensor:
llama-quantize --imatrix imatrix_unsloth.gguf_file \
--tensor-type-file q6dense.types \
Qwen3.6-35B-A3B-BF16-00001-of-00002.gguf Qwen3.6-35B-A3B-Q6dense.gguf \
Q4_K_M 16
Apache-2.0, as the original model. This is a modified (requantized) version of Qwen3.6-35B-A3B by the Qwen team, using Unsloth's BF16 conversion and importance matrix. All credit for the model belongs to them.
A GGUF of Qwen/Qwen3.6-35B-A3B with higher-precision dense weights and the model's native MTP (multi-token prediction) block, for fast speculative decoding.
| Tensors | Type |
|---|---|
| Trunk dense weights (attention, DeltaNet, shared expert, output head) | Q6_K |
| Routed experts | as Unsloth UD-Q4_K_XL (gate/up Q4_K, down Q5_K; a few layers Q5_K/Q6_K) |
| Token embeddings, MTP block dense weights | Q8_0 |
| Norms, router, SSM parameters | F32 (MTP router BF16) |
File: Qwen3.6-35B-A3B-Q6dense.gguf, 22,388,168,960 bytes, SHA-256
842ed2bf58034c2f4856239d04de57acec27183ef4012e469d1a423a8de23108.
On AMD Strix Halo (gfx1151) with gufo, the Q6_K dense weights gave the same task quality as UD-Q4_K_XL and were faster with MTP than other dense quantizations we tried. Greedy next-token agreement with llama.cpp on this file: 443/445 positions, mean KL 0.002.
# gufo (native MTP from the same file)
gufo serve llm --model Qwen3.6-35B-A3B-Q6dense.gguf --speculative mtp
# llama.cpp (b11069 or later)
llama-server -m Qwen3.6-35B-A3B-Q6dense.gguf -ngl 999 -fa on --spec-type draft-mtp
Source: the BF16 GGUF and importance matrix of
unsloth/Qwen3.6-35B-A3B-MTP-GGUF
at revision 5bc3e238d916f48a861bac2f8a1990a0e9b7e98d. Quantized with llama.cpp
b11069 (commit 68d9053a). q6dense.types in this repo sets the type of
every tensor:
llama-quantize --imatrix imatrix_unsloth.gguf_file \
--tensor-type-file q6dense.types \
Qwen3.6-35B-A3B-BF16-00001-of-00002.gguf Qwen3.6-35B-A3B-Q6dense.gguf \
Q4_K_M 16
Apache-2.0, as the original model. This is a modified (requantized) version of Qwen3.6-35B-A3B by the Qwen team, using Unsloth's BF16 conversion and importance matrix. All credit for the model belongs to them.