ligamentexceed/Qwen3.6-35B-A3B-Q6dense-GGUF

Model

Qwen3.6-35B-A3B Q6dense GGUF

0

2 commits

2 linked in READMEs

updated Oct 3, 2026

See the code

README

Qwen3.6-35B-A3B Q6dense GGUF

A GGUF of Qwen/Qwen3.6-35B-A3B with higher-precision dense weights and the model's native MTP (multi-token prediction) block, for fast speculative decoding.

TensorsType
Trunk dense weights (attention, DeltaNet, shared expert, output head)Q6_K
Routed expertsas Unsloth UD-Q4_K_XL (gate/up Q4_K, down Q5_K; a few layers Q5_K/Q6_K)
Token embeddings, MTP block dense weightsQ8_0
Norms, router, SSM parametersF32 (MTP router BF16)

File: Qwen3.6-35B-A3B-Q6dense.gguf, 22,388,168,960 bytes, SHA-256 842ed2bf58034c2f4856239d04de57acec27183ef4012e469d1a423a8de23108.

Why

On AMD Strix Halo (gfx1151) with gufo, the Q6_K dense weights gave the same task quality as UD-Q4_K_XL and were faster with MTP than other dense quantizations we tried. Greedy next-token agreement with llama.cpp on this file: 443/445 positions, mean KL 0.002.

Use

# gufo (native MTP from the same file)
gufo serve llm --model Qwen3.6-35B-A3B-Q6dense.gguf --speculative mtp

# llama.cpp (b11069 or later)
llama-server -m Qwen3.6-35B-A3B-Q6dense.gguf -ngl 999 -fa on --spec-type draft-mtp

Recipe

Source: the BF16 GGUF and importance matrix of unsloth/Qwen3.6-35B-A3B-MTP-GGUF at revision 5bc3e238d916f48a861bac2f8a1990a0e9b7e98d. Quantized with llama.cpp b11069 (commit 68d9053a). q6dense.types in this repo sets the type of every tensor:

llama-quantize --imatrix imatrix_unsloth.gguf_file \
  --tensor-type-file q6dense.types \
  Qwen3.6-35B-A3B-BF16-00001-of-00002.gguf Qwen3.6-35B-A3B-Q6dense.gguf \
  Q4_K_M 16

License and attribution

Apache-2.0, as the original model. This is a modified (requantized) version of Qwen3.6-35B-A3B by the Qwen team, using Unsloth's BF16 conversion and importance matrix. All credit for the model belongs to them.

conversational
endpoints_compatible
gguf
gufo
imatrix
qwen35moe
qwen3.6
text-generation

ligamentexceed/Qwen3.6-35B-A3B-Q6dense-GGUF

Model

Qwen3.6-35B-A3B Q6dense GGUF

0

2 commits

2 linked in READMEs

updated Oct 3, 2026

See the code

README

Qwen3.6-35B-A3B Q6dense GGUF

A GGUF of Qwen/Qwen3.6-35B-A3B with higher-precision dense weights and the model's native MTP (multi-token prediction) block, for fast speculative decoding.

TensorsType
Trunk dense weights (attention, DeltaNet, shared expert, output head)Q6_K
Routed expertsas Unsloth UD-Q4_K_XL (gate/up Q4_K, down Q5_K; a few layers Q5_K/Q6_K)
Token embeddings, MTP block dense weightsQ8_0
Norms, router, SSM parametersF32 (MTP router BF16)

File: Qwen3.6-35B-A3B-Q6dense.gguf, 22,388,168,960 bytes, SHA-256 842ed2bf58034c2f4856239d04de57acec27183ef4012e469d1a423a8de23108.

Why

On AMD Strix Halo (gfx1151) with gufo, the Q6_K dense weights gave the same task quality as UD-Q4_K_XL and were faster with MTP than other dense quantizations we tried. Greedy next-token agreement with llama.cpp on this file: 443/445 positions, mean KL 0.002.

Use

# gufo (native MTP from the same file)
gufo serve llm --model Qwen3.6-35B-A3B-Q6dense.gguf --speculative mtp

# llama.cpp (b11069 or later)
llama-server -m Qwen3.6-35B-A3B-Q6dense.gguf -ngl 999 -fa on --spec-type draft-mtp

Recipe

Source: the BF16 GGUF and importance matrix of unsloth/Qwen3.6-35B-A3B-MTP-GGUF at revision 5bc3e238d916f48a861bac2f8a1990a0e9b7e98d. Quantized with llama.cpp b11069 (commit 68d9053a). q6dense.types in this repo sets the type of every tensor:

llama-quantize --imatrix imatrix_unsloth.gguf_file \
  --tensor-type-file q6dense.types \
  Qwen3.6-35B-A3B-BF16-00001-of-00002.gguf Qwen3.6-35B-A3B-Q6dense.gguf \
  Q4_K_M 16

License and attribution

Apache-2.0, as the original model. This is a modified (requantized) version of Qwen3.6-35B-A3B by the Qwen team, using Unsloth's BF16 conversion and importance matrix. All credit for the model belongs to them.

conversational
endpoints_compatible
gguf
gufo
imatrix
qwen35moe
qwen3.6
text-generation