3
stars
9
commits
1
linked in READMEs
Jun 23, 2026
updated
INT8 (W8A8) quantization of Qwen/Qwen3.6-35B-A3B — a hybrid Mixture-of-Experts model (256 experts, top-8, ~3B active) with GatedDeltaNet linear-attention, a vision tower and an MTP head.
self_attn, the router (mlp.gate), shared_expert, GatedDeltaNet (linear_attn), lm_head, embeddings, vision tower, MTP head.compressed-tensors (int-quantized). Full recipe: recipe.yaml.INT8 W8A8 is near-lossless; for the smallest footprint use the INT4-W4A16 variant (GSM8K 96.8% / MMLU-Pro 80.2%).
vllm serve Avesed/Qwen3.6-35B-A3B-INT8-W8A8 \
--tensor-parallel-size 2 --trust-remote-code --reasoning-parser qwen3
Served via vLLM's INT8 MoE path (works on Ampere sm_80 / sm_86).
Quantized with vllm-ampere-optimized/quantize.
9 commits
3
stars
9
commits
1
linked in READMEs
Jun 23, 2026
updated
INT8 (W8A8) quantization of Qwen/Qwen3.6-35B-A3B — a hybrid Mixture-of-Experts model (256 experts, top-8, ~3B active) with GatedDeltaNet linear-attention, a vision tower and an MTP head.
self_attn, the router (mlp.gate), shared_expert, GatedDeltaNet (linear_attn), lm_head, embeddings, vision tower, MTP head.compressed-tensors (int-quantized). Full recipe: recipe.yaml.INT8 W8A8 is near-lossless; for the smallest footprint use the INT4-W4A16 variant (GSM8K 96.8% / MMLU-Pro 80.2%).
vllm serve Avesed/Qwen3.6-35B-A3B-INT8-W8A8 \
--tensor-parallel-size 2 --trust-remote-code --reasoning-parser qwen3
Served via vLLM's INT8 MoE path (works on Ampere sm_80 / sm_86).
Quantized with vllm-ampere-optimized/quantize.
9 commits