julianmb/DeepSeek-V4-Flash-0731-IQ2XXS-STRIX

Model

DeepSeek-V4-Flash-0731-IQ2XXS-STRIX

1

3 commits

2 linked in READMEs

updated Aug 4, 2026

See the code

README

DeepSeek-V4-Flash-0731-IQ2XXS-STRIX

Quantized DeepSeek V4 Flash (0731) GGUF for AMD Strix Halo (gfx1151).

Details

  • Base model: DeepSeek V4 Flash 0731 (284B MoE)
  • Quantization source: tekosML (DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix)
  • File size: 86.72 GB
  • Architecture: 43 routed layers, 256 experts/layer (6 active), 4-stream hyper-connections

Quant Recipe

Tensor groupType
Attention projectionsQ8_0
Shared expertsQ8_0
Output headQ8_0
Token embeddingF16
Routed gate/up expertsIQ2_XXS
Routed down expertsQ2_K

Run with the julianmb/ds4fa engine on 128 GB Strix Halo:

DS4_ROCM_STREAM_MODEL_CACHE_GB=48 ./ds4 -m DeepSeek-V4-Flash-0731-IQ2XXS-STRIX.gguf -c 512 \
  --ssd-streaming --ssd-streaming-cache-experts 32GB \
  -p "What is the capital of France?" --think --tokens 60
conversational
deepseek
endpoints_compatible
gguf
imatrix
iq2_xxs
q2_k
q8_0
strix-halo
text-generation

julianmb/DeepSeek-V4-Flash-0731-IQ2XXS-STRIX

Model

DeepSeek-V4-Flash-0731-IQ2XXS-STRIX

1

3 commits

2 linked in READMEs

updated Aug 4, 2026

See the code

README

DeepSeek-V4-Flash-0731-IQ2XXS-STRIX

Quantized DeepSeek V4 Flash (0731) GGUF for AMD Strix Halo (gfx1151).

Details

  • Base model: DeepSeek V4 Flash 0731 (284B MoE)
  • Quantization source: tekosML (DeepSeek-V4-Flash-0731-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-imatrix)
  • File size: 86.72 GB
  • Architecture: 43 routed layers, 256 experts/layer (6 active), 4-stream hyper-connections

Quant Recipe

Tensor groupType
Attention projectionsQ8_0
Shared expertsQ8_0
Output headQ8_0
Token embeddingF16
Routed gate/up expertsIQ2_XXS
Routed down expertsQ2_K

Run with the julianmb/ds4fa engine on 128 GB Strix Halo:

DS4_ROCM_STREAM_MODEL_CACHE_GB=48 ./ds4 -m DeepSeek-V4-Flash-0731-IQ2XXS-STRIX.gguf -c 512 \
  --ssd-streaming --ssd-streaming-cache-experts 32GB \
  -p "What is the capital of France?" --think --tokens 60
conversational
deepseek
endpoints_compatible
gguf
imatrix
iq2_xxs
q2_k
q8_0
strix-halo
text-generation