AutomatosX/AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP

Model

0

stars

8

commits

1

linked in READMEs

Aug 21, 2026

updated

2-bit
apple-silicon
axq
axquant
conversational
deepseek
deepseek-v4
deepseek_v4
development
experimental
mixed-precision
mlx
mtp
quantized
safetensors
text-generation
Browse cluster: Model Compression for Edge AI

README

AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP

Development / experimental AXQuant 2-bit pack of deepseek-ai/DeepSeek-V4-Flash-0731 @ 7872f01b1d1fe23eabc4c98b48bffcef5a386062.

Converted on df-macstudio-m2 (Apple M2 Ultra, 192 GB) from the native FP8 0731 source (quant_method=fp8). Product class 2bit-experimental. Recipe: AXQuant manual deepseek-v4-experimental-2bit-v0.1.yaml (uniform 2-bit trunk). MTP sidecar is packaged (mtp.safetensors).

This is not the older DeepSeek-V4-Flash Hub pack. Do not treat AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP certificates as evidence for this 0731 revision.

Recipe (uniform v0.1)

This is the AXQuant assignment that scored best among 2-bit-class converts on the factory v-extract suite. Later mixed / attention-6 / shared-4-bit recipes scored worse and are not this pack.

TensorsBitsMethod
Routed experts + MLP2affine, group 32
Attention4affine, group 32
Embeddings / routers8affine, group 32
Norms, LM head, MTP16bf16

Measured precision

PropertyValue
Target class2bit-experimental
Measured main BPW3.1328993873020314
Measured total BPW3.2142055528774454
Weight bytes122,212,298,775
Sourcedeepseek-ai/DeepSeek-V4-Flash-0731@7872f01b1d1fe23eabc4c98b48bffcef5a386062
Convert hostdf-macstudio-m2
AXQuant1.9.0

Claims

ClaimStatus
Converted on Studio from the pinned 0731 revisionYes
Official DSV4 chat_template.jinjaIn pack
Checkpoint Tier 1 (generation viability suite)Not certified — 7.1.5 native 15+15 combined 0.633; v-extract on AX Engine HEAD 80f2a3e6 combined 0.887 (floor 0.90). Distinct 2-bit recipe converts scored worse.
AX Engine 7.1.5 native loadPassed on df-macstudio-m2 (Hub commit cb1a34b4, --stream-experts off, chat smoke Okay.)
Decode-128 (informational)15.535 tok/s on 7.1.5; not a Tier 1 claim
MTP assets (mtp.safetensors)Packaged — Hub name uses -MTP
MTP accelerationNot certified (T1 below 0.90; MTP A/B not run)

Requires AX_ENGINE_2BIT_EXPERIMENTAL=1 for AX Engine native serve. Certificate: deepseek-v4-flash-0731-axq2-tier1.md. Comparison vs OptiQ 2-bit: optiq2-vs-axq2-v190.

Runtime and MTP policy

Download the complete snapshot before serving it:

hf download AutomatosX/AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP \
  --local-dir ./AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP
AX_ENGINE_2BIT_EXPERIMENTAL=1 \
  ax-engine serve ./AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP --port 31418

mtp.safetensors is the native DeepSeek V4 nextn sidecar. It is not a Qwen qwen3-next-mtp sidecar, so the oMLX/MTPLX Qwen import workflow does not apply. AX Engine 7.1.5 recognizes the sidecar, but keeps this checkpoint on direct fallback because no revision-bound Tier 2 MTP acceptance, exactness, or speed evidence exists. The internal DeepSeek MTP certification-candidate switch is for the formal harness, not normal serving. Stock MLX-LM can run the text backbone but does not activate this sidecar.

Attribution

Base weights © DeepSeek. Quantization by AXQuant (development).

Contributors

AutomatosX

8 commits

AutomatosX/AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP

Model

0

stars

8

commits

1

linked in READMEs

Aug 21, 2026

updated

2-bit
apple-silicon
axq
axquant
conversational
deepseek
deepseek-v4
deepseek_v4
development
experimental
mixed-precision
mlx
mtp
quantized
safetensors
text-generation
Browse cluster: Model Compression for Edge AI

README

AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP

Development / experimental AXQuant 2-bit pack of deepseek-ai/DeepSeek-V4-Flash-0731 @ 7872f01b1d1fe23eabc4c98b48bffcef5a386062.

Converted on df-macstudio-m2 (Apple M2 Ultra, 192 GB) from the native FP8 0731 source (quant_method=fp8). Product class 2bit-experimental. Recipe: AXQuant manual deepseek-v4-experimental-2bit-v0.1.yaml (uniform 2-bit trunk). MTP sidecar is packaged (mtp.safetensors).

This is not the older DeepSeek-V4-Flash Hub pack. Do not treat AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP certificates as evidence for this 0731 revision.

Recipe (uniform v0.1)

This is the AXQuant assignment that scored best among 2-bit-class converts on the factory v-extract suite. Later mixed / attention-6 / shared-4-bit recipes scored worse and are not this pack.

TensorsBitsMethod
Routed experts + MLP2affine, group 32
Attention4affine, group 32
Embeddings / routers8affine, group 32
Norms, LM head, MTP16bf16

Measured precision

PropertyValue
Target class2bit-experimental
Measured main BPW3.1328993873020314
Measured total BPW3.2142055528774454
Weight bytes122,212,298,775
Sourcedeepseek-ai/DeepSeek-V4-Flash-0731@7872f01b1d1fe23eabc4c98b48bffcef5a386062
Convert hostdf-macstudio-m2
AXQuant1.9.0

Claims

ClaimStatus
Converted on Studio from the pinned 0731 revisionYes
Official DSV4 chat_template.jinjaIn pack
Checkpoint Tier 1 (generation viability suite)Not certified — 7.1.5 native 15+15 combined 0.633; v-extract on AX Engine HEAD 80f2a3e6 combined 0.887 (floor 0.90). Distinct 2-bit recipe converts scored worse.
AX Engine 7.1.5 native loadPassed on df-macstudio-m2 (Hub commit cb1a34b4, --stream-experts off, chat smoke Okay.)
Decode-128 (informational)15.535 tok/s on 7.1.5; not a Tier 1 claim
MTP assets (mtp.safetensors)Packaged — Hub name uses -MTP
MTP accelerationNot certified (T1 below 0.90; MTP A/B not run)

Requires AX_ENGINE_2BIT_EXPERIMENTAL=1 for AX Engine native serve. Certificate: deepseek-v4-flash-0731-axq2-tier1.md. Comparison vs OptiQ 2-bit: optiq2-vs-axq2-v190.

Runtime and MTP policy

Download the complete snapshot before serving it:

hf download AutomatosX/AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP \
  --local-dir ./AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP
AX_ENGINE_2BIT_EXPERIMENTAL=1 \
  ax-engine serve ./AX-DeepSeek-V4-Flash-0731-MLX-AXQ-2bit-MTP --port 31418

mtp.safetensors is the native DeepSeek V4 nextn sidecar. It is not a Qwen qwen3-next-mtp sidecar, so the oMLX/MTPLX Qwen import workflow does not apply. AX Engine 7.1.5 recognizes the sidecar, but keeps this checkpoint on direct fallback because no revision-bound Tier 2 MTP acceptance, exactness, or speed evidence exists. The internal DeepSeek MTP certification-candidate switch is for the formal harness, not normal serving. Stock MLX-LM can run the text backbone but does not activate this sidecar.

Attribution

Base weights © DeepSeek. Quantization by AXQuant (development).

Contributors

AutomatosX

8 commits