AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP

Model

0

stars

4

commits

1

linked in READMEs

Aug 21, 2026

updated

2-bit
2bit
apple-silicon
axq
axquant
deepseek-v4
deepseek_v4
development
mixed-precision
mlx
mtp
quantized
safetensors
text-generation

README

AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head is preserved at BF16 in the checkpoint (or a bound sidecar when present).

Checkpoint Tier 1 certified (experimental) on df-macstudio-m2 at Hub revision e22b117aa812b29943b160bb0fbf0b962d0d3819. Safetensors fingerprints are unchanged; this metadata repair is not a new MTP acceleration certificate.

Model details

PropertyValue
Base modeldeepseek-ai/DeepSeek-V4-Flash
Source revision60d8d70770c6776ff598c94bb586a859a38244f1
Product familydeepseek-v4
Source architectureDeepseekV4ForCausalLM (mixture of experts (MoE)); text path optimized
Main-model parameters284.33B logical parameters
QuantizerAXQuant 1.5.1
Hub budget class2bit
AXQuant base precision class2bit-experimental
Planned storage-adjusted BPW3.4232
Measured main-model BPW3.1329
Measured total BPW, including MTP3.1605
Safetensors weight size114.94 GB
Approximate complete download115.02 GB
Configured maximum context1,048,576 tokens; practical limits depend on unified memory
Primary MLX runtimeMLX-LM
AX Engine native executionDirect runtime smoke passed with AX Engine 6.15.0; MTP remains direct fallback
MTP presentTrue
Vision presentFalse
Audio presentFalse

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models. When that collapse happens, AutomatosX does not publish a separate misleading 4bit sibling for that base.

SiblingIntended trade-off
This 2bit packLowest-storage AXQ budget; check its exact BPW
4bit siblingHigher average precision near the 4-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP --local-dir ./AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP

Allow at least 115.02 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine and DeepSeek MTP status

The checkpoint Tier 1 certificate is bound to Hub revision e22b117aa812b29943b160bb0fbf0b962d0d3819; every Safetensors LFS fingerprint is unchanged on this metadata-only revision. AX Engine 6.15.0 passed direct load, chat, stream, and context-retrieval smoke tests on df-macstudio-m2 with AX_ENGINE_2BIT_EXPERIMENTAL=1. That checkpoint result does not certify speculative decode.

The packaged mtp.safetensors is the native DeepSeek V4 nextn sidecar, not a Qwen qwen3-next-mtp sidecar. AX Engine 7.1.5 recognizes this layout but keeps the product route on direct fallback until a revision-bound Tier 2 MTP acceptance, exactness, and speed certificate exists. Stock MLX-LM runs the backbone without activating the sidecar, and the oMLX/MTPLX Qwen import workflow does not apply. The internal DeepSeek MTP certification-candidate switch is for the formal harness, not normal serving.

Quantization layout

Main-weight precisionParametersShare
2bit278.11B95.59%
4bit3.64B1.25%
8bit529.53M0.18%
bf168.67B2.98%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32.
  • MTP sidecar: 1575 tensors, 6.61B parameters, 3.59 GB, BF16, F32, F8_E4M3, F8_E8M0, I8.
  • Vision sidecar: not included.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior
Calibrationnone; the allocation is based on architecture priors
Quantizer execution33492/33492 recorded module conversions succeeded; 0 fallbacks
AX Engine direct runtimePassed on df-macstudio-m2 with AX Engine 6.15.0
Quality versus BF16 or uniform baselinesNot published; no quality-retention claim
MTP acceptance and speednot measured; no MTP speedup claim
AX Engine kernel evidenceunmeasured
Vision-language qualityNot applicable (no vision tower in this package)
Speech-recognition qualityNot applicable
Long-context quality1,048,576-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified (experimental) at e22b117a; MTP Tier 2 not certified

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP uses the DeepSeek V4 nextn contract. Product serving remains direct fallback until a revision-bound Tier 2 certificate exists.

  • The configured context window can require substantially more memory as the KV cache grows.

  • Direct AX Engine runtime passed at the certificate revision; this package still has no in-repo native manifest, and MTP Tier 2 remains unverified.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the deepseek-ai/DeepSeek-V4-Flash model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

4 commits

AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP

Model

0

stars

4

commits

1

linked in READMEs

Aug 21, 2026

updated

2-bit
2bit
apple-silicon
axq
axquant
deepseek-v4
deepseek_v4
development
mixed-precision
mlx
mtp
quantized
safetensors
text-generation

README

AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP

An AXQuant (AXQ) mixed-precision MLX checkpoint for Apple Silicon, converted directly from the BF16 source model. The language path is quantized while the multi-token-prediction (MTP) head is preserved at BF16 in the checkpoint (or a bound sidecar when present).

Checkpoint Tier 1 certified (experimental) on df-macstudio-m2 at Hub revision e22b117aa812b29943b160bb0fbf0b962d0d3819. Safetensors fingerprints are unchanged; this metadata repair is not a new MTP acceleration certificate.

Model details

PropertyValue
Base modeldeepseek-ai/DeepSeek-V4-Flash
Source revision60d8d70770c6776ff598c94bb586a859a38244f1
Product familydeepseek-v4
Source architectureDeepseekV4ForCausalLM (mixture of experts (MoE)); text path optimized
Main-model parameters284.33B logical parameters
QuantizerAXQuant 1.5.1
Hub budget class2bit
AXQuant base precision class2bit-experimental
Planned storage-adjusted BPW3.4232
Measured main-model BPW3.1329
Measured total BPW, including MTP3.1605
Safetensors weight size114.94 GB
Approximate complete download115.02 GB
Configured maximum context1,048,576 tokens; practical limits depend on unified memory
Primary MLX runtimeMLX-LM
AX Engine native executionDirect runtime smoke passed with AX Engine 6.15.0; MTP remains direct fallback
MTP presentTrue
Vision presentFalse
Audio presentFalse

This repository contains MLX Safetensors. It does not contain PyTorch or GGUF weights.

Choosing an AXQ pack

AXQ names describe a storage-budget product class, not one uniform precision applied to every tensor. Protected tensors remain at higher precision, so the exact measured BPW is authoritative. In particular, a 6bit-named mixed plan may retain 4bit as its base precision while selecting 6-bit, 8-bit, or BF16 for other tensors to meet an approximately 6-BPW total budget. Protection floors can also raise a 4bit-named pack close to (or above) a 6bit budget on small or heavily protected models. When that collapse happens, AutomatosX does not publish a separate misleading 4bit sibling for that base.

SiblingIntended trade-off
This 2bit packLowest-storage AXQ budget; check its exact BPW
4bit siblingHigher average precision near the 4-BPW budget

See the AutomatosX MLX model catalog for related MLX and OptiQ alternatives.

Download

python -m pip install -U huggingface_hub
hf download AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP --local-dir ./AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP

Allow at least 115.02 GB of free disk space. Pin the resulting Hub commit in reproducible deployments rather than relying indefinitely on main.

Run with MLX-LM

python -m pip install -U mlx-lm
mlx_lm.generate \
  --model AutomatosX/AX-DeepSeek-V4-Flash-MLX-AXQ-2bit-MTP \
  --prompt "Explain mixed-precision quantization in three sentences." \
  --max-tokens 128 \
  --temp 0.0

MLX-LM compatibility covers standard text/backbone inference. It may ignore AXQuant runtime metadata and optional sidecars (vision.safetensors, mtp.safetensors); this command therefore does not establish MTP acceleration or vision-language quality. The artifact records MLX 0.32.0 and MLX-LM 0.31.3 from conversion.

AX Engine and DeepSeek MTP status

The checkpoint Tier 1 certificate is bound to Hub revision e22b117aa812b29943b160bb0fbf0b962d0d3819; every Safetensors LFS fingerprint is unchanged on this metadata-only revision. AX Engine 6.15.0 passed direct load, chat, stream, and context-retrieval smoke tests on df-macstudio-m2 with AX_ENGINE_2BIT_EXPERIMENTAL=1. That checkpoint result does not certify speculative decode.

The packaged mtp.safetensors is the native DeepSeek V4 nextn sidecar, not a Qwen qwen3-next-mtp sidecar. AX Engine 7.1.5 recognizes this layout but keeps the product route on direct fallback until a revision-bound Tier 2 MTP acceptance, exactness, and speed certificate exists. Stock MLX-LM runs the backbone without activating the sidecar, and the oMLX/MTPLX Qwen import workflow does not apply. The internal DeepSeek MTP certification-candidate switch is for the formal harness, not normal serving.

Quantization layout

Main-weight precisionParametersShare
2bit278.11B95.59%
4bit3.64B1.25%
8bit529.53M0.18%
bf168.67B2.98%
  • Quantization methods: affine, bf16.
  • Group sizes used by quantized assignments: 32.
  • MTP sidecar: 1575 tensors, 6.61B parameters, 3.59 GB, BF16, F32, F8_E4M3, F8_E8M0, I8.
  • Vision sidecar: not included.
  • Optimization scope: text-path.
  • Support tier: convertible.

BF16 sidecars, when present, are included in total download size. Their presence does not by itself establish MTP acceleration or vision-language quality.

Evidence and validation status

CheckStatus
Planning evidencearchitecture_prior
Calibrationnone; the allocation is based on architecture priors
Quantizer execution33492/33492 recorded module conversions succeeded; 0 fallbacks
AX Engine direct runtimePassed on df-macstudio-m2 with AX Engine 6.15.0
Quality versus BF16 or uniform baselinesNot published; no quality-retention claim
MTP acceptance and speednot measured; no MTP speedup claim
AX Engine kernel evidenceunmeasured
Vision-language qualityNot applicable (no vision tower in this package)
Speech-recognition qualityNot applicable
Long-context quality1,048,576-token capacity is config metadata, not a validated claim
Release certificationCheckpoint Tier 1 certified (experimental) at e22b117a; MTP Tier 2 not certified

Intended use and limitations

  • Intended for local development and evaluation on Apple Silicon with MLX-compatible runtimes.

  • No minimum unified-memory figure is claimed; loadability depends on model size, context length, KV-cache policy, runtime buffers, and other processes using unified memory.

  • Architecture-prior allocation is not measured sensitivity. It must not be presented as measured model quality.

  • MTP uses the DeepSeek V4 nextn contract. Product serving remains direct fallback until a revision-bound Tier 2 certificate exists.

  • The configured context window can require substantially more memory as the KV cache grows.

  • Direct AX Engine runtime passed at the certificate revision; this package still has no in-repo native manifest, and MTP Tier 2 remains unverified.

  • Upstream capabilities, limitations, biases, and responsible-use guidance still apply.

Provenance and audit files

All published provenance uses repository-relative paths. Local source paths are stripped before publication. The checkpoint was converted from BF16 rather than re-quantized from an OptiQ artifact. If an OptiQ repository is published separately, it uses a different quantizer and should not be assumed to have identical BPW or quality.

License

The checkpoint follows the upstream model license where applicable (often Apache License 2.0). See the deepseek-ai/DeepSeek-V4-Flash model card for license terms, model limitations, and responsible-use guidance.

Contributors

AutomatosX

4 commits